Systems and methods for multiplexed rendering of light fields
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-23
- Publication Date
- 2026-08-11
AI Technical Summary
光场数据的处理可能成为网络的瓶颈
Smart Images

Figure CN113826402B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application is a non-provisional filing of U.S. Provisional Patent Application Serial No. 62 / 823,714, entitled “System and Method for Multiplexed Rendering of Light Fields”, filed on March 26, 2019, and claims the benefit thereto under 35 U.S. SC § 119(e), which is incorporated herein by reference in its entirety. Background Technology
[0003] Many display technologies are emerging that enable adjustable-viewpoint viewing of video content. The light field used by such adjustable-viewpoint displays can include a large amount of data. The processing of this light field data can become a bottleneck for networks. Summary of the Invention
[0004] Example methods according to some embodiments may include: receiving a media manifest file identifying a plurality of representations of a multi-view video, at least a first representation of the plurality of representations comprising a first subsample of a view, and at least a second representation of the plurality of representations comprising a second subsample of a view different from the first subsample of the view; selecting a selected representation from the plurality of representations; retrieving the selected representation; and rendering the selected representation.
[0005] In some embodiments of the example method, each of the plurality of representations may have a corresponding view density oriented toward a particular corresponding direction.
[0006] In some embodiments of the example method, the media manifest file may identify the corresponding view density and the specific corresponding orientation for one or more of the plurality of representations.
[0007] In some embodiments of the example method, the two or more distinct subsamples of the view may differ at least with respect to the view density of the subsamples of the view oriented toward a particular direction.
[0008] Some embodiments of the example method may further include: tracking the user's viewing direction, wherein selecting a selected representation may include a view subsample of two or more subsamples of a selected view, the selected view subsample having a high view density toward the user's tracked viewing direction.
[0009] Some embodiments of the example method may also include: tracking the user's viewing direction, wherein selecting the selected representation may include selecting the selected representation based on the tracked user's viewing direction.
[0010] Some embodiments of the example method may further include: obtaining the user's viewing direction, wherein selecting the selection includes selecting the selection based on the obtained viewing direction of the user.
[0011] In some embodiments of the example method, the selection of the chosen representation may be based on the user's location.
[0012] In some embodiments of the example method, the selection of the chosen representation may be based on bandwidth constraints.
[0013] In some embodiments of the example method, at least one of the plurality of representations may include a view with a higher density for a first viewing direction than for a second viewing direction.
[0014] Some embodiments of the example method may further include: using the rendered representation to generate a signal for display.
[0015] Some embodiments of the example method may further include: tracking the user's head position, wherein the selection of the selected representation may be based on the user's head position.
[0016] Some embodiments of the example method may further include: tracking the user's gaze direction, wherein the selection of the selected representation may be based on the user's gaze direction.
[0017] Some embodiments of the example method may further include: using the user's gaze direction to determine the user's viewpoint, wherein selecting the selected representation may include selecting the selected representation based on the user's viewpoint.
[0018] Some embodiments of the example method may further include: determining the user's viewpoint using the user's gaze direction; and selecting at least one subsample of the view of the multiview video, wherein selecting the at least one subsample of the view may include: selecting at least one subsample of the view within a threshold viewpoint angle of the user's viewpoint.
[0019] Some embodiments of the example method may also include: interpolating at least one view, wherein selecting a selected representation can be chosen from the plurality of representations and the at least one view.
[0020] In some embodiments of the example method, the media manifest file may include priority data for one or more views, and the at least one view is interpolated using the priority data.
[0021] In some embodiments of the example method, the media manifest file may include priority data for one or more views, and the selection of the view indicates the use of the priority data.
[0022] Some embodiments of the example method may also include: obtaining light field content associated with a selected representation; decoding a frame of the light field content; combining two or more views represented in the frame to generate a composite view result; and rendering the composite view result to a display.
[0023] In some embodiments of the example method, the frame may include a frame-packed representation of two or more views corresponding to the selected representation.
[0024] In some embodiments of the example method, for at least one of the plurality of representations, the media manifest file may include information corresponding to two or more views of the light field content.
[0025] For some embodiments of the example method, the selected representation may be selected based on at least one of the following criteria: user head position, gaze tracking, complexity of the content, display capability, and bandwidth available for retrieving the light field content.
[0026] For some embodiments of the example method, selecting the selected representation may include: predicting the user's viewpoint; and selecting the selected representation based on the predicted user viewpoint.
[0027] Some embodiments of the example method may also include: obtaining light field content associated with the selected representation; generating a generated view of the light field content from the obtained light field content; and rendering the generated view to a display.
[0028] For some embodiments of the example method, the generated view of generating light field content may include: decoding a frame of the light field content; and combining two or more views represented in the frame to generate the generated view, wherein the obtained light field content may include the frame of the light field content.
[0029] Some embodiments of the example method may further include: decoding a frame of light field content; and combining two or more views represented in the frame to generate a combined view composition result, wherein the plurality of representations of the multi-view video may include plurality of subsamples of the views of the light field content, and wherein rendering the selected representation may include rendering the combined view composition result to a display.
[0030] Some embodiments of the example method may further include: requesting the media manifest file from a server; and requesting the light field content associated with a selected subset of views, wherein obtaining the light field content associated with the selected subset of views may include performing a process selected from the group consisting of: retrieving the light field content associated with the selected subset of views from the server, requesting the light field content associated with the selected subset of views from the server, and receiving the light field content associated with the selected subset of views.
[0031] For some embodiments of the example method, combining two or more views represented in the frame may include using view composition techniques.
[0032] For some embodiments of the example method, the selection of one of a plurality of subsamples of the view can be based on at least one of the following criteria: user head position, gaze tracking, complexity of the content, display capability, and bandwidth available for retrieving the light field content.
[0033] For some embodiments of the example method, selecting one of the plurality of subsamples of the view may include: predicting the user's viewpoint; and selecting the subset of views based on the predicted user viewpoint.
[0034] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0035] Example methods according to some embodiments may include: receiving a media manifest file, the media manifest file including information on a plurality of subsampled representations of views for light field video content; selecting one of the plurality of subsampled representations; obtaining the selected subsampled representation; interpolating the one or more views using the information in the manifest file that corresponds to one or more subviews from the selected subsampled representation; synthesizing one or more synthesized views from the one or more subviews; and displaying the one or more synthesized views.
[0036] Some embodiments of the example method may also include: estimating the bandwidth available for streaming the light field video content, such that the selection of the sub-sample representation among the plurality of sub-sample representations is based on the estimated bandwidth.
[0037] Some embodiments of the example method may also include: tracking the user's location such that the selection of the subsampled representation among the plurality of subsampled representations is based on the user's location.
[0038] Some embodiments of the example method may also include: requesting the light field video content from a server, wherein obtaining the selected subsample representation may include performing a process of selecting from a group consisting of: retrieving the selected subsample representation from the server, requesting the selected subsample representation from the server, and receiving the selected subsample representation.
[0039] For some embodiments of the example method, the information in the manifest file may include location data for two or more views.
[0040] For some embodiments of the example method, the information in the manifest file may include interpolation priority data for one or more of the plurality of views, and the selection of one of the plurality of subsampled representations may be based on the interpolation priority data for one or more of the plurality of views.
[0041] Some embodiments of the example method may also include: tracking the user's head position, wherein selecting one of a plurality of subsampled representations is based on the user's head position.
[0042] Some embodiments of the example method may also include: tracking the user's gaze direction, wherein selecting one of a plurality of sub-sample representations may be based on the user's gaze direction.
[0043] Some embodiments of the example method may further include: determining the user's viewpoint from the user's gaze direction; and selecting one or more subviews of the light field video content from a group including one or more interpolated views and selected subsampled representations, wherein synthesizing one or more composite views from the one or more interpolated subviews may include: synthesizing the one or more composite views of the light field using the one or more selected subviews and the user's viewpoint.
[0044] Some embodiments of the example method may also include: displaying the synthesized view of the light field.
[0045] For some embodiments of the example method, selecting one or more subviews of the light field may include selecting one or more subviews within a threshold viewpoint angle of the user's viewpoint.
[0046] Some embodiments of the example method may further include: determining the user's viewpoint from the user's gaze direction, wherein selecting one of the plurality of subsampled representations may include: selecting the subsampled representation based on the user's viewpoint.
[0047] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0048] Example methods according to some embodiments may include: receiving a media manifest file including information of a plurality of subsamples of a view for light field content; selecting one of the plurality of subsamples of the view; obtaining light field content associated with the selected view subsample; decoding a frame of the light field content; combining two or more views represented in the frame to generate a combined view composition result; and rendering the combined view composition result to a display, wherein the obtained light field content includes the frame of the light field content.
[0049] Some embodiments of the example method may further include: requesting the media manifest file from a server; and requesting the light field content associated with a selected subset of views, wherein obtaining the light field content associated with the selected subset of views may include performing a process selected from the group consisting of: retrieving the light field content associated with the selected subset of views from the server, requesting the light field content associated with the selected subset of views from the server, and receiving the light field content associated with the selected subset of views.
[0050] In some embodiments of the example method, the frame may include a frame-packed representation of two or more views corresponding to a selected subset of views.
[0051] In some embodiments of the example method, the media manifest file may include information corresponding to two or more views corresponding to a selected subset of views.
[0052] For some embodiments of the example method, for at least one of multiple subsamples of a view, the media manifest file may include information corresponding to two or more views of the light field content.
[0053] For some embodiments of the example method, selecting one of the plurality of subsamples of the view may include parsing the information of the plurality of subsamples of the view for light field content in the media manifest file.
[0054] For some embodiments of the example method, combining two or more views represented in the frame may include using view composition techniques.
[0055] For some embodiments of the example method, the selection of one of the plurality of subsamples of the view may be based on at least one of the following criteria: gaze tracking, complexity of the content, display capability, and bandwidth available for retrieving the light field content.
[0056] For some embodiments of the example method, selecting one of the plurality of subsamples of the view may include: predicting the user's viewpoint; and selecting the subset of views based on the predicted user viewpoint.
[0057] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0058] Example methods according to some embodiments may include: receiving a media manifest file including information of a plurality of subsamples of a view for light field content; selecting one of the plurality of subsamples of the view; obtaining light field content associated with the selected subset of the view; generating a view from the obtained light field content; and rendering the generated view to a display.
[0059] For some embodiments of the example method, generating a view from the obtained light field content may include: interpolating the views from the light field content associated with a selected subset of views to generate the generated view, the interpolation being performed using information in the manifest file corresponding to the views respectively.
[0060] For some embodiments of the example method, generating one or more views from the obtained light field content may include: decoding a frame of the light field content; and combining two or more views represented in the frame to generate the generated view, wherein the obtained light field content may include the frame of the light field content.
[0061] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0062] Example methods according to some embodiments may include: receiving a media manifest file identifying a plurality of subsamples of views of a multi-view video, the plurality of subsamples of the views including two or more views of different densities; selecting a selected subsample from the plurality of subsamples of the views; retrieving the selected subsample; and rendering the selected subsample.
[0063] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0064] Example methods according to some embodiments may include: rendering a representation of a view comprising a full array of light field video content; sending the rendered full array representation of the view; obtaining the current viewpoint of a client; using the current viewpoint and a viewpoint motion model to predict future viewpoints; prioritizing a plurality of subsampled representations of the view of the light field video content; rendering the prioritized plurality of subsampled representations of the view of the light field video content; and sending the prioritized plurality of subsampled representations of the view.
[0065] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0066] Example methods according to some embodiments may include: selecting a plurality of subviews of light field video content; generating streaming data for each of the plurality of subviews of the light field video content; and generating a media manifest file including the streaming data for each of the plurality of subviews of the light field video content.
[0067] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0068] Example methods according to some embodiments may include: receiving a request for information about light field video content; when the request is a new session request, sending a media manifest file that includes information on a plurality of subsampled representations of a view of the light field video content; and when the request is a subset data fragment request, sending a data fragment comprising a subset of the light field video content.
[0069] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0070] An example signal according to some embodiments may include a signal carrying a view representation of a full array of light field video content and a plurality of subsampled representations of the view of the light field video content.
[0071] Example signals according to some embodiments may include a signal carrying multiple sub-views of light field video content.
[0072] Example signals according to some embodiments may include a signal carrying streaming data of each of a plurality of sub-views of light field video content.
[0073] An example signal according to some embodiments may include a signal that carries information about a plurality of subsampled representations of a view of light field video content.
[0074] Example signals according to some embodiments may include a signal carrying a data segment comprising a subset of light field video content. Attached Figure Description
[0075] Figure 1A This is a system diagram of an example system of an example communication system according to some embodiments.
[0076] Figure 1B This is a system diagram of an example system according to some embodiments, illustrating that it can be... Figure 1A Example wireless transmit / receive unit (WTRU) used in the communication system shown.
[0077] Figure 2 This is a schematic diagram illustrating an example full light field rendered using a 5×5 subview according to some embodiments.
[0078] Figure 3 This is a schematic diagram illustrating an example subsampling using a 3×3 subview according to some embodiments.
[0079] Figure 4 This is a schematic diagram illustrating an example 5×5 subview that is assigned priority based on an estimated viewpoint according to some embodiments.
[0080] Figure 5 This is a schematic diagram illustrating an exemplary adjustment of sub-viewpoint priority based on an updated viewpoint motion model according to some embodiments.
[0081] Figure 6A It is shown that, according to some embodiments, the corresponding Figure 2 An illustration of an exemplary full-image 5×5 array light field configuration.
[0082] Figure 6B It is shown that, according to some embodiments, it corresponds to Figure 3 An illustration of an example of a 3×3 array optical field configuration with uniform subsampling.
[0083] Figure 6C It is shown that, according to some embodiments, the corresponding Figure 4 An illustration of an example center-biased sub-sampling 3×3 array optical field configuration.
[0084] Figure 6D It is shown that, according to some embodiments, it corresponds to Figure 5 An illustration of an exemplary left-biased sub-sampling 3×3 array optical field configuration.
[0085] Figure 7 This is a schematic diagram illustrating an example numbered configuration of a light field subview according to some embodiments.
[0086] Figure 8 This is an illustration of an example MPEG-DASH Media Presentation Description (MPD) file according to some embodiments.
[0087] Figure 9 This is an illustration of an example MPD file with light field configuration elements according to some embodiments.
[0088] Figure 10 This is a system diagram illustrating a set of example interfaces for multiplexed light field rendering according to some embodiments.
[0089] Figure 11 This is a message sequence diagram illustrating an example process for multiplexed light field rendering using an example client to pull a model, according to some embodiments.
[0090] Figure 12 This is a message sequence diagram illustrating an example process for reusing light field rendering using an example server to push a model, according to some embodiments.
[0091] Figure 13 This is a message sequence diagram illustrating an example process for multiplexed light field rendering using an example client to pull a model, according to some embodiments.
[0092] Figure 14 This is a flowchart illustrating an example process for a content server according to some embodiments.
[0093] Figure 15 This is a flowchart illustrating an example process for a viewing client according to some embodiments.
[0094] Figure 16 This is a schematic diagram showing an example light field rendered in a 5×5 subview according to some embodiments.
[0095] Figure 17 This is a schematic diagram illustrating an example light field rendered with a 3×3 subview according to some embodiments.
[0096] Figure 18 This is a schematic diagram illustrating an example light field rendered with a 2×3 subview according to some embodiments.
[0097] Figure 19A It is shown that, according to some embodiments, it corresponds to Figure 16 A diagram of an example 5×5 array light field configuration.
[0098] Figure 19B It is shown that, according to some embodiments, it corresponds to Figure 17 A diagram of an example 3×3 array light field configuration.
[0099] Figure 19C It is shown that, according to some embodiments, it corresponds to Figure 18 A diagram of an example 2×3 array light field configuration.
[0100] Figure 20 This is a schematic diagram illustrating an example subsampling using a 3×3 subview according to some embodiments.
[0101] Figure 21 This is a schematic diagram illustrating an example numbered configuration of a light field subview according to some embodiments.
[0102] Figure 22 This is a diagram illustrating an example MPEG-DASH Media Presentation Description (MPD) file according to some embodiments.
[0103] Figure 23 This is a diagram illustrating an example MPD file with light field configuration elements according to some embodiments.
[0104] Figure 24 This is a system diagram illustrating a set of example interfaces for adaptive optical field streaming according to some embodiments.
[0105] Figure 25 This is a message sequence diagram illustrating an example process for adaptive optical field streaming using estimated bandwidth and view interpolation, according to some embodiments.
[0106] Figure 26 This is a message sequence diagram illustrating an example process for adapting light field streaming using predicted view locations, according to some embodiments.
[0107] Figure 27 This is a flowchart illustrating an example process for content server preprocessing according to some embodiments.
[0108] Figure 28 This is a flowchart illustrating an example process for runtime processing of a content server according to some embodiments.
[0109] Figure 29 This is a flowchart illustrating an example process for a viewing client according to some embodiments.
[0110] Figure 30This is a flowchart illustrating an example process for a viewing client according to some embodiments.
[0111] Figure 31 This is a flowchart illustrating an example process for a viewing client according to some embodiments.
[0112] Figure 32 This is a flowchart illustrating an example process for a viewing client according to some embodiments.
[0113] The entities, connections, arrangements, etc., depicted and described in connection with the various figures are given as examples rather than as limitations. Therefore, any and all statements or other indications regarding what a particular figure “depicts,” what a particular element or entity in a particular figure “is” or “has,” and any and all similar statements—which may be isolated and interpreted as absolute and therefore limiting outside the context—may be properly interpreted only if they are preceded by a constructive prefix such as “in at least one embodiment, …”. For the sake of brevity and clarity, this implied prefix is not repeated in the detailed description.
[0114] Example network for implementation of the embodiments
[0115] In some embodiments described herein, the wireless transmit / receive unit (WTRU) can be used as, for example, a viewing client.
[0116] Figure 1A This diagram illustrates an example communication system 100 that can implement one or more of the disclosed embodiments. The communication system 100 can be a multiple access system providing voice, data, video, messaging, broadcasting, and other content to multiple wireless users. The communication system 100 enables multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 100 can use one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT-Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtering OFDM, and Filter Bank Multicarrier (FBMC), etc.
[0117] like Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106, public switched telephone network (PSTN) 108, Internet 110, and other networks 112. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network components. Each WTRU 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, any WTRU 102a, 102b, 102c, or 102d may be referred to as a “station” and / or “STA”, and may be configured to transmit and / or receive wireless signals. It may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain environments), consumer electronics devices, and devices operating on commercial and / or industrial wireless networks, etc. Any of WTRU 102a, 102b, 102c, or 102d may be interchangeably referred to as a UE.
[0118] The communication system 100 may also include base stations 114a and / or 114b. Each base station 114a, 114b may be any type of device configured to enable its access to one or more communication networks (e.g., CN 106, Internet 110, and / or other networks 112) by wirelessly interfacing with at least one of WTRUs 102a, 102b, 102c, 102d. For example, base stations 114a, 114b may be base transceiver stations (BTS), node B, e-node B, home node B, home e-node B, gNB, NR node B, site controller, access point (AP), and wireless routers, etc. Although each base station 114a, 114b is described as a single component, it should be understood that base stations 114a, 114b may include any number of interconnected base stations and / or network components.
[0119] Base station 114a may be part of RAN 104 / 113, and the RAN may also include other base stations and / or network components (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies called cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide radio service coverage for a specific geographic area that is relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, that is, each transceiver corresponds to one sector of the cell. In embodiments, base station 114a may use multiple-input multiple-output (MIMO) technology and may use multiple transceivers for each sector of the cell. For example, by using beamforming, signals can be transmitted and / or received in a desired spatial direction.
[0120] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116, wherein the air interface can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 can be established using any suitable radio access technology (RAT).
[0121] More specifically, as described above, the communication system 100 can be a multiple access system and can use one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA, etc. For example, base station 114a in RAN 104 / 113 and WTRUs 102a, 102b, and 102c can implement a certain radio technology, such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), wherein the technology can use Wideband CDMA (WCDMA) to establish the air interface 116. WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0122] In an embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement a certain radio technology, such as Evolved UMTS Terrestrial Radio Access (E-UTRA), wherein the technology may use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTA Pro (LTE-A Pro) to establish air interface 116.
[0123] In an embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement a certain radio technology, such as NR radio access, wherein the radio technology may use a novel radio (NR) to establish air interface 116.
[0124] In this embodiment, base station 114a and WTRUs 102a, 102b, and 102c can implement various radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can jointly implement LTE radio access and NR radio access (e.g., using the dual connectivity (DC) principle). Therefore, the air interface used by WTRUs 102a, 102b, and 102c can be characterized by various types of radio access technologies and / or transmissions sent to / from various types of base stations (e.g., eNBs and gNBs).
[0125] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement the following radio technologies, such as IEEE 802.11 (i.e., WiFi), IEEE 802.16 (Global Microwave Access Interoperability (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN), etc.
[0126] Figure 1ABase station 114b can be a wireless router, home node B, home e node B, or access point, and can use any suitable RAT to facilitate wireless connectivity in a local area, such as a business premises, residence, vehicle, campus, industrial facility, air corridor (e.g., for use by drones), and road, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can establish a wireless local area network (WLAN) by implementing radio technology such as IEEE 802.11. In another embodiment, base station 114b and WTRUs 102c, 102d can establish a wireless personal area network (WPAN) by implementing radio technology such as IEEE 802.15. In yet another embodiment, base station 114b and WTRUs 102c, 102d can establish a picocell or femtocell by using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). Figure 1A As shown, base station 114b can be directly connected to the Internet 110. Therefore, base station 114b does not need to access the Internet 110 via CN 106.
[0127] RAN 104 / 113 can communicate with CN 106, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more WTRUs 102a, 102b, 102c, 102d. This data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements, etc. CN 106 can provide call control, billing services, location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or can perform advanced security functions such as user authentication. Although in Figure 1A While not shown, it should be understood that RAN 104 / 113 and / or CN 106 can communicate directly or indirectly with other RANs that use the same RAT or a different RAT as RAN 104 / 113. For example, in addition to connecting with RAN 104 / 113 using NR radio technology, CN 106 can also communicate with other RANs (not shown) using GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technologies.
[0128] CN 106 can also act as a gateway for WTRUs 102a, 102b, 102c, and 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Simple Old-Style Telephone Service (POTS). The Internet 110 may include a globally interconnected computer network equipment system using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the other network 112 may include another CN connected to one or more RANs, wherein the one or more RANs may use the same RAT or a different RAT as RAN 104 / 113.
[0129] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multi-mode capability (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers communicating with different wireless networks on different wireless links). For example... Figure 1A The WTRU 102c shown can be configured to communicate with base station 114a, which can use cellular-based radio technology, and with base station 114b, which can use IEEE 802 radio technology.
[0130] Figure 1B This is a system diagram illustrating an example of WTRU 102. (See diagram below.) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive unit 122, a speaker / microphone 124, a keyboard 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and other peripheral devices 138. It should be understood that, while remaining consistent with the embodiments, WTRU 102 may also include any sub-combination of the foregoing components.
[0131] Processor 118 can be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. Processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other function that enables WTRU 102 to operate in a wireless environment. Processor 118 can be coupled to transceiver 120, and transceiver 120 can be coupled to transmitting / receiving unit 122. Although Figure 1B While the processor 118 and transceiver 120 are described as separate components, it should be understood that the processor 118 and transceiver 120 can also be integrated into a single electronic component or chip.
[0132] Transmit / receive component 122 may be configured to transmit or receive signals to or from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmit / receive component 122 may be an antenna configured to transmit and / or receive RF signals. As an example, in an embodiment, transmit / receive component 122 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals. In an embodiment, transmit / receive component 122 may be configured to transmit and / or receive RF and optical signals. It should be understood that transmit / receive component 122 may be configured to transmit and / or receive any combination of wireless signals.
[0133] Although Figure 1B The transmit / receive component 122 is described as a single component, but the WTRU 102 may include any number of transmit / receive components 122. More specifically, the WTRU 102 may use MIMO technology. Thus, in an embodiment, the WTRU 102 may include two or more transmit / receive components 122 (e.g., multiple antennas) that transmit and receive radio signals via the air interface 116.
[0134] Transceiver 120 can be configured to modulate signals to be transmitted by transmitter / receiver 122 and demodulate signals received by transmitter / receiver 122. As described above, WTRU 102 can have multimode capability. Therefore, transceiver 120 can include multiple transceivers that allow WTRU 102 to communicate using various RATs (e.g., NR and IEEE 802.11).
[0135] The processor 118 of WTRU 102 can be coupled to a speaker / microphone 124, a keyboard 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit) and can receive user input data from these components. The processor 118 can also output user data to the speaker / microphone 124, keyboard 126, and / or display / touchpad 128. Furthermore, the processor 118 can access and store information from any suitable memory, such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 can include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 can include a subscriber identification module (SIM) card, a memory stick, a secure digital card (SD) memory card, etc. In other embodiments, the processor 118 can access and store information from memory that is not actually located in WTRU 102; for example, such memory could be located in a server or home computer (not shown).
[0136] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control power for other components in the WTRU 102. The power supply 134 can be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell battery packs (such as nickel-cadmium (Ni-Cd), nickel-zinc (Ni-Zn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, and fuel cells, etc.
[0137] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) related to the current location of the WTRU 102. As a supplement or replacement to the information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via the air interface 116, and / or determine its location based on signal timing received from two or more nearby base stations. It should be understood that, while remaining consistent with the embodiments, the WTRU 102 may acquire location information using any suitable positioning method.
[0138] The processor 118 can also be coupled to other peripheral devices 138, which may include one or more software and / or hardware modules providing additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or video), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, etc. Modules, FM radio units, digital music players, media players, video game console modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, and activity trackers, etc. Peripheral devices 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors, geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.
[0139] WTRU 102 may include a full-duplex wireless device, wherein the reception or transmission of some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous for the wireless device. The full-duplex wireless device may include an interference management unit that reduces and / or substantially eliminates self-interference by means of hardware (e.g., choke coils) or by means of a processor (e.g., a separate processor (not shown) or by means of processor 118) for signal processing. In embodiments, WTRU 102 may include a half-duplex wireless device that transmits and receives some or all signals (e.g., associated with specific subframes for UL (e.g., for transmission) or downlink (e.g., for reception).
[0140] Given Figure 1A-1B and Figure 1A-1B The corresponding descriptions may be performed by one or more emulation devices (not shown) that perform one or more or all of the functions described herein with respect to: WTRU 102a-d, base stations 114a-b, and / or any other devices (one or more) described herein. The emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation device may be used to test other devices and / or simulate network and / or WTRU functions.
[0141] The simulation device may be designed to perform one or more tests on other devices in a laboratory environment and / or a carrier network environment. For example, the one or more simulation devices may perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. The one or more simulation devices may perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The simulation device may be directly coupled to another device for testing purposes and / or may use over-the-air wireless communication to perform tests.
[0142] The one or more simulation devices may perform one or more functions (including all) when implemented / deployed not as part of a wired and / or wireless communication network. For example, the simulation devices may be used in test scenarios in test labs and / or non-deployed (e.g., testing) wired and / or wireless communication networks to test one or more components. The one or more simulation devices may be test equipment. The simulation devices may transmit and / or receive data using direct RF coupling and / or wireless communication via RF circuitry (e.g., which may include one or more antennas). Detailed Implementation
[0143] Light fields generate a large amount of data, which can include a description of the amount of light emanating from points in space and flowing in a set of directions. High-fidelity light fields, as representations of 3D scenes, can contain a vast amount of data. To support real-time transmission and visualization, efficient data distribution optimization methods may be needed, and the amount of light field data to be rendered and transmitted may need to be reduced.
[0144] To compress traditional 2D video, various lossless and lossy bitrate reduction and compression methods have been developed. Besides current spatiotemporal compression methods, one class of bitrate reduction methods transmits portions of information that are time-multiplexed. For CRT displays, multiplexing is widely used to simulate the interlaced image lines format in TV transmission.
[0145] Another class of compression algorithms consists of various prediction methods, which can generally be applied similarly to both the transmission side (encoder or server) and the receiving side (decoder). These predictions can be spatial (intra-frame) or temporal (inter-frame). Some of the methods mentioned above have also been applied to light fields. As an example, the following journal article describes how the subjective quality of light field rendering is affected by simple quality switching and stopping (frame freeze) methods, as well as the balance between related trade-offs between transmission bit rate, light field angular resolution, and spatial resolution: Kara, Peter A. et al., Evaluation of the Concept of Dynamic Adaptive Streaming of Light Field Video, IEEE T RANSACTIONS ON B ROADCASTING (2018).
[0146] Typically, methods for applying predictive coding to real-time transmission of light fields remain rare. Examples of light field compression are discussed in the following journal article: Ebrahimi, Touradj et al., JPEG Pleno: Toward an Efficient Representation of Visual Reality, 23:4 IEEE M ULTIMEDIA 14-20 (October-December 2016). This article describes how existing multi-view coding methods (e.g., MPEG HEVC or 3D HEVC) can be used for light field compression. 3D HEVC is an extension of HEVC for supporting depth images.
[0147] H.264MVC and its later counterpart, MFPEG HEVC (H.265), and their derivatives support several important 3D functions. These include viewing content on an external 3D display, which projects a set of viewpoints onto different angular directions in space. For this purpose, the entire 3D information can typically be sent and decoded. The aforementioned standards also support dynamic viewpoints and motion parallax using traditional 2D displays, including HMDs, but in these applications, receiving a complete set of spatial views from one or a few time-varying user viewpoints may not be optimal.
[0148] The H.264MVC, HEVC, and 3D HEVC standards can be used to encode light field data. The traditional and commonly used light field format is the view matrix (integral format), which represents different viewpoints of a scene as a matrix / tessellation of 2D views from neighboring viewpoints. For example, the HEVC standard supports up to 1024 multiple views, which can be applied to light fields represented by multiple sub-views.
[0149] MPEG video codecs can be used to compress light fields in multiview formats. In these codecs, the Network Abstraction Layer (NAL) defines the upper-layer data structure, but also imposes limitations on the utilization of subview redundancy over time. Specifically, according to the following journal article, this NAL structure does not allow prediction of a subview image at a given time from another subview image at different times: Vetro, Anthony et al., Overview of the Stereo and Multiview Video Coding Extensions of the H.264 / MPEG-4 AVC Standard, 99:4P ROCIEEE 1-15 (April 2011) (“Vetro”). Furthermore, for backward compatibility reasons, Vetro is understood to mean that compressed MVC multiple views must include the base view bitstream. In use cases where the base view is not used or is only used occasionally, this leads to excessive bandwidth usage during transmission.
[0150] In existing multi-view coding standards, such as Sullivan, Gary J. et al., Standardized Extensions of High Efficiency Video Coding (HEVC), 7:6 IEEE JS ELECTED T OPICS IN S IGNAL P ROC As discussed in .1001-16 (December 2013), all views typically need to be decoded even when viewing a specific subview. Accordingly, decoding only one subview at a time may not be possible unless that subview is a mandatory base view.
[0151] Journal article Kurutepe, Engin et al., Client-Driven Selective Streaming of Multiview Video for Interactive 3DTV, 17:11 IEEE T RANS.ON C IRCUITS AND S YSTEMS FOR V IDEO T ECHNOLOGY 1558-65 (2007) (“Kurutepe”) proposed a modification to the MVC structure to allow individual views of a multi-view video to be distributed as enhancement layers. The enhancement layers are linked to the MVC base layer, which contains low-bitrate spatially downsampled frames for all views of the multi-view video. Thus, the proposed alternative MVC structure is understood to distribute all views, but only as low-quality versions, whereas, in contrast to the standard MVC, individual views or pairs of views in high-resolution versions can be selectively downloaded as enhancement layers. In addition to proposing the modification to the MVC structure, Kurutepe also described prefetching views based on the predicted future user head position inferred from collected head tracking data.
[0152] For the distribution of large data files, the end-to-end (P2P) distribution model offers a robust alternative to the strict client-server model, reducing server connection and bandwidth requirements as clients share portions of the data they have downloaded among all other clients. P2P distribution has also been considered for visual content distribution. The following article proposes an adapted P2P multi-view content distribution solution: Ozcinar et al., Adaptive 3D Multi-View Video Streaming Over P2P Networks, C ONFERENCE P ROCEEDINGS OF 2014 IEEE I NTERNATIONAL C ONFERENCE ON I MAGE P ROCESSING (ICIP) 2462-66(2014) (“Ozcinar”) and Gürler, C. Goktug & Tekalp, Murat, Peer-to-Peer System Design for Adaptive 3D Video Streaming, 51:5 IEEE C OMM .108-114 (2013) (“Gürler”). In the proposed method, all clients connected to the streaming session are connected to a mesh network, where each client downloads missing content fragments from the server or any other client in the mesh network, while also allowing other clients to download fragments they have already downloaded. A tracker is used to collect and share information about the available content fragments on each client. Gürler proposes two variations of the multiview adaptation method, both of which perform adaptation by asymmetrically reducing the image quality between individual subviews of the multiview content following observations from previous studies on asymmetric image quality of stereoscopic views. Ozcinar and Gürler both discuss the distribution of depth information along with color images of the multiview content. Ozcinar proposes the distribution of additional metadata that describes which multiview images to drop in advance under network congestion, as a practical mechanism for handling adaptation on the client side. Experimental results show that the robustness of P2P multiview streaming using the proposed adaptation scheme is significantly increased under congestion.
[0153] In addition to camera arrays that produce high-resolution subviews, lens array optics can be used in capturing light field images in microlens format. If the light field is in microlens format, the content can be converted to a multiview light field (see JPEG Pleno) before compression and can be converted back to the microlens format used by the receiver.
[0154] The resolution and frame rate of a video sequence can be adapted to enable content streaming via the DASH protocol. Many other devices do not adapt the angle of the light field content in the video stream. Many other devices also do not change the number of views in the multi-view representation bundled into frames.
[0155] Multiplexed light field rendering
[0156] Using many existing multi-view coding standards to support dynamic user viewpoints and motion parallax can often lead to excessive bandwidth usage, especially when only a single view or stereoscopic view for display is generated in a single time step. Figure 2 Applications of 2D images. Examples of such applications include 3DoF+ or 6DoF applications based on viewing content on an HMD or other 2D display.
[0157] To optimize rendering and data distribution, a subset of the complete overall light field data can be generated and transmitted, and the receiver can synthesize additional views from the transmitted views. Furthermore, the impact of the selection of subviews for rendering and transmission on perceived image quality can guide the subview rendering process on the server side or the client side to ensure quality of experience.
[0158] In some embodiments, data transmission can be optimized using multiplexed rendering of the light field (which can be done at the content server) and view composition (which can be done at the viewing client), which can mitigate the temporal variation of the light field data caused by multiplexing.
[0159] The viewing client can send its current viewpoint to a content server or analyze content locally and request specific light field subviews. A viewpoint motion estimation model can be used to predict the client's viewpoint for future time steps. Based on this viewpoint prediction, subviews of the overall light field can be prioritized. Priority controls the rendering and transmission of each subview of the overall light field.
[0160] In some embodiments, the process of multiplexing light field data can be changed from transmitting the entire light field image at each time step to transmitting a subset of the subviews. This reduces the content delivery bandwidth used. Temporal multiplexing of subviews can be detected in the transmission of data and associated metadata (e.g., streaming subview and virtual camera specifications and timestamps) and in the current view signaling at the client.
[0161] According to some embodiments, server-side processing and content delivery bandwidth requirements can be reduced by dynamically limiting the number of light field subviews that are rendered and delivered. By prioritizing the rendering of subviews that maintain the image quality perceived by the viewer on the client side through content analytics, the content server can render individual light field subviews sequentially, rather than rendering the entire overall image at each time step in a sequential manner.
[0162] In some embodiments, the viewer uses an HMD or similar mobile device that provides a monocular or stereoscopic view by synthesizing a new view onto the light field data based on tracking and user input.
[0163] Some embodiments use, for example, an example server-push model or an example client-pull model for streaming. In the server-push model, for example, the example server process may prioritize subview rendering and streaming based on the viewpoint signaled by the client, and if the subviews of the overall light field are time-varying, the example client process may synthesize the viewpoint based on the light field data. In the client-pull model, for example, similar to MPEG-Dash model operation, the server may provide the client with a manifest file indicating versions of the array of images for different viewing positions, and in some embodiments, the client may: determine the number of subviews based on bandwidth and the viewer's position for the set of subviews described in the manifest; prioritize the subviews; and pull data from the content server according to priority.
[0164] Utilizing existing multi-view coding standards to support dynamic viewpoints and motion parallax can often lead to excessive bandwidth usage, especially in applications that require only one viewpoint at a time. Such applications may include, for example, certain 3DoF+ or 6DoF applications, which are based on viewing real-time or stored content on an HMD or other 2D display. The systems and methods described herein, according to some embodiments, can avoid existing limitations by taking into account the impact of the current and predicted user viewpoint, as well as the selected rendered and transported subviews, on the perceived image quality.
[0165] Significant bitrate savings can typically be achieved by tracking the user and predicting their movement and viewpoint. In the following article, temporal prediction of the user's viewpoint in a real-time 3D capture scene resulted in approximately 70% bitrate savings (compared to an average of 2-3 simultaneous captures for encoding and transmission if viewpoint prediction is used instead of eight 3D captures per scene): Yang, Zhengyu et al., Enabling Multi-party 3D Tele-immersive Environments with ViewCast, 6:2ACM TRANS.ON M ULTIMEDIA C OMPUTING C OMM ' S And A PPL ' S 111-139 (March 2010) (“Yang”). Yang provides some examples of user viewpoint prediction techniques. According to some embodiments, the example methods and systems disclosed herein apply, for example, user viewpoint prediction to, for example, light field rendering. Yang is not to be construed as applying user viewpoint prediction to light field rendering.
[0166] Some embodiments reduce the number of subviews transmitted and rendered through content analysis and reused light field rendering. In some embodiments, the content server prioritizes full-light field subviews by analyzing the content based on content features, estimated user viewpoint location, and the prediction accuracy of the estimated viewpoint location. This, in turn, determines the order and frequency in which the subview is rendered and submitted. Dynamic selection of the set of view subsampling locations allows for the allocation of greater view density near the predicted viewpoint, enabling improved interpolation of views near the predicted viewpoint as rendered frames arrive at the viewing client and are displayed. Coarser view sampling can be used at locations far from the signaled viewpoint to allow data to be displayed for unfocused areas or focused areas (not the predicted viewpoint) due to inaccuracies in viewpoint motion estimation. Figures 2 to 5 An exemplary priority is shown for assigning to a light field sub-image.
[0167] Figure 2 This is a schematic diagram illustrating an example full light field rendered using a 5x5 subview according to some embodiments. Figure 2 In this example, a 5×5 virtual camera array of subview 202 is used to render an exemplary full light field 200. Figure 2 It shows Figure 3-5 The example full-field 200 on which the subsampling example is based.
[0168] Figure 3 This is a schematic diagram illustrating an example subsampling of 3×3 subviews according to some embodiments. To reduce bandwidth, a 3×3 array of images can be selected from the exemplary full-light field 300 for transmission to the client. In some embodiments, the server can prioritize subviews using estimated viewpoint positions predicted by a viewpoint motion model. Figure 3 An exemplary 3×3 array of subviews 302, 304, 306, 308, 310, 312, 314, 316, and 318 selected in a checkerboard pattern is shown, wherein the selected subviews are shown in solid black ellipses.
[0169] According to some embodiments, subsampling a view can often reduce transmission bandwidth by utilizing view interpolation at the receiver to generate additional views. In some embodiments, the location of the subsampled views is specified relative to the full grid. Different subsampling priorities can be assigned based on user gaze. The priority of each light field subview can be determined based on content analysis and contextual information (e.g., user / gaze position and display capabilities). The individual sets of subviews can be generated and / or transmitted sequentially based on said priorities. On the client side, a cache can receive subviews, and the client can use temporal coherent interpolation of the subviews to synthesize a viewpoint, e.g., a new viewpoint.
[0170] Figure 4 This is a schematic diagram illustrating an example 5×5 subview that is assigned priority based on an estimated viewpoint according to some embodiments. Figure 5 This is a schematic diagram illustrating exemplary adjustments to sub-viewpoint priorities based on an updated viewpoint motion model according to some embodiments. For some embodiments, in Figure 4 and Figure 5 In the middle, the subviews 418 and 518 that are closest to the estimated viewpoint are assigned the highest priority (marked with a solid white ellipse).
[0171] Using the current velocity of the viewpoint motion, the accuracy of the viewpoint position estimate is determined, and subviewpoint priorities are set. In some embodiments, subview priority can be an indicator of the extent of neighboring subviews to be rendered. In some embodiments, visual differences between previously rendered subviews can be used to weight the priorities. Because small shifts in viewpoint position can cause visual differences (e.g., specular reflections or large depth changes in content), denser sampling around the estimated most probable viewpoint position can be used, and subview priorities can be assigned accordingly. In some embodiments, the analysis process can use these metrics and assign a second-highest priority to the region around the most probable viewpoint with appropriate sampling density. Figure 4 and 5 In the example shown, frames assigned the second highest priority are marked as dashed black ellipses 410, 412, 414, 416, 510, 512, 514, 516.
[0172] Based on the example, for any given viewpoint, the region outside the focal region (in this example, the regions with the highest and second highest priorities) is assigned the third highest priority and rendered with a lower sampling density. As a result, most of the entire light field can be rendered with a lower sampling density. Figure 4 and 5 In the example shown, the subviews assigned the third highest priority are marked as solid black ellipses 402, 404, 406, 408, 502, 504, 506, and 508.
[0173] Figure 5 This shows that the user's viewpoint is estimated to be higher than... Figure 4 The viewpoint is shifted to the left. Subviews 510, 512, 514, and 516, directly adjacent to the estimated user viewpoint, are assigned the second highest priority. (Comparison) Figure 4 and Figure 5 The second high-priority subview set 410, 412, 414, 416, 510, 512, 514, and 516 also shift to the left, corresponding to the estimated user viewpoint shifting to the left.
[0174] In some embodiments, the example rendering process may use an assigned priority to determine how often to render a subview (or, for example, update the rendering of a subview). For example, the rendering process may send the rendered view along with a timestamp to the client. This timestamp may indicate the synchronization time between the viewing client and the virtual camera used by the subview. For some embodiments, the light field compression and distribution process or apparatus may use a prioritized subset of views (e.g., a 3×3 array) and the original position (or, for some embodiments, the relative position) of each subview within a larger array (e.g., a 5×5 array).
[0175] In some embodiments, the example viewing client process may include: determining the user's viewpoint from the user's gaze direction, and selecting a subsampled representation of the light field content based on the user's viewpoint. In some embodiments, the example viewing client process may include determining the user's viewpoint from the user's gaze direction, such that selecting one of the plurality of subsampled representations may include selecting the subsampled representation based on the user's viewpoint. For example, the center point of the subsampled representation may be selected as being within a threshold angle or distance of the user's viewpoint.
[0176] Figure 6A It is shown that, according to some embodiments, the corresponding Figure 2 An illustration of an exemplary 5×5 full-image array light field configuration. For some embodiments, a manifest file such as a Media Presentation Description (MPD) (e.g., a media manifest file) may include details of the full (or entire) light field array, such as the number of views, the indices of these views, and the sampling positions of the views within the adapter set. For example, the full light field array may have N×M view positions. Figure 6A An example naming convention for a 5x5 array 600 with a full light field is shown. The first number of each coordinate position indicates the row, and the second number indicates the column. Position (1, 1) is at the top left corner, and position (5, 5) is at the bottom right corner. Of course, this is just an example, and other embodiments may use different naming conventions, such as... Figure 7 The example shown.
[0177] Figure 6B It is shown that, according to some embodiments, it corresponds to Figure 3 A diagram of an example uniform subsampling 3×3 array optical field configuration. Figure 6B The 3×3 array 620 shown corresponds to Figure 3 A solid black ellipse in the middle. For example, Figure 6B A uniform subsampled 3×3 array represents a subview (5, 3) at the lower center position. This subview corresponds to Figure 3 The black ellipse in the fifth (bottom) row and the third (center) column.
[0178] Figure 6C It is shown that, according to some embodiments, the corresponding Figure 4 The diagram shows an example of a 3×3 array optical field configuration with an eccentric sub-sampling center. Figure 6C The 3×3 array 640 shown corresponds to Figure 4 The ellipse in the middle, for example, Figure 6C The center-eccentric sub-sampling 3×3 array indicates the subview (3, 4) at the middle right position. This subview corresponds to... Figure 4 The black dashed ellipse in the third (middle) row and the fourth (middle right) column.
[0179] Figure 6D It is shown that, according to some embodiments, it corresponds to Figure 5 An illustration of an exemplary left-biased 3×3 array optical field configuration. Figure 6D The 3×3 array 660 shown corresponds to Figure 5 An ellipse in the image. For example... Figure 6D The left-biased 3×3 array represents the subview (2, 2) at the top center position. This subview corresponds to the second (top-middle) row and the second (left-middle) column. Figure 5 The black dashed ellipse.
[0180] In some embodiments, the representation may include a full-field array. In some embodiments, the representation may include a set of views of the light field array. In some embodiments, the representation may include subsamples of views of the light field array. In some embodiments, the representation may include a subset of views selected from a full-field array or a subset of views of the light field array. In some embodiments, the representation (which may be designated as a subsampled representation) may include subsamples of another representation of the light field array. In some embodiments, subsamples of views may include views corresponding to a specific direction and other directions, and may include one or more uniform subsamples, such as... Figure 6B , 6C And the example shown in 6D. In some embodiments, the set of views of the light field content may contain one or more views with different sampling rates, such that the set of views can correspond to multiple sampling rates.
[0181] Figure 7 This is a schematic diagram illustrating an example numbered configuration of a light field subview according to some embodiments. Figure 7 The example subview numbering configuration shown is used in the example MPD shown in Table 1 below. Figure 7 An example subview number configuration for a full-field 5×5 array 700 is shown. It can be... Figure 7 Subview numbering configuration and Figure 6A The subview numbering configurations are compared. Figure 6A An exemplary configuration of coordinate subview numbers is shown, while Figure 7 An exemplary sequential subview numbering configuration is shown. Figure 7 In a 5x5 array of 700, the top row numbers subviews 1 to 5 from left to right. The second row numbers subviews 6 to 10 from left to right, continuing until the last row, which is numbered from left to right. Figures 21 to 25 serial number. Figure 7-9 Along with Table 1, an example MPD of the example client pull model is shown.
[0182] Figure 8 This is a diagram illustrating an example MPEG-DASH Media Presentation Description (MPD) file according to some embodiments. An example client pull model (e.g., using MPEG-DASH) can be used... Figure 8 The standard structure of MPEG-DASH Media Presentation Description (MPD) 802 is shown. Figure 8 The MPD file format 800 shown can be used as part of the initialization of a streaming session to transmit the overall media description downloaded by the viewing client. The organizational structure of the MPEG-DASH MPD is as follows: Figure 8 As shown.
[0183] Top-level time period fields 804 and 806 can indicate the start time and duration. MPD 802 may include one or more time period fields 804 and 806. Time period fields 804 and 806 may include one or more adapter sets 808 and 810. Adapter sets 808 and 810 may include one or more representation fields 812 and 814. Each representation 812 and 814 within an adapter set 808 and 810 may contain the same content encoded with different parameters. Representation fields 812 and 814 may include one or more segments 816 and 818. Segments 816 and 818 may include one or more sub-segments 824 and 826, which include DASH media files. Representation fields 812 and 814 may be divided into one or more sub-representation fields 820 and 822. Sub-representation fields 820 and 822 may include information applicable only to one media stream.
[0184] In some embodiments, one or more representation fields of the MPD may include a higher view density for a first viewing direction compared to a second viewing direction. In some embodiments, a media manifest file (e.g., an MPD) may include priority data for one or more views, and the process of interpolating views may use this priority data (such as to select which view can be interpolated). In some embodiments, a media manifest file (e.g., an MPD) may include priority data for one or more views, and the process of selecting a representation may use this priority data (such as selecting a representation with a higher priority). In some embodiments, a media manifest file (e.g., an MPD) may include at least one representation having information corresponding to two or more views of light field content. In some embodiments, the information in the manifest file may include position data for two or more views. In some embodiments, the information in the manifest file may include interpolation priority data for one or more of the plurality of views, and wherein the selection of one of the plurality of subsampled representations is based on the interpolation priority data for one or more of the plurality of views. In some embodiments, a process may include: tracking the user's viewing direction such that the selection of a representation may select a representation having a view density associated with the user's viewing direction. In some embodiments, the representation may include two or more views having a density related to the user's viewing direction, such as the density of views related to the left viewing direction or the density of views related to the right viewing direction.
[0185] Figure 9 This is a diagram illustrating an example MPD file with light field configuration elements according to some embodiments. Figure 9 This demonstrates how MPD data 902 is organized according to the MPEG-DASH protocol structure, which enables a sample client pull model for multiplexed optical field streaming.
[0186] In some embodiments, the MPD structure 900 uses time periods 904 and 906 as top-level entities. Each time period 904 and 906 can provide information about a single light field scene. A single scene can be, for example, continuous light field rendering, where the virtual camera array used for rendering remains constant. The entire experience can include several scenes, each specified in a separate time period block. Each time period block can include light field rendering settings, which are... Figure 9The light field description block 908 is labeled as light field description 908 and is associated with the first adapter set. Light field rendering settings block 908 may include the number of subviews, the placement of the virtual camera used by the subviews, and an overview of the scene layout, such as the size of the views and the size and placement of elements within the views. Each time-segment block 904, 906 may include one or more subset blocks 910, 912, 914, each subset block being associated with an adapter set. Each subset adapter set 910, 912, 914 may contain one or more different versions of content encoded with different configuration parameters. For example, the content may be encoded with different resolutions, different compression rates, different bitrates, and / or different supported codecs. Each subset 910, 912, 914 may include a set number of subviews depending on the adapter set, with a resolution depending on the representation block, encoded using a specific codec at a specific bitrate for the segment block. Each subset 910, 912, 914 may be divided into short clips, which are contained within sub-segment blocks having links to the actual video data.
[0187] The adapter set within time periods 904 and 906 of MPD 902 may include subsets 910, 912, and 914 of the full array of views, varying the number of views and the sampling positions of the views in the adapter set. The adapter set may contain the number of existing views and the indices of available views. The adapter set may indicate the priority of subviews. The adapter set (or, for some embodiments, subsets 910, 912, and 914) may include one or more resolutions 918 and 920, each resolution including one or more bitrates 922, 924, and 926. For each resolution 918 and 920, a series of time steps 1, 2, ..., N (928, 930, 932) may exist. Each time step 928, 930, and 932 may have a separate URL 934, 936, 938, 940, 942, and 944 for each supported bitrate. Figure 9 The example shown has N resolutions, each supporting N bitrates. Additionally, the adapter sets within time slots 904 and 906 of MPD 902 can be used for audio, such as... Figure 9 The example shows audio block 916.
[0188] In some embodiments, a manifest file (such as a media manifest file or MPD) may be used. Information in the manifest file may include location data for two or more views. In some embodiments, information in the manifest file may include interpolation priority data for one or more views. In some embodiments, selection subsampling representation may be based on interpolation priority data for one or more views. For example, interpolation priority data may be used to determine how often a subview is interpolated or updated.
[0189] Table 1 shows the table with Figure 9The fields shown and Figure 7 The example MPD with a subview configuration of the 5×5 array light field is shown in pseudocode. Table 1 illustrates three different light field view adapter sets: full, center, and left. The number of sampled views and their positions vary between adapter sets. For a given resolution, fewer views result in more pixels per view. Traditional DASH rate and resolution adapters can be used within each view category. An example for UHD / HD is shown in the "center" light field view adapter set.
[0190] ●Time Period
[0191] ○ Adaptation Set 1: Light Field Scene Description:
[0192] ■Sparse view array. Number of subviews: 5x5. Virtual camera position.
[0193] ○ Adaptation set 2: Subset 1 (Full Light Field): Number of subviews 5x5 (subviews 1-25)
[0194] ■Resolution 1: 1920x1080px per subview
[0195] ●Bitrate 1: Using codec XX, a transmission capacity of 35Mbps is required.
[0196] ●Bitrate 2: Using codec XX, a transmission capacity of 31Mbps is required.
[0197] ■Resolution 2: 1280x720px per subview
[0198] ●Bitrate 1: Using codec XX, a transmission capacity of 22Mbps is required.
[0199] ● Bitrate 2: Using codec XX, a transmission capacity of 19Mbps is required.
[0200] ○ Fit set 3: Subset 2 (sparse sampling): Number of subviews 2x2 (subviews 1, 5, 21, 25)
[0201] ■Resolution 1: 1920x1080px per subview
[0202] ●Bitrate 1: Using codec XX, a transmission capacity of 9Mbps is required.
[0203] ● Bitrate 2: Using codec XX, a transmission capacity of 8Mbps is required.
[0204] ■Resolution 2: 1280x720px per subview
[0205] ●Bitrate 1: Using codec XX, a transmission capacity of 6Mbps is required.
[0206] ● Bitrate 2: Using codec XX encoding, a transmission capacity of 5Mbps is required.
[0207] ○ Adaptation set 4: Subset 3 (semi-dense partial sampling): Number of subviews 2x2 (subviews) Figure 2 ,6,8,12)
[0208] ■Resolution 1: 1920x1080px per subview
[0209] ●Bitrate 1: Using codec XX, a transmission capacity of 9Mbps is required.
[0210] ● Bitrate 2: Using codec XX, a transmission capacity of 8Mbps is required.
[0211] ■Resolution 2: 1280x720px per subview
[0212] ●Bitrate 1: Using codec XX, a transmission capacity of 6Mbps is required.
[0213] ● Bitrate 2: Using codec XX encoding, a transmission capacity of 5Mbps is required.
[0214] ○ Adaptation set 5: Subset 4 (semi-dense partial sampling): Number of subviews 2x2 (subviews) Figure 4 ,8,10,14)
[0215] ■Resolution 1: 1920x1080px per subview
[0216] ●Bitrate 1: Using codec XX, a transmission capacity of 9Mbps is required.
[0217] ● Bitrate 2: Using codec XX, a transmission capacity of 8Mbps is required.
[0218] ■Resolution 2: 1280x720px per subview
[0219] ●Bitrate 1: Using codec XX, a transmission capacity of 6Mbps is required.
[0220] ● Bitrate 2: Using codec XX encoding, a transmission capacity of 5Mbps is required.
[0221] ○ Adaptation set 6: Subset 5 (semi-dense partial sampling): Number of subviews 2x2 (subviews) Figure 12 ,16,18,22)
[0222] ■Resolution 1: 1920x1080px per subview
[0223] ●Bitrate 1: Using codec XX, a transmission capacity of 9Mbps is required.
[0224] ● Bitrate 2: Using codec XX, a transmission capacity of 8Mbps is required.
[0225] ■Resolution 2: 1280x720px per subview
[0226] ●Bitrate 1: Using codec XX, a transmission capacity of 6Mbps is required.
[0227] ● Bitrate 2: Using codec XX encoding, a transmission capacity of 5Mbps is required.
[0228] ○ Adaptation set 7: Subset 6 (semi-dense partial sampling): Number of subviews 2x2 (subviews) Figure 14 ,18,20,24)
[0229] ■Resolution 1: 1920x1080px per subview
[0230] ●Bitrate 1: Using codec XX, a transmission capacity of 9Mbps is required.
[0231] ● Bitrate 2: Using codec XX, a transmission capacity of 8Mbps is required.
[0232] ■Resolution 2: 1280x720px per subview
[0233] ●Bitrate 1: Using codec XX, a transmission capacity of 6Mbps is required.
[0234] ● Bitrate 2: Using codec XX encoding, a transmission capacity of 5Mbps is required.
[0235] ○ Fit set 8: Subset 7 (single subview): Number of subviews 1x1 (subview 1)
[0236] ■Resolution 1: 1920x1080px per subview
[0237] ●Bitrate 1: Using codec XX, a transmission capacity of 3Mbps is required.
[0238] ● Bitrate 2: Using codec XX, a transmission capacity of 2.5 Mbps is required.
[0239] ■Resolution 2: 1280x720px per subview
[0240] ●Bitrate 1: Using codec XX, a transmission capacity of 2Mbps is required.
[0241] ● Bitrate 2: Using codec XX, a transmission capacity of 1.8 Mbps is required.
[0242]
[0243] ○ Fit set 32: Subset 31 (single subview): Number of subviews 1 x 1 (subview Figure 25 )
[0244] ■Resolution 1: 1920x1080px per subview
[0245] ●Bitrate 1: Using codec XX, a transmission capacity of 3Mbps is required.
[0246] ● Bitrate 2: Using codec XX, a transmission capacity of 2.5 Mbps is required.
[0247] ■Resolution 2: 1280x720px per subview
[0248] ●Bitrate 1: Using codec XX, a transmission capacity of 2Mbps is required.
[0249] ● Bitrate 2: Using codec XX, a transmission capacity of 1.8 Mbps is required.
[0250] Table 1
[0251] Figure 10 This is a system diagram illustrating a set of example interfaces for multiplexed light field rendering according to some embodiments. In some embodiments, the example system 1000 having a server (e.g., content server 1002) may include a database 1010 of spatial content. The content server 1002 may be able to run an example process 1012 for analyzing content and performing multiplexed light field rendering. One or more content servers 1002 may be connected to a client (e.g., viewing client 1004) via a network 1018. In some embodiments, the viewing client 1004 may include a local cache 1016. The viewing client 1004 may be allowed to perform an example process 1014 for compositing views (or subviews in some embodiments) onto content. The viewing client 1004 may be connected to a display 1006 for displaying content (including a composited view of the content). In some embodiments, the viewing client 1004 may include an input 1008 for receiving tracking data (such as user gaze tracking data, or, for example, user location data, such as user head position data from, for example, headphones or head-mounted devices (HMDs)) and input data. In some embodiments, the viewing client 1004 (or, in some embodiments, a sensor) can be used to track the user's gaze direction. The user's gaze direction can be used to select a subsampled representation of the content. For example, the user's gaze direction can be used to determine the user's viewpoint (and associated subviews), such as... Figure 4 and Figure 5Example subviews are indicated by white ellipses. In some embodiments, the distance between the user's gaze point and the center of each subview in the full-field array can be minimized to determine the subview closest to the user's gaze. The subview associated with this minimum distance can be selected as the viewpoint subview.
[0252] In some embodiments, the selection of a representation (such as a subsampled view of a light field array) can be based on the user's position. In some embodiments, the viewing client can track the user's head position, and the selection of a representation can be based on the tracked user head position. In some embodiments, the user's gaze direction can be tracked, and the representation can be selected based on the tracked user's gaze direction.
[0253] Figure 11 This is a message sequence diagram illustrating an example process for multiplexed light field rendering using an example client to pull a model, according to some embodiments. For some embodiments of the exemplary light field (LF) generation process 1100, server 1102 may generate 1106 subsampled light fields and MPD descriptions, such as regarding... Figure 2-9Some examples are described below. In some embodiments, client 1104 (or viewing client) may track 1108 the viewer's gaze and use the viewer's gaze to select content. A content request may be sent 1110 from client 1104 to server 1102 (or content server), and the MPD may be sent 1112 in response to client 1104. In some embodiments, server 1102 may send the MPD 1112 to the client without a request. In some embodiments, client 1104 may estimate 1114 the available bandwidth of the network between server 1102 and client 1104 and / or estimate the bandwidth to be used in one or more subsampling scenarios. Client 1104 may select 1116 light field subsamplings in part based on the estimated bandwidth(s). For example, 1116LF subsamplings may be selected from the sample set that minimizes the bandwidth used while keeping one or more light field parameters above a threshold (e.g., using a resolution and bit rate above a specific threshold for the first two priority subviews). Client 1104 may send 1118 a light field representation request to server 1102. Server 1102 can retrieve the requested LF representation and send the requested representation to client 1120. Client 1104 can receive the requested representation as one or more content fragments. In some embodiments, client 1104 can use the received content fragments to interpolate one or more interpolated subviews 1124. Viewing client 1104 can display the received and interpolated subview light field content data 1126. In some embodiments, one or more composite views (e.g., interpolated composite views) can be composited from one or more interpolated subviews, and viewing client 1104 can display 1126 the one or more composite views (e.g., interpolated composite views).
[0254] For example, the client can receive Figure 4 The content fragments of the highest, second-highest, and third-highest priority subviews are shown, and the client can interpolate the content data of one or more additional subviews. In some embodiments, for example, the client can request the content fragment of the highest priority subview at the request rate shown in Equation 1:
[0255]
[0256] Here, variable x equals the time unit. The client can request content fragments of the second highest priority subview at the request rate shown in Equation 2:
[0257]
[0258] The client can request content fragments of the third high-priority subview at the request rate shown in Equation 3:
[0259]
[0260] Some embodiments may interpolate subview content data for subviews that have not received content data for a specific time step. Some embodiments may store subview content data and may use the stored subview content data for time steps occurring between subview request rates of a specific priority. The request rates shown in Equations 1-3 are examples, and in some embodiments, request rates may be assigned differently.
[0261] In some embodiments, the selection of a representation of the light field (which may be a subsample of the light field in some embodiments) may be based on bandwidth constraints. For example, a representation below the bandwidth constraint may be selected. In some embodiments, the user's head position may be tracked, and the representation may be selected based on the user's head position. In some embodiments, the user's gaze direction may be tracked, and the representation may be selected based on the user's gaze direction. In some embodiments, the client process may include interpolating at least one view such that the selection of a representation can be chosen from a group including the interpolated views. In some embodiments, the client process may include requesting light field video content from a server such that obtaining the selected subsampled representation includes performing a process of selecting from a group consisting of: retrieving the selected subsampled representation from the server, requesting the selected subsampled representation from the server, and receiving the selected subsampled representation. In some embodiments, the client process may include: requesting a media manifest file from a server; and requesting light field content associated with a selected subset of views, such that obtaining the light field content associated with the selected subset of views may include: performing a process selected from the group consisting of: retrieving the light field content associated with the selected subset of views from a server, requesting the light field content associated with the selected subset of views from a server, and receiving the light field content associated with the selected subset of views.
[0262] Figure 12 This is a message sequence diagram illustrating an example process for reusing light field rendering using an example server to push a model according to some embodiments. Client and server processing will be described in further detail in this application.
[0263] In some embodiments, the example server push process 1200 may include, for example, a preprocessing process and a runtime process 1222. In some embodiments of the example preprocessing process, the viewing client 1202 may send a content request 1206 to the content server 1204, and the content server 1204 may respond 1208 using a first full-field frame. The viewing client 1202 may update the viewpoint 1210 and composite the view, which can be done using the first full-field frame.
[0264] In some embodiments of exemplary runtime process 1222, viewing client 1202 may send an updated viewpoint 1212 to content server. Content server 1204 may use a motion model of the viewpoint motion to predict the viewpoint position for the next time step. In some embodiments, content server 1204 may construct the motion model 1214 in part using the user's tracked motion. Content server 1204 may analyze 1216 the content and may prioritize the light field subviews to be rendered for the next time step based on the predicted view position. Content server 1204 may send 1218 the selected light field subview, the virtual camera position associated with the selected subview, and the timestamp of the content to viewing client 1202. Viewing client 1202 may update 1220 the viewpoint and use the temporal variability of the light field data to synthesize one or more subviews. In some embodiments, a frame may include a frame-packed representation of two or more views corresponding to the selected representation. In some embodiments, the light field data may include 2D image frame data sent from the content server to the client device. In some embodiments, the light field content may include frames of light field content (which may be transmitted from a content server to a client device).
[0265] Figure 13 This is a message sequence diagram illustrating an example process for multiplexed light field rendering using an example client-pull model, according to some embodiments. For some embodiments of the example client-pull model process 1300, client 1302 may perform future viewpoint prediction and perform content analysis to determine which subviews are prioritized. For some embodiments, client 1302 may pull a specific subview stream from content server 1304 based on the priority assigned to the subviews according to the analysis performed by viewing client 1302. In this model, client 1302 may select how many subviews to pull and which light field sampling to use by selecting which spatial distribution to use across the entire light field array. Viewing client 1302 may dynamically change the number and sampling distribution of the individual pulled subviews based on viewpoint motion, content complexity, and display capabilities. For some embodiments, a subset of views may be collected at the server and provided as alternative light field subsamplings. The client may select one of these subsets based on local criteria such as available memory, display rendering capabilities, and / or processing power. Distributing a subset of views allows the encoder to leverage redundancy between views for greater compression compared to compressing individual views. Figure 13 An overview of the sequence of processes and communications between viewing client 1302 and content server 1304 is shown in example client pull session 1300.
[0266] In some embodiments, content server 1304 may render a subset of the light field in a sequential order (1306). Viewing client 1302 may request content (1308), and content server 1304 may send back an MPD file (1310). Viewing client 1302 may select a subset to pull (1312). Viewing client 1302 may request a subset (1314), and content server 1304 may respond with subset content data, virtual camera position, and timestamp of capturing the content data (1316). Viewing client 1302 may store the received data (1318) in a local cache. Viewing client 1302 may update the user's viewpoint (1320) and synthesize the view using the temporal variability of the light field data stored in the local cache. Viewing client 1302 may use (or, in some embodiments, construct 1322) a motion model of the viewpoint motion and predict the viewpoint position in the next time step. Viewer 1302 can analyze the content of 1324 and prioritize the subset of light field to be pulled based on the predicted viewpoint position.
[0267] For some embodiments of the example client-side pull model, the client receives a Media Rendering Description (MPD) from a content server. The MPD can indicate available subsets and may include alternative light field subsamples compiled by the content server. Subsets can include individual subviews or collections of subviews, and the sampling density, distribution, and number of the included subviews can vary for each subset. By compiling several subviews into subsets, the content server can provide good candidate subsamples based on the content. If the subset is delivered as a single stream, the encoder can leverage the redundancy between views for greater compression. Each subset can be available in one or more spatial resolutions and compression versions.
[0268] In some embodiments of the client-side pull model, the content server may render and / or serve each individual light field subview at a fixed frame rate. In some embodiments, the content server may analyze the content to determine the subview update frequency. In some embodiments, selecting a representation may include: predicting the user's viewpoint; and selecting the chosen representation may be based on the user's predicted viewpoint.
[0269] Figure 14This is a flowchart illustrating an example process for a content server according to some embodiments. For some embodiments, example process 1400 performed by the content server may include: receiving 1402 a content request from a client (e.g., "waiting for a content request from a client"). The viewing client may use tracking (e.g., user and / or eye position tracking) to determine the user's viewpoint and send that viewpoint to the content server while displaying content. Process 1400 may also include rendering 1404 a full-length image of a first frame sequence of the complete light field. Spatial data 1406 may be retrieved from memory (e.g., a cache local to the server). The rendered light field data may be sent 1408 to the client. 1410 The client's current viewpoint may be received and stored in memory as viewpoint data 1412. 1414 A motion estimation model may be constructed and / or updated using the viewpoint position data received from the client. The motion estimation model may be used to estimate the predicted viewpoint position for the next time step (which may be equal to the estimated time when the rendered content arrives at the viewing client). In some embodiments, the predicted viewpoint position may be based on a model that estimates network communication latency between the server and client, as well as the continuity of viewpoint motion. To estimate motion, Kalman filtering or any other suitable motion estimation model may be used. The predicted viewpoint for the next time step can be used to analyze the content at 1416 to determine the rendering priority of the light field subview. Using the predicted viewpoint, the content server can perform content analysis to determine the highest priority subview for rendering and delivery to the viewing client in the next time step. The subview at 1418 can be rendered based on the rendering priority. The rendered light field subview content, timestamp, and virtual camera position can be streamed to the client at 1420. A determination at 1422 whether to request session termination can be performed. If no session termination is requested, process 1400 can repeat, starting from receiving the client's current viewpoint at 1410. A determination at 1424 whether to request processing termination can be performed. If a processing termination signal is received, process 1400 exits at 1426. Otherwise, process 1400 can repeat, starting from waiting for a content request from the client at 1402. In some embodiments, the content server continuously streams the multiplexed light field data throughout the duration of the content or until the end of the viewing client's signal session, prioritizing the rendering of subviews.
[0270] Figure 15This is a flowchart illustrating an example process for a viewing client according to some embodiments. If a user launches the viewing client application, the user also indicates the content to be viewed. The content can be a link to that content, for example, the content is located on and / or streamed from a content server. The link to the content can be a URL identifying the content server and the specific content. The viewing client application can be launched by an explicit command from the user, or automatically by the operating system based on identifying the content type request and the application associated with the specific content type. In addition to being a standalone application, the viewing client can be integrated with a web browser or social media client, or the viewing client can be part of the operating system. If the viewing client application is launched, the viewing client can initialize sensors for device, user, and / or gaze tracking.
[0271] Some embodiments of the example process 1500 performed by the viewing client may include: requesting content from the content server 1502, and initializing user gaze tracking 1504. The viewing client may receive a first full-scale light field image frame 1506 from the content server. The light field data may be stored in a local cache 1508 and displayed by the viewing client. An initial user viewpoint may be set using a default position, which can be used to display the first light field frame. The viewpoint may be updated 1510 based on device tracking and user input. The current viewpoint may be sent 1512 to the content server. The content server may use the received viewpoint to render one or more subviews of the light field, which may be streamed to the client. In step 1518, the client may receive these subviews, a timestamp indicating a shared time step of the rendering, the position of the virtual camera relative to the full image array, and optical parameters for rendering the subviews.
[0272] One or more views 1514 of the light field can be composited based on the current viewpoint. This composited view can be a new view, for example, one that lacks light field content specifically associated with that particular view. The composited view(s) can be displayed by a viewing client 1516. At 1518, the viewing client can receive light field subview content rendered by a content server, which can be stored in memory such as a local light field cache 1508. The viewing client can use tracking and user input to composite the viewpoint for one or more rendering steps using the latest viewpoint update 1510. A determination 1520 can be performed to determine whether a processing end signal has been received. If a processing end signal is received, process 1500 can end 1522. Otherwise, process 1500 can be repeated by updating the user viewpoint 1510.
[0273] In some embodiments, the client process may include: generating a signal for display using the rendered representation. In some embodiments, the client process may include: determining a user's viewpoint using the user's gaze direction, such that selecting a selected representation may include: selecting the selected representation based on the user's viewpoint. In some embodiments, the client process may include: determining a user's viewpoint using the user's gaze direction; and selecting at least one subsample of a view of a multi-view video, such that selecting a subsample of the view may include: selecting a subsample of the view within a threshold viewpoint angle of the user's viewpoint.
[0274] In some embodiments, the viewing client renders the image to be displayed by synthesizing a viewpoint that matches the current viewpoint using locally cached light field data. If a viewpoint is synthesized, the viewing client can select a subview to use in the synthesis. In some embodiments, for selection, the viewing client can examine the local light field cache to identify a subview that, for example, is close to the current viewpoint, has the latest available data, has sufficient data from various time steps to enable interpolation or prediction to produce a good estimate of the subview's appearance at a given time step, etc. To mitigate temporal variations between subviews (which may be received in separate subset streams), the viewing client can use techniques developed for video frame interpolation, such as those in Nikilaus, S et al., Video Frame Interpolation via Adaptive Separable Convolution, P ROC.OF THE IEEE I NT ' L C ONF.ON C OMP .V ISION The techniques mentioned in 261-270 (2017) (“Nikilaus”) are used to generate frames corresponding to the rendering time when “future” frames are available from the content server, or when only frames from previous time steps are available for a particular subview. Vedran et al., One-Step Time-Dependent Future Video Frame Prediction with a Convolutional Encoder-Decoder Neural Network, P ROC.OF I NT ' L C ONF.ON I MAGE A NALYSISAND P ROCESSING Video frame prediction as described in 140-151 (2017). In some embodiments, the viewing client may use one of these two methods to estimate a subview at a specific time step, which can be used for rendering and displaying the view. In some embodiments, the viewing client uses subview images at some time steps stored in a local cache. In some embodiments, if the viewing client has all the subviews for synthesizing a viewpoint at a given time step, the viewing client may use the subviews described in Kalantari, Nima Khademi et al., Learning-Based View Synthesis for Light Field Cameras, 37.6 ACM T. RANSACTIONS ON G RAPHICS The process described in (TOG)193(2016)(“Kalantari”) is used to synthesize, for example, a new viewpoint from a sparse global light field formed by selected subviews.
[0275] In some embodiments, an example viewing client process may include: determining the user's viewpoint from the user's gaze direction. The process may also include selecting one or more subviews of the light field video content from a group comprising one or more interpolated views and retrieved subsampled representations. The process may further include: synthesizing a view of the light field using one or more selected subviews and the user's viewpoint. In some embodiments, the process may include: determining the user's viewpoint from the user's gaze direction; and selecting one of the plurality of subsampled representations based on the user's viewpoint. In some embodiments, the example viewing client process may also include: displaying the synthesized view of the light field. In some embodiments, selecting one or more subviews of the light field may include: selecting one or more subviews within a threshold viewpoint angle of the user's viewpoint.
[0276] In some embodiments, the client process may include: determining the user's viewpoint from the user's gaze direction, such that selecting one of the plurality of sub-sampled representations may be further based on the user's viewpoint. In some embodiments, the client process may include: displaying a composite view of the light field. In some embodiments, selecting one or more sub-views of the light field may include: selecting one or more sub-views within a threshold viewpoint angle of the user's viewpoint.
[0277] Adapted for optical field streaming
[0278] Using many existing multi-view coding standards to support dynamic user viewpoints and motion parallax can lead to excessive bandwidth usage, especially when only a single view or stereoscopic view is generated for display in a single time step. Figure 2Applications of 2D images. Examples of such applications include 3DoF+ or 6DoF applications based on viewing content on an HMD or other 2D display.
[0279] In some embodiments, to optimize rendering and data distribution, a subset of the complete overall light field data can be generated and transmitted, and the receiver can synthesize additional views from the transmitted views. Furthermore, the impact of the selection of subviews for rendering and transmission on perceived image quality can guide the subview rendering process on the server side or the client side to ensure quality of experience.
[0280] In some embodiments, optical field streaming can be adapted to available resources by varying the number of optical field sub-viewpoints during streaming, thereby optimizing data transmission.
[0281] The content server can provide multiple versions of light field data, characterized by a variable number of light field subviews. In addition to several versions of the streaming content, the server can provide metadata describing the available streams as a streaming adaptation manifest. At the start of a content streaming session, the client can download the manifest metadata and, based on this metadata, begin downloading content fragments. While downloading content fragments, the client can observe session characteristics, network and processing performance, and can adapt the streaming quality by switching between available content streams.
[0282] In some implementations, the client can optimize the quality of experience while dynamically adapting the light field streaming to changing performance and network conditions. By allowing the viewing client to dynamically adjust the number of light field subviews transmitted from the content server to the viewing client, content delivery bandwidth requirements can be reduced.
[0283] Some embodiments may use an exemplary client-pull model for streaming, which operates similarly to the MPEG-Dash model. The server can generate several versions of the light field content by varying the number of subviews included in each individual version of the stream. The server may provide the client with a manifest file (e.g., MPD) indicating the number and location of subviews included with each available version of the content stream. The client can continuously analyze session characteristics and performance metrics to determine which stream to download, thereby utilizing given network and computing resources to deliver Quality of Experience (QoE).
[0284] Utilizing existing multi-view coding standards to support dynamic viewpoints and motion parallax can lead to excessive bandwidth usage, especially in applications where only a small number of viewpoints are needed at any given time. Example methods and systems according to some embodiments can avoid these limitations by taking into account the impact of the current and predicted user viewpoints, as well as the selected number of subviews packed into frames for delivery, on the perceived image quality.
[0285] Figure 16 This is a schematic diagram illustrating an example light field rendered with 5×5 subviews according to some embodiments. In some embodiments, subsampling of the views can be used to reduce transmission bandwidth, where view interpolation is used at the receiver to generate the desired view. In some embodiments, varying the number of packed views can be used to adapt to bandwidth or processing constraints at the client. Figure 16 An exemplary full light field 1600 rendered using a 5×5 virtual camera array with subview 1602 is shown.
[0286] Figure 17 This is a schematic diagram illustrating an example light field rendered with 3×3 subviews according to some embodiments. The priority of each light field subview can be determined based on content analysis and contextual information (e.g., user / gaze position and display capabilities). The individual subview sets can be generated based on sequential priority order. A client can receive a subview, store it in a cache, and use temporal coherent interpolation of the subviews to synthesize a new viewpoint. Figure 17 This example shows nine views encapsulated within a 3×3 grid of 1700. Figure 17 The view in the image is based on the priority settings described above for some embodiments. Figure 16 The 5×5 array is selected in a checkerboard pattern within the full light field. Figure 19B An example is shown below, in which a specific subview is selected for the example 3×3 array.
[0287] Figure 18 This is a schematic diagram illustrating an example light field rendered with a 2×3 subview according to some embodiments. Figure 18 An example showing six views packed in a 2×3 grid of 1800. Figure 18 The view in the image is based on the priority settings described above for some embodiments. Figure 16 The full light field shown in the 5×5 array is selected in a rectangular pattern. Figure 19C An example is shown below, in which a specific subview is selected for the example 2×3 array.
[0288] Figure 19A It is shown that, according to some embodiments, it corresponds to Figure 16 A diagram illustrating an example 5×5 array light field configuration. For some embodiments, the Media Presentation Description (MPD) may include details of the full (or entire) light field array, such as the number of views, the indexes of those views, and the sampling positions of the views within the adaptation set. For example, a full light field array 1900 may have N×M view positions. Figure 19AAn example naming convention for a 5x5 array 1900 with a full light field is shown. The first number of each coordinate position indicates the row, and the second number indicates the column. Position (1, 1) is at the top left corner, and position (5, 5) is at the bottom right corner. Other embodiments may use different naming conventions, such as... Figure 21 The example shown.
[0289] Figure 19B It is shown that, according to some embodiments, it corresponds to Figure 17 A diagram of an example 3×3 array light field configuration. Figure 19B The 3×3 array 1920 shown corresponds to Figure 20 A solid black ellipse in, for example, Figure 19B The uniform subsampled 3×3 array 1920 represents the subview (5, 3) at the lower center position. This subview corresponds to Figure 20 The black ellipse in the fifth (bottom) row and the third (center) column.
[0290] Figure 19C It is shown that, according to some embodiments, it corresponds to Figure 18 A diagram of an example 2×3 array light field configuration. Figure 19C The 2×3 array 1940 shown corresponds to from Figure 16 A rectangular view pattern selected from the middle portion of the full light field of a 5×5 array. For example, Figure 19C The 2×3 array 1940 indicates the subview (2, 4) at the upper right position. This subview corresponds to... Figure 16 The second (middle-top) row and the fourth (middle-right) column. Figure 19C Example 2×3 array 1940 can be as follows Figure 20 It is shown in black ellipse (not shown).
[0291] Figure 20 This is a schematic diagram illustrating example subsampling using 3×3 subviews according to some embodiments. The content server can generate several subsampled versions of the original light field data by using several subsampling configurations. For example, the content server can reduce the number of subviews of the full light field 2000 to subsampling of 3×3 subviews 2002, 2004, 2206, 2208, 2010, 2012, 2014, 2016, and 2018, as shown below. Figure 17 and 19B As shown. Similarly, the content server can generate, for example, 10×10, 5×5, 2×2, and 1×1 array subsamples.
[0292] The content server can generate metadata that describes the available streams in the MPD file. For example, if the MPD describes two subsampling configurations (a 10x10 array and a 5x5 array), the 10x10 array can use the first stream, while the 5x5 array can use the second stream. In some embodiments, support for a 10x10 subview array may have, for example, only one stream (not 100 separate streams). Similarly, in some embodiments, a 5x5 array may have only five streams (or another number of streams, less than 25).
[0293] Figure 21 This is a schematic diagram illustrating an example numbered configuration of a light field subview according to some embodiments. Figure 21 The example subview number configuration 2100 shown is used in the example MPD shown in Table 2. Figure 21 An example subview number configuration is shown for a full-field 5×5 array. Figure 21 The subview number configuration can be combined with Figure 19A The subview numbering configurations are compared. Figure 19A An exemplary configuration of coordinate subview numbers is shown, while Figure 21 An exemplary sequential subview numbering configuration is shown. Figure 21 The top row of the 5x5 array is numbered from left to right as subviews 1 to 5. The second row is numbered from left to right as subviews 6 to 10, and so on until the last row, which is numbered from left to right as subviews... Figures 21 to 25 serial number. Figure 21-23 Together with Table 2, an example MPD of the example client pull model is shown.
[0294] Figure 22 This is a diagram illustrating an example MPEG-DASH Media Presentation Description (MPD) file according to some embodiments. An exemplary client-side pull model can be used... Figure 22 The standard structure of the MPEG-DASH Media Presentation Description (MPD) is shown. Figure 22 The MPD file format shown can be used to transmit the overall media description downloaded by the viewing client as part of the streaming session initialization. The organization structure of MPEG-DASH MPD is as follows: Figure 22 As shown.
[0295] Top-level time period fields 2204 and 2206 can indicate the start time and duration. MPD 2202 may include one or more time period fields 2204 and 2206. Time period fields 2204 and 2206 may contain one or more adapter sets 2208 and 2210. Adapter sets 2208 and 2210 may contain one or more representation fields 2212 and 2214. Each representation 2212 and 2214 within an adapter set 2208 and 2210 may include the same content encoded using different parameters. Representation fields 2212 and 2214 may include one or more segments 2216 and 2218. Segments 2216 and 2218 may include one or more sub-segments 2224 and 2226, which include DASH media files. Representation fields 2212 and 2214 may be divided into one or more sub-representation fields 2220 and 2222. Sub-representation fields 2220 and 2222 may include information applicable only to one media stream.
[0296] Figure 23 This is a diagram illustrating an example MPD file with light field configuration elements according to some embodiments. Figure 23 This demonstrates how to organize MPD data of an example client model according to the MPEG-DASH protocol structure to enable optical field streaming.
[0297] In some embodiments, the MPD structure 2300 uses time periods 2304 and 2306 as top-level entities. Each time period 2304 and 2306 can provide information about a single light field scene. A single scene can be, for example, continuous light field rendering, in which the virtual camera array used for rendering remains constant. The entire experience may include several scenes, each specified in a separate time period block. Each time period block may include light field rendering settings, which are... Figure 23The light field description block 2308 is labeled as light field description 2308 and is associated with the first adapter set. Light field rendering settings block 2308 may include the number of subviews, the placement of the virtual camera used by the subview, and an overview of the scene layout, such as the size of the view and the size and placement of elements within the view. Each time-period block 2304, 2306 may include one or more subset blocks 2310, 2312, 2314, each subset block being associated with an adapter set. Each subset adapter set 2310, 2312, 2314 may include one or more different versions of content encoded using different configuration parameters. For example, the content may be encoded with different resolutions, different compression rates, different bitrates, and / or different supported codecs. Each subset 2310, 2312, 2314 may contain a selected number of subviews in different adapter sets, with resolutions depending on the representation block and encoded with a specific codec at a specific bitrate for the fragment block. Each subset 2310, 2312, 2314 can be divided into short clips, which are contained in sub-segment blocks with links to the actual video data.
[0298] The adapter set within time periods 2304 and 2306 of MPD 2302 may include subsets 2310, 2312, and 2314 of the full array of views, which vary the number of views and the sampling positions of the views in the adapter set. The adapter set may contain the number of existing views and the indexes of available views. The adapter set may indicate the priority of subviews. The adapter set (or, for some embodiments, subsets 2310, 2312, and 2314) may include one or more resolutions 2318 and 2320, each including one or more bit rates 2322, 2324, and 2326. For each resolution 2318 and 2320, a series of time steps 1, 2, ..., N (2328, 2330, 2332) may exist. Each time step 2328, 2330, 2332 can have a separate URL 2334, 2336, 2338, 2340, 2342, 2344 for each supported bit rate. For Figure 23 The example shown has N resolutions, each supporting N bitrates. Additionally, the adapter sets within time slots 2304 and 2306 of MPD 2302 can be used for audio, such as... Figure 23 The example shows audio block 2316.
[0299] In some embodiments, the MPD may include information corresponding to two or more views, which correspond to a selected subset of views. In some embodiments, for at least one of the subsets of views, the media manifest file (e.g., the MPD) may include information corresponding to two or more views of the light field content.
[0300] Table 2 shows the uses for having Figure 23 The fields shown and Figure 21 The example MPD with a subview configuration for a 5×5 array light field is shown in pseudocode. Table 2 illustrates three different sets of light field view adapters with varying numbers of views packed into each frame. Conventional DASH rate and resolution adapters can be used within each view category.
[0301] ●Time Period
[0302] ○ Adaptation Set 1: Light Field Scene Description:
[0303] ■Sparse view array. Number of subviews: 5 x 5. Virtual camera position.
[0304] ○ Adaptation Set 2: Subset 1 (Full Light Field): Number of Subviews 5 x 5 (Including Subviews 1-25)
[0305] ■Resolution 1: 1920x1080px per subview
[0306] ●Bitrate 1: Using codec XX, a transmission capacity of 35Mbps is required.
[0307] ●Bitrate 2: Using codec XX, a transmission capacity of 31Mbps is required.
[0308] ■Resolution 2: 1280x720px per subview
[0309] ●Bitrate 1: Using codec XX, a transmission capacity of 22Mbps is required.
[0310] ● Bitrate 2: Using codec XX, a transmission capacity of 19Mbps is required.
[0311] ○ Fit set 3: Subset 2: Number of subviews 3 x 3 (including subviews 1, 3, 5, 11, 13, 15, 21, 23, 25)
[0312] ■Resolution 1: 1920x1080px per subview
[0313] ●Bitrate 1: Using codec XX for encoding, a transmission capacity of 10Mbps is required.
[0314] ● Bitrate 2: Using codec XX, a transmission capacity of 9Mbps is required.
[0315] ■Resolution 2: 1280x720px per subview
[0316] ●Bitrate 1: Using codec XX, a transmission capacity of 7Mbps is required.
[0317] ● Bitrate 2: Using codec XX, a transmission capacity of 6Mbps is required.
[0318] ○ Fit set 4: Subset 3: Number of subviews 2 x 2 (including subviews 1, 5, 21, 25)
[0319] ■Resolution 1: 1920x1080px per subview
[0320] ●Bitrate 1: Using codec XX, a transmission capacity of 7Mbps is required.
[0321] ● Bitrate 2: Using codec XX, a transmission capacity of 6Mbps is required.
[0322] ■Resolution 2: 1280x720px per subview
[0323] ● Bit rate 1: Using codec XX, a transmission capacity of 4Mbps is required.
[0324] ● Bitrate 2: Using codec XX encoding, a transmission capacity of 3Mbps is required.
[0325] ○ Adaptation set 5: Subset 4: Number of subviews 1 x 1 (including subviews) Figure 13 )
[0326] ■Resolution 1: 1920x1080px per subview
[0327] ●Bitrate 1: Using codec XX, a transmission capacity of 2Mbps is required.
[0328] ● Bitrate 2: Using codec XX, a transmission capacity of 1.5 Mbps is required.
[0329] ■Resolution 2: 1280x720px per subview
[0330] ● Bit rate 1: Using codec XX, a transmission capacity of 1 Mbps is required.
[0331] ● Bitrate 2: Encoded using codec XX, requiring a transmission capacity of 0.7 Mbps.
[0332] Table 2
[0333] Figure 24This is a system diagram illustrating a set of example interfaces for adapting light field streaming according to some embodiments. In some embodiments, the example system 2400 having a server (e.g., content server 2402) may include a database 2410 of light field content and one or more MPDs. The content server 2402 is capable of running an example process 2412 for content preprocessing. One or more content servers 2402 may be connected to a client (e.g., viewing client 2404) via a network 2418. In some embodiments, the viewing client 2404 may include a local cache 2416. The viewing client 2404 may be enabled to execute an example process 2414 for compositing views (or, in some embodiments, subviews) into content. The viewing client 2404 may be connected to a display 2406 for displaying content (including a composite view of the content). In some embodiments, the viewing client 2404 may include an input 2408 for receiving tracking data (such as user eye-tracking data) and input data. For example, Figure 24 Example content server 2402 can be with, for example Figure 10 The example content servers are configured similarly or differently.
[0334] Figure 25 This is a message sequence diagram illustrating an example process for adapting optical field streaming using estimated bandwidth and view interpolation, according to some embodiments. For some embodiments of the example adapting LF streaming process 2500, server 2502 can generate 2506 sub-sampled optical field and MPD descriptions, such as those for... Figure 16-23Some related examples are described. In some embodiments, client 2504 (or viewing client) may track 2508, for example, a viewer's gaze, and use the viewer's gaze to select content. In some embodiments, tracking may be used, for example, the user's location (e.g., in combination with or instead of a gaze). A content request 2510 may be sent from client 2504 to server (or content server 2502), and an MPD may be sent to the client in response. In some implementations, server 2502 may send the MPD 2512 to the client without a request. In some embodiments, client 2504 may estimate 2514 the available bandwidth of the network between server 2502 and client 2504 and / or estimate the bandwidth to be used in one or more subsampling scenarios. Client 2504 may select 2516 the number of packaged views or array configuration in part based on the estimated bandwidth(s). For example, the number of packaged views or array configuration may be selected from a set of samples 2516 that minimizes the bandwidth used while keeping one or more light field parameters above a threshold (e.g., using a resolution and bit rate above a specific threshold for the two highest priority subviews). Client 2504 may send a light field representation request (2518) to server 2502. Server 2502 may retrieve the requested LF representation (2520) and transmit the requested representation (2522) to client 2504. Client 2504 may receive the requested representation as one or more content fragments. In some embodiments, client 2504 may use the received content fragments to interpolate (2524) one or more interpolated subviews. Viewing client 2504 may display the received and interpolated subview light field content data (2526).
[0335] In some embodiments, the client process may include: obtaining light field content associated with a selected representation; decoding a frame of the light field content; combining two or more views represented in the frame to generate a composite view result; and rendering the composite view result to a display.
[0336] Figure 26This is a message sequence diagram illustrating an example process for adaptive light field streaming using predicted view locations, according to some embodiments. In some embodiments, client 2602 may perform future viewpoint prediction and perform content analysis to determine which subviews are prioritized. In some embodiments, this prioritization may be based on, for example, delivery bandwidth and view interpolation requirements and resources. In some embodiments, client 2602 may pull a specific representation characterized by a selected number of subviews from a content server based on the priority assigned to subviews according to, for example, analysis performed by viewing client 2602. In this model, client 2602 can select which light field sampling to use by selecting how many subviews to pull. Viewing client 2602 may dynamically change the number of individual subviews pulled based on viewpoint motion, content complexity, and display capabilities. In some embodiments, content representations with different numbers of views may be collected at the server and provided as alternative light field subsampling. Client 2602 may select one of these subsets based on local criteria such as available memory, display rendering capabilities, and / or processing power. Figure 26 An overview of the sequence of processes and communications used between the viewing client 2602 and the content server 2604 is shown in example client pull session 2600.
[0337] In some embodiments, content server 2604 may render a subset of the light field in a sequential manner (2606). Viewing client 2602 may request content (2608), and content server 2604 may send back an MPD file (2610). Viewing client 2602 may select (2612) the number of views representing the content to be fetched. Viewing client 2602 may request a subset (2614), and content server 2604 may respond (2616) with subset content data, virtual camera position, and timestamps of the captured content data. Viewing client 2602 may store the received data (2618) in a local cache. Viewing client 2602 may update the user's viewpoint (2620) and synthesize a view using the temporal variability of the light field data stored in the local cache. Viewing client 2602 may use (or, in some embodiments, construct (2622)) a motion model of the viewpoint motion and predict the viewpoint position in the next time step. Viewer 2602 can analyze the content of 2624 and prioritize the subset of light field to be pulled based on the predicted viewpoint location.
[0338] In some embodiments, selecting one of a plurality of viewpoint subsets may include: parsing information from a media manifest file (e.g., MPD) for a plurality of viewpoint subsets for light field content.
[0339] Figure 27This is a flowchart illustrating an example process for content server preprocessing according to some embodiments. In some embodiments of the preprocessing stage, the server may generate multiple subsets of light field data and metadata describing those subsets. Multiple subsets of light field data can be configured by selecting the number of subviews in various subsamples to be used to reduce the number of subviews included in the streaming data.
[0340] In some embodiments, the content server may execute the example preprocessing method multiple times for different subsampling configurations. In some configurations of the example preprocessing method, raw light field data 2704 may be received or read from a memory such as a local content server cache. The content server may select 2702 a variant with a different number of subviews 2708 to be generated. The content server may generate 2706 streaming data for each subview configuration (or, in some embodiments, a subsampling configuration). The content server may generate 2710 MPD (or metadata within the MPD) for each streaming data version or variant. The MPD 2712 may be stored by the content server in a local cache.
[0341] Figure 28 This is a flowchart illustrating an example process for content server runtime processing according to some embodiments. For some embodiments, the example adapts to a content distribution runtime process 2800, such as… Figure 28 The example shown can be executed in ways such as Figure 27 The example preprocessing method is executed after preprocessing. In some embodiments, the content server may receive a content request 2802 from a viewing client. Runtime process 2800 may determine 2804 the request type. If the content request indicates a new session, the content server may retrieve 2806 the MPD corresponding to the content request from memory (e.g., a local cache) and send the MPD 2808 to the viewing client. In some embodiments, the MPD may be generated at runtime, which may be, for example, similar to the method described above. Figure 27 The described process. If the content request indicates a request for a subset of data fragments, the content server can retrieve the light field subset stream from storage 2812 and send the data fragment 2810 to the viewing client. The runtime process 2800 can determine 2814 whether an end-of-process request has been made. If an end-of-process signal is received, the process can terminate 2816. Otherwise, the process can be repeated by waiting 2802 for a request for content.
[0342] Figure 29This is a flowchart illustrating an example process for a viewing client according to some embodiments. In some embodiments of the example viewing client process 2900, a content request may be sent to a content server at 2902. In some embodiments, a content request may be sent if a user launches the viewing client application and indicates content to be viewed. In some embodiments, the viewing client application may be launched by executing an explicit command or script file from the user, or automatically by the operating system recognizing the content request and the application associated with the content type in the content request. In some embodiments, the viewing client may be, for example, a standalone application, an integrated web browser or social media client, or part of an operating system. In some embodiments, launching the viewing client application causes sensor initialization, such as sensors used by the viewing client device, the user, or gaze tracking.
[0343] The viewing client may receive, 2904, a manifest file (e.g., a media manifest file), such as an MPD (or, in some embodiments, an adapter manifest) corresponding to the requested content. In some embodiments, the viewing client initializes, 2906, device tracking. The viewing client may select, 2908, an initial number of subviews (or representations) to request (or download from) the content server. The viewing client may sequentially download sub-segments from the content server. In some embodiments, the viewing client may determine an initial set of adapters (or, in some embodiments, the number of packaged subsets) and a specific representation to request from the content server based on, for example, display settings, tracking application settings, resolution, and bitrate.
[0344] The subviews can be packaged into each frame of the content to be pulled by the viewing client. In some embodiments, the viewing client can request 2910 streams with a selected number of subviews. In some embodiments, the viewing client can download media segments of a selected subset from a content server. In some embodiments, the content can be a URL link identifying the content data and / or an MPD file. The viewing client receives 2912 these subview sets. In some embodiments, upon receiving a representation of the first segment, the client can begin a continuous runtime process. This runtime process can include updating the viewpoint 2914. The viewpoint can be updated based on device tracking and user input. The viewpoint can be synthesized using the received light field data and the current viewpoint 2916. Based on the received data, the client renders the light field in a display-specific format, interpolating additional subviews for some viewing conditions.
[0345] The number of selected subviews can be updated based on, for example, user tracking (e.g., user location, user gaze, user viewpoint, user view tracking), content analysis, performance metrics, viewpoint motion, content complexity, display capabilities, network and processing performance, or any other suitable criteria. In some embodiments, performance metrics (e.g., bandwidth used, available bandwidth, and client capabilities) can be measured or observed, and the number of selected subviews can be updated accordingly. As part of an exemplary runtime process, the viewing client can update the number of selected subviews. The runtime process 2900 may include determining whether 2920 has requested termination of processing. If a termination signal is received, process 2900 may terminate 2922. Otherwise, process 2900 may repeat by requesting 2910 (or, in some embodiments, downloading) the selected subview stream.
[0346] In some embodiments, the frame may include a frame-packed representation of two or more views corresponding to a selected subset of views. In some embodiments, combining two or more views represented in a frame may involve using view composition techniques. In some embodiments, selecting one subset of views from the plurality of view subsets is based on at least one of the following criteria: gaze tracking, content complexity, display capability, and bandwidth available for retrieving light field content. In some embodiments, selecting one of the plurality of view subsets may include: predicting the user's viewpoint; and selecting the view subset based on the predicted user viewpoint. For example, the viewpoint subset may be selected as being within a user's viewpoint angle threshold.
[0347] In some embodiments, the selection of a representation may be based on at least one of the following criteria: user head position, gaze tracking, complexity of the content, display capability, and bandwidth available for retrieving the light field content. In some embodiments, the selection of a subset of views may be based on at least one of the following criteria: user head position, gaze tracking, complexity of the content, display capability, and bandwidth available for retrieving the light field content.
[0348] In some embodiments, a frame may include a frame-packed representation of two or more views corresponding to a selected subset of views. In some embodiments, combining two or more views represented in a frame may include using view composition techniques. For some embodiments, selecting a subset of views may include: predicting the user's viewpoint; and selecting the subset of views based on the predicted user viewpoint. In some embodiments, generating a view from the obtained light field content may include: interpolating the views from the light field content associated with the selected subset of views to generate the generated view, the interpolation being performed using information corresponding to the views respectively in the manifest file.
[0349] Figure 30 This is a flowchart illustrating an example process for a viewing client according to some embodiments. Some embodiments of the example method 3000 for the viewing client may include: receiving 3002 a media manifest file that includes information on a plurality of subsampled representations of a view of light field video content. In some embodiments, the media manifest file may be an MPD. The example method 3000 may further include: estimating 3004 the bandwidth available for streaming the light field video content, selecting 3006 one of the plurality of subsampled representations, and obtaining 3008 the selected subsampled representation. Some embodiments of the example method 3000 may include: interpolating 3010 one or more interpolated subviews from the selected subsampled representation, the interpolation being performed using information in the manifest file corresponding to the one or more views respectively. Some embodiments of the example method 3000 may further include: synthesizing 3012 one or more synthesized views from the one or more interpolated subviews. Some embodiments of the example method 3000 may further include: displaying 3014 the one or more synthesized views.
[0350] In some embodiments, example method 3000 may further include: requesting light field video content from a server. In some embodiments of example method 3000, retrieving a selected subsampled representation from a server may include: requesting the selected subsampled representation from the server and receiving the selected subsampled representation. In some embodiments, the example apparatus may include a processor and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods described above. Some embodiments of the example method may include: obtaining light field content associated with the selected representation; generating a view of the light field content from the obtained light field content; and rendering the generated view to a display. Some embodiments of the example method may include: receiving a media manifest file including information on a plurality of subsampled representations for the view of the light field video content; estimating the bandwidth available for streaming the light field video content; selecting one of the plurality of subsampled representations; obtaining the selected subsampled representation; interpolating one or more interpolated subviews from the selected subsampled representation, the interpolation being performed using information in the manifest file corresponding to the one or more views respectively; synthesizing one or more synthesized views from the one or more interpolated subviews; and displaying the one or more synthesized views. Some embodiments of the apparatus may include a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to perform any of the example methods listed above.
[0351] Figure 31This is a flowchart illustrating an example process for a viewing client according to some embodiments. Some embodiments of the example method 3100 for the viewing client may include: receiving 3102 a media manifest file, which includes information on a plurality of view subsets of light field content. Example method 3100 may further include: selecting 3104 one of the plurality of view subsets and obtaining the light field content associated with the selected view subset. In some embodiments, example method 3100 may include: decoding 3108 a frame of the light field content. Example method 3100 may include: combining 3110 two or more views represented in the frame to generate a combined view composition result. Example method 3100 may further include: rendering 3112 the combined view composition result to a display. For some implementations of the exemplary method 3100, retrieving the MPD may include: requesting and receiving the MPD. For some embodiments, regarding... Figure 30 and 31 The described example methods can be executed by a client device. In some embodiments, the light field content can be a three-dimensional light field content. In some embodiments, the example apparatus may include a processor and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods described above.
[0352] In some embodiments, generating a view of light field content may include: decoding a frame of the light field content; and combining two or more views represented in the frame to generate a generated view such that the obtained (or, in some embodiments, retrieved) light field content includes the frame of the light field content. In some embodiments, an example method may include: receiving a media manifest file containing information of a plurality of view subsets of the light field content; selecting one view subset from the plurality of view subsets; obtaining light field content associated with the selected view subset; decoding a frame of the light field content; combining two or more views represented in the frame to generate a composite view result; and rendering the composite view result to a display. In some embodiments, an example method may include: receiving a media manifest file containing information of a plurality of view subsets of light field content; selecting one view subset from the plurality of view subsets; obtaining light field content associated with the selected view subset; generating a view from the obtained light field content; and rendering the generated view to a display. For some embodiments of the sample methods, generating one or more views from the obtained light field content may include: decoding a frame of the light field content; and combining two or more views represented in the frame to generate a generated view such that the obtained light field content includes the frame of the light field content. Some embodiments of the apparatus may include a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the example methods listed above.
[0353] Figure 32 This is a flowchart illustrating an example process for a viewing client according to some embodiments. In some embodiments, example process 3200 may include: receiving 3202 a media manifest file identifying a plurality of representations of a multi-view video, such that at least a first representation of the plurality of representations may include a first subsample of a view, and at least a second representation of the plurality of representations includes a second subsample of a view different from the first subsample of the view. In some embodiments, example process 3200 may further include: selecting 3204 a selected representation from the plurality of representations. In some embodiments, example process 3200 may further include: retrieving 3206 the selected representation. In some embodiments, example process 3200 may further include: rendering 3208 the selected representation. In some embodiments, an apparatus may include a processor and a non-transitory computer-readable medium storing instructions that, when executed by the processor, operate to perform an embodiment of example process 3200.
[0354] Although methods and systems according to some embodiments have been discussed in the context of displays, some embodiments can also be applied to the context of virtual reality (VR), mixed reality (MR), and augmented reality (AR). Furthermore, although the term "head-mounted display (HMD)" is used herein according to some embodiments, for some embodiments, some embodiments can be applied to wearable devices capable of, for example, VR, AR, and / or MR (which may or may not be attached to the head).
[0355] Example methods according to some embodiments may include: receiving a media manifest file identifying a plurality of representations of a multi-view video, at least a first representation of the plurality of representations comprising a first subsample of a view, and at least a second representation of the plurality of representations comprising a second subsample of a view that is different from the first subsample of the view; selecting a selected representation from the plurality of representations; retrieving the selected representation; and rendering the selected representation.
[0356] In some embodiments of the example method, each of the plurality of representations may have a corresponding view density oriented toward a particular corresponding direction.
[0357] In some embodiments of the example method, the media manifest file may identify the corresponding view density and the specific corresponding orientation for one or more of the plurality of representations.
[0358] In some embodiments of the example method, two or more distinct subsamples of a view may differ at least with respect to the view density of the subsamples of the view oriented toward a particular direction.
[0359] Some embodiments of the example method may further include: tracking the user's viewing direction, wherein selecting the selected representation may include: selecting a view subsample from two or more subsamples of the view, the selected view subsample having a high view density toward the user's tracked viewing direction.
[0360] Some embodiments of the example method may also include: tracking the user's viewing direction, wherein selecting the selected representation may include: selecting the selected representation based on the tracked user's viewing direction.
[0361] Some embodiments of the example method may further include: obtaining the user's viewing direction, wherein selecting the representation of the selection includes: selecting the representation of the selection based on the obtained viewing direction of the user.
[0362] In some embodiments of the example method, the selection of the chosen representation may be based on the user's location.
[0363] In some embodiments of the example method, the selection of the chosen representation may be based on bandwidth constraints.
[0364] In some embodiments of the example method, at least one of the plurality of representations may include a view with a higher density for a first viewing direction than for a second viewing direction.
[0365] Some embodiments of the example method may further include: using the rendered representation to generate a signal for display.
[0366] Some embodiments of the example method may further include: tracking the user's head position, wherein the selection of the selected representation may be based on the user's head position.
[0367] Some embodiments of the example method may further include: tracking the user's gaze direction, wherein the selection of the selected representation may be based on the user's gaze direction.
[0368] Some embodiments of the example method may further include: using the user's gaze direction to determine the user's viewpoint, wherein selecting the selected representation may include: selecting the selected representation based on the user's viewpoint.
[0369] Some embodiments of the example method may further include: determining the user's viewpoint using the user's gaze direction; and selecting at least one subsample of a view of the multiview video, wherein selecting the at least one subsample of the view may include: selecting at least one subsample of the view within a threshold viewpoint angle of the user's viewpoint.
[0370] Some embodiments of the example method may also include: interpolating at least one view, wherein selecting a selected representation can be chosen from the plurality of representations and the at least one view.
[0371] In some embodiments of the example method, the media manifest file may include priority data for one or more views, and the at least one view is interpolated using the priority data.
[0372] In some embodiments of the example method, the media manifest file may include priority data for one or more views, and the selection of the view indicates the use of the priority data.
[0373] Some embodiments of the example method may further include: obtaining light field content associated with a selected representation; decoding a frame of the light field content; combining two or more views represented in the frame to generate a composite view result; and rendering the composite view result to a display.
[0374] In some embodiments of the example method, the frame may include a frame-packed representation of two or more views corresponding to the selected representation.
[0375] In some embodiments of the example method, for at least one of the plurality of representations, the media manifest file may include information corresponding to two or more views of the light field content.
[0376] For some embodiments of the example method, the selected representation may be selected based on at least one of the following criteria: user head position, gaze tracking, complexity of the content, display capability, and bandwidth available for retrieving the light field content.
[0377] For some embodiments of the example method, selecting the selected representation may include: predicting the user's viewpoint; and selecting the selected representation based on the predicted user viewpoint.
[0378] Some embodiments of the example method may also include: obtaining light field content associated with the selected representation; generating a generated view of the light field content from the obtained light field content; and rendering the generated view to a display.
[0379] For some embodiments of the example method, generating a view of the light field content may include: decoding a frame of the light field content; and combining two or more views represented in the frame to generate the generated view, wherein the obtained light field content may include the frame of the light field content.
[0380] Some embodiments of the example method may further include: decoding a frame of light field content; and combining two or more views represented in the frame to generate a combined view composition result, wherein the plurality of representations of the multi-view video may include plurality of subsamples of the views of the light field content, and wherein rendering the selected representation may include rendering the combined view composition result to a display.
[0381] Some embodiments of the example method may further include: requesting the media manifest file from a server; and requesting the light field content associated with a selected subset of views, wherein obtaining the light field content associated with the selected subset of views may include: performing a process selected from the group consisting of: retrieving the light field content associated with the selected subset of views from the server, requesting the light field content associated with the selected subset of views from the server, and receiving the light field content associated with the selected subset of views.
[0382] For some embodiments of the example method, combining two or more views represented in the frame may include using view composition techniques.
[0383] For some embodiments of the example method, the selection of one of the plurality of subsamples of the view may be based on at least one of the following criteria: user head position, gaze tracking, complexity of the content, display capability, and bandwidth available for retrieving the light field content.
[0384] For some embodiments of the example method, selecting one of the plurality of subsamples of the view may include: predicting the user's viewpoint; and selecting the subset of views based on the predicted user viewpoint.
[0385] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0386] Example methods according to some embodiments may include: receiving a media manifest file, the media manifest file including information on a plurality of subsampled representations of views for light field video content; selecting one of the plurality of subsampled representations; obtaining the selected subsampled representation; interpolating one or more interpolated subviews from the selected subsampled representation, the interpolation being performed using the information in the manifest file corresponding to the one or more views respectively; synthesizing one or more synthesized views from the one or more interpolated subviews; and displaying the one or more synthesized views.
[0387] Some embodiments of the example method may also include: estimating the bandwidth available for streaming the light field video content, such that the selection of a plurality of subsample representations is based on the estimated bandwidth.
[0388] Some embodiments of the example method may also include: tracking the user's location such that the selection of the subsampled representation among the plurality of subsampled representations is based on the user's location.
[0389] Some embodiments of the example method may also include: requesting the light field video content from a server, wherein obtaining the selected subsample representation may include: performing a process of selecting from a group consisting of: retrieving the selected subsample representation from the server, requesting the selected subsample representation from the server, and receiving the selected subsample representation.
[0390] For some embodiments of the example method, the information in the manifest file may include location data for two or more views.
[0391] For some embodiments of the example method, the information in the manifest file may include interpolation priority data for one or more of the plurality of views, and the selection of one of the plurality of subsampled representations may be based on the interpolation priority data for one or more of the plurality of views.
[0392] Some embodiments of the example method may also include: tracking the user's head position, wherein selecting one of the plurality of subsampled representations is based on the user's head position.
[0393] Some embodiments of the example method may also include: tracking the user's gaze direction, wherein selecting one of the plurality of sub-sample representations may be based on the user's gaze direction.
[0394] Some embodiments of the example method may further include: determining the user's viewpoint from the user's gaze direction; and selecting one or more subviews of the light field video content from a group including the one or more interpolated views and selected subsampled representations, wherein synthesizing the one or more synthesized views from the one or more interpolated subviews may include: synthesizing the one or more synthesized views of the light field using the one or more selected subviews and the user's viewpoint.
[0395] Some embodiments of the example method may also include: displaying the synthesized view of the light field.
[0396] For some embodiments of the example method, selecting one or more subviews of the light field may include selecting one or more subviews within a threshold viewpoint angle of the user's viewpoint.
[0397] Some embodiments of the example method may further include: determining the user's viewpoint from the user's gaze direction, wherein selecting one of the plurality of subsampled representations may include selecting the subsampled representation based on the user's viewpoint.
[0398] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0399] Example methods according to some embodiments may include: receiving a media manifest file including information of a plurality of subsamples of a view for light field content; selecting one of the plurality of subsamples of the view; obtaining light field content associated with the selected view subsample; decoding a frame of the light field content; combining two or more views represented in the frame to generate a combined view composition result; and rendering the combined view composition result to a display, wherein the obtained light field content includes the frame of the light field content.
[0400] Some embodiments of the example method may further include: requesting the media manifest file from a server; and requesting the light field content associated with a selected subset of views, wherein obtaining the light field content associated with the selected subset of views may include performing a process selected from the group consisting of: retrieving the light field content associated with the selected subset of views from the server, requesting the light field content associated with the selected subset of views from the server, and receiving the light field content associated with the selected subset of views.
[0401] In some embodiments of the example method, the frame may include a frame-packed representation of two or more views corresponding to a selected subset of views.
[0402] In some embodiments of the example method, the media manifest file may include information corresponding to two or more views that correspond to a selected subset of views.
[0403] For some embodiments of the example method, for at least one of the plurality of subsamples of the view, the media manifest file may include information corresponding to two or more views of the light field content.
[0404] For some embodiments of the example method, selecting one of the plurality of subsamples of the view may include: parsing the information of the plurality of subsamples of the view for light field content in the media manifest file.
[0405] For some embodiments of the example method, combining two or more views represented in the frame may include using view composition techniques.
[0406] For some embodiments of the example method, the selection of one of the plurality of subsamples of the view may be based on at least one of the following criteria: gaze tracking, complexity of the content, display capability, and bandwidth available for retrieving the light field content.
[0407] For some embodiments of the example method, selecting one of the plurality of subsamples of the view may include: predicting the user's viewpoint; and selecting the subset of views based on the predicted user viewpoint.
[0408] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0409] Example methods according to some embodiments may include: receiving a media manifest file including information of a plurality of subsamples of a view for light field content; selecting one of the plurality of subsamples of the view; obtaining light field content associated with the selected subset of the view; generating a view from the obtained light field content; and rendering the generated view to a display.
[0410] For some embodiments of the example method, generating a view from the obtained light field content may include: interpolating the views from the light field content associated with a selected subset of views to generate the generated view, the interpolation being performed using information in the manifest file corresponding to the views respectively.
[0411] For some embodiments of the example method, generating one or more views from the obtained light field content may include: decoding a frame of the light field content; and combining two or more views represented in the frame to generate the generated view, wherein the obtained light field content may include the frame of the light field content.
[0412] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0413] Example methods according to some embodiments may include: receiving a media manifest file identifying a plurality of subsamples of views of a multiview video, the plurality of subsamples of the views including two or more views of different densities; selecting a selected subsample from the plurality of subsamples of the views; retrieving the selected subsample; and rendering the selected subsample.
[0414] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0415] Example methods according to some embodiments may include: rendering a representation of a view comprising a full array of light field video content; sending the rendered full array representation of the view; obtaining the current viewpoint of a client; using the current viewpoint and a viewpoint motion model to predict future viewpoints; prioritizing a plurality of subsampled representations of the view of the light field video content; rendering the prioritized plurality of subsampled representations of the view of the light field video content; and sending the prioritized plurality of subsampled representations of the view.
[0416] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0417] Example methods according to some embodiments may include: selecting a plurality of subviews of light field video content; generating streaming data for each of the plurality of subviews of the light field video content; and generating a media manifest file including the streaming data for each of the plurality of subviews of the light field video content.
[0418] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0419] Example methods according to some embodiments may include: receiving a request for information about light field video content; when the request is a new session request, sending a media manifest file containing information about a plurality of subsampled representations of a view of the light field video content; and when the request is a subset data fragment request, sending a data fragment containing a subset of the light field video content.
[0420] Example apparatuses according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operable, when executed by the processor, to perform any of the methods listed above.
[0421] An example signal according to some embodiments may include a signal carrying a view representation of the full array of the light field video content and a plurality of subsampled representations of the view of the light field video content.
[0422] Example signals according to some embodiments may include a signal carrying multiple sub-views of light field video content.
[0423] Example signals according to some embodiments may include a signal carrying streaming data of each of a plurality of sub-views of light field video content.
[0424] An example signal according to some embodiments may include a signal that carries information about a plurality of subsampled representations of a view of light field video content.
[0425] Example signals according to some embodiments may include a signal carrying a data segment comprising a subset of light field video content.
[0426] Note that the various hardware elements of the one or more embodiments described are referred to as “modules,” and their implementation (i.e., execution, operation, etc.) is in conjunction with the various functions described herein for each module. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices) that a person skilled in the art would consider suitable for a given implementation. Each described module may also include instructions executable to perform one or more functions described as being performed by the corresponding module, and note that these instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, and / or software instructions, or include them, and may be stored in any suitable non-transitory computer-readable medium or media, such as commonly referred to as RAM, ROM, etc.
[0427] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor memory devices, magnetic media (e.g., internal hard disks and removable disks), magneto-optical media, and optical media (e.g., CD-ROMs and DVDs). The processor associated with the software can be used to implement radio frequency transceivers for WTRUs, UEs, terminals, base stations, RNCs, and / or any host computer.
Claims
1. A method for rendering light fields, comprising: Receive a media manifest file that identifies multiple representations of light field video content, the light field video content comprising multiple subviews of a complete light field image, at least a first representation of the multiple representations comprising a first subset of the multiple subviews, and at least a second representation of the multiple representations comprising a second subset of the multiple subviews, wherein each of the first subset and the second subset comprises multiple subviews of the multiple subviews, and wherein the second subset is different from the first subset. Choose one representation from the plurality of representations; Retrieve the selected representation; as well as Based on the retrieved selected representation, render the content to the display.
2. The method of claim 1, wherein each of the plurality of representations has a corresponding view density oriented toward a particular corresponding direction.
3. The method of claim 2, wherein, The media manifest file identifies the corresponding view density and the specific corresponding orientation for one or more of the plurality of representations.
4. The method of claim 1, wherein the first subview subset and the second subview subset differ at least in terms of the view density of the subview subset oriented in a particular direction.
5. The method according to any one of claims 1 to 4, further comprising: Track the user's viewing direction, The selection of the selected representation includes: selecting a subset of subviews from the first subset of subviews and the second subset of subviews, wherein the selected subset of subviews has a high view density toward the viewing direction of the tracked user.
6. The method of claim 1, wherein the selection of the selected representation is based on the user's location.
7. The method of claim 1, wherein, At least one of the plurality of representations includes a view with a higher density for a first viewing direction than for a second viewing direction.
8. The method according to claim 1, further comprising: The user's viewpoint is determined using the user's gaze direction. The selection of the selected representation includes: selecting the selected representation based on the user's viewpoint.
9. The method according to claim 1, further comprising: Interpolate at least one view. The selected representation is chosen from the plurality of representations and the at least one view.
10. The method according to claim 9, The media manifest file includes priority data for one or more views, and The priority data is used for interpolation of at least one view.
11. The method according to claim 1, The media manifest file includes priority data for one or more views, and The selected option represents the use of the priority data.
12. The method according to claim 1, further comprising: Obtain the light field content associated with the selected representation; Decode the frames containing the light field content; Combine two or more views represented in the frame to generate a composite view result; as well as The composite view is rendered onto the display.
13. The method of claim 12, wherein the frame comprises a frame-packed representation of two or more views corresponding to the selected representation.
14. The method of claim 1, wherein the selected representation is selected based on at least one of the following criteria: user head position, gaze tracking, content complexity, display capability, and bandwidth available for retrieving light field content.
15. The method according to any one of claims 12-14, wherein selecting the selected representation comprises: Predict the user's viewpoint; as well as The selected representation is chosen based on the predicted viewpoint of the user.
16. The method of claim 1, further comprising: Determine the viewpoint position at the first moment; as well as Based on the motion estimation model of the device, a prediction of the viewpoint position at the second moment is obtained; The selection of one representation from the plurality of representations is based on the prediction of the viewpoint position.
17. An apparatus for rendering light fields, comprising: processor; as well as A non-transitory computer-readable medium storing instructions that, when executed by the processor, operate to perform the method according to any one of claims 1 to 16.
18. A method for rendering light fields, comprising: Receive a media manifest file, the media manifest file including information on multiple subsample representations of a view for light field video content; Select a sub-sampling representation from the plurality of sub-sampling representations; Obtain the selected subsample representation; Using the information in the manifest file that corresponds to the one or more views respectively, interpolate one or more subviews from the selected subsampled representation; One or more composite views are synthesized from the one or more subviews; as well as Display the one or more composite views.
19. The method of claim 18, further comprising: Estimate the bandwidth available for streaming the light field video content. The selection of a subsample representation from the plurality of subsample representations is based on the estimated bandwidth.
20. The method of claim 18, further comprising: Tracking user location, The selection of a subsample representation from the plurality of subsample representations is based on the user's location.
21. The method according to any one of claims 18-20, wherein the information in the manifest file includes location data of two or more views.
Citation Information
Patent Citations
Light Field Display Device and Method
US20140035959A1
Methods and apparatus for light-field imaging
US8290358B1
Metrics and messages to improve experience for 360-degree adaptive streaming
WO2018175855A1
Image processing method, terminal, and server
WO2019024521A1