Light Field / Immersive Media Split Rendering Using Edge Cloud Architecture and Peer-to-Peer Streaming

An edge-cloud based split rendering architecture dynamically allocates tasks between edge and cloud architectures to address the computational challenges in immersive media, resulting in improved user experience and efficient content streaming.

JP7693825B2Active Publication Date: 2025-06-17TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023558857
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-07
Filing Date
2022-10-14
Publication Date
2025-06-17
Estimated Expiration
2042-10-14

AI Technical Summary

Technical Problem

Current immersive media technologies face challenges in efficiently processing and rendering high-quality light field and holographic content, particularly in terms of computational power and latency, which affects user experience.

Method used

The implementation of an edge-cloud based architecture for split rendering, where tasks are dynamically divided between edge and cloud architectures based on factors like sampling delay, computational delay, and networking delay, using adaptive streaming methods such as depth priority or asset priority.

Benefits of technology

This approach reduces the processing requirements on end devices, enhances user experience by minimizing latency, and allows for the efficient streaming of high-quality light field and immersive media content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693825000001
    Figure 0007693825000001
  • Figure 0007693825000002
    Figure 0007693825000002
  • Figure 0007693825000003
    Figure 0007693825000003
Patent Text Reader

Abstract

Systems and methods for split rendering for light field or immersive media by using edge cloud and peer-to-peer based architectures. The systems and methods include using a combination of cloud-based devices and edge devices to provide distributed processing related to streaming of media, particularly light field or immersive media, to end user devices. The systems and methods further include using multiple cloud and edge devices to provide parallel streaming of a given media package to end user devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 272,654, filed Oct. 27, 2021, and U.S. Patent Application No. 17 / 962,122, filed Oct. 7, 2022, the disclosures of which are incorporated herein by reference in their entireties.

[0002] The disclosed subject matter relates to methods and systems for an edge - cloud - based architecture for split rendering of immersive media.

Background Art

[0003] Immersive media is defined by immersive technologies that attempt to create or simulate the physical world through digital simulation, thereby stimulating any or all of the human sensory systems and creating the perception that the user is physically present in the scene.

[0004] Current ongoing immersive media technologies include various types such as virtual reality (VR), augmented reality (AR), mixed reality (MR), light field / holographic, etc. VR refers to replacing the user's physical environment with a digitally generated world by placing the user in it using a headset. On the other hand, AR incorporates digital media and overlays them onto the real - world environment around the user using either clear - vision glasses or a smartphone. MR refers to creating an environment where the technology and the physical world can co - exist by fusing the real world and the digital world.

[0005] Light field / holographic technology is composed of light rays in 3D space, and the light rays come from each point and each direction. This technology is based on the concept that everything visible to the eye is illuminated by light coming from all light sources, travels through space, hits the surface of an object, where the light is partially absorbed and partially reflected onto another surface before reaching the eye. The light rays that exactly reach our eyes depend on the exact position of the user within the light field. As the user moves around, the user perceives a part of the light field and uses the perceived part to obtain an idea about the position of the object.

[0006] Light rays can be defined by a 5D plenoptic operation where each ray can be defined by three coordinates (3D) in 3D space and two angles specifying the direction in 3D space.

[0007] Conventional cameras can only capture a 2D representation of the light rays reaching the camera lens at a given position. The image sensor records the sum of the brightness and color of all the light rays reaching each pixel.

[0008] When capturing content for a light field or holographic-based display, a light field camera that can capture not only the brightness and color but also the direction of all the light rays reaching the camera sensor is required. Using this information, a digital scene that accurately represents the origin of each light ray can be reconstructed, making it possible to accurately digitally reconstruct the captured scene in 3D.

[0009] Currently, two main technologies are being used to capture such a stereoscopic scene. The first approach is to use an array of cameras or camera modules to capture different light rays / fields of view from each direction. The second approach is to use a depth camera that can capture 3D information in a single exposure without the need for structured illumination by measuring the depth of multiple objects under controlled lighting conditions. Summary of the Invention

Means for Solving the Problem

[0010] The following presents a simplified overview of such embodiments in order to provide a basic understanding of one or more embodiments of the present disclosure. This overview is not an extensive overview of all contemplated embodiments, nor is it intended to identify key or important elements of all embodiments or to delineate the scope of any or all embodiments. Its sole purpose is to present, in a simplified form, some concepts of one or more embodiments of the present disclosure as a prelude to the more detailed description presented later.

[0011] According to an exemplary embodiment, a method for performing light field or immersive media streaming in a network including a cloud architecture, an edge architecture, and an end-user device, the method being executed by at least one processor. This method of light field or immersive media streaming Priority of Adaptive Streaming Method includes dividing tasks associated with the light field or immersive media streaming into a plurality of computing tasks based on the above, a first set of the plurality of computing tasks being executed on the cloud architecture, a second set of the plurality of computing tasks being executed on the edge architecture, and the second set of the plurality of computing tasks being The priority is different from and non-overlapping with the first set of the plurality of computing tasks. The adaptive streaming method includes adaptive streaming based on depth priority or adaptive streaming based on asset priority. This method further includes streaming the light field or immersive media from the cloud architecture and the edge architecture to the end-user device. Based on the adaptive streaming method

[0012] ​According to an exemplary embodiment, a system for light field or immersive media streaming comprising a network, wherein the network comprises a cloud architecture and an edge architecture, the network being operable to stream light field or immersive media from the cloud architecture and the edge architecture to an end-user device, and comprising at least one memory configured to store computer program code, and at least one processor configured to access the at least one memory and operate as instructed by the computer program code. In this system, the computer program code comprises split code configured to cause the at least one processor to split a plurality of calculation tasks associated with the light field or immersive media streaming into a plurality of calculation tasks Priority of Adaptive Streaming Method based on Priority of Adaptive Streaming Method , the first set of the plurality of calculation tasks being executed on the cloud architecture and the second set of the plurality of calculation tasks being executed on the edge architecture, the second set of the plurality of calculation tasks being The priority is different from and non-overlapping with the first set of the plurality of calculation tasks. The adaptive streaming method includes adaptive streaming based on depth priority or adaptive streaming based on asset priority. The system streams the light field or immersive media from the cloud architecture and the edge architecture to the end-user device based on the adaptive streaming method.

[0013] According to an exemplary embodiment, a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method for light field or immersive media streaming. This method for light field or immersive media streaming Priority of Adaptive Streaming Method includes splitting tasks associated with the light field or immersive media streaming into a plurality of calculation tasks based on Priority of Adaptive Streaming Method , the first set of the plurality of calculation tasks being executed on the cloud architecture and the second set of the plurality of calculation tasks being executed on the edge architecture, the second set of the plurality of calculation tasks being The priority is different from and non-overlapping with the first set of the plurality of calculation tasks. The adaptive streaming method includes adaptive streaming based on depth priority or adaptive streaming based on asset priority.This method further includes streaming the light field or immersive media from the cloud architecture and the edge architecture to the end user device. Based on the adaptive streaming method

[0014] Additional embodiments are described in the following description, and will become apparent, in part, from the description and / or learned by practice of the presented embodiments of the disclosure.

[0015] The foregoing and other objects, features, and advantages of the disclosure will be apparent from the following detailed description taken in conjunction with the accompanying drawings in which like parts are labeled with like reference numerals. BRIEF DESCRIPTION OF THE DRAWINGS

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

[0017] The images in the drawings are simplified for illustrative purposes and are not drawn to scale. In the description of the drawings, like elements are labeled with like names and reference numerals as in the previous drawings. The specific numbers assigned to the elements are provided merely to assist the description and do not imply any limitation (structural or functional) to the present invention.

[0018] The following detailed description of the exemplary embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0019] ​The foregoing disclosure provides examples and explanations, but is not intended to be exhaustive or to limit the implementation forms strictly to the disclosed forms. Modifications and variations are possible in light of the above disclosure, or may be obtained from the practice of the implementation forms. Further, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Further, in the flowcharts and descriptions of operations provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be executed simultaneously (at least partially), and the order of one or more operations may be switched.

[0020] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the embodiments. Accordingly, the operations and behaviors of the systems and / or methods are described herein without reference to specific software code, and it is understood that the software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0021] Even if specific combinations of features are recited in the claims and / or disclosed herein, these combinations are not intended to limit the disclosure of possible implementation forms. In fact, many of these features may not be specifically recited in the claims and / or may be combined in ways not disclosed herein. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible implementation forms includes each dependent claim in combination with all other claims in the claim set.

[0022] Elements, operations, or instructions used in this specification should not be construed as important or essential unless explicitly described as such. Also, as used in this specification, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." When only one item is intended, the term "one" or similar words are used. Also, terms such as "has," "have," "having," "include," "including," etc. used in this specification are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise specified. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0023] Throughout this specification, references to "one embodiment," "an embodiment," or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present solution. Thus, appearances of the phrases "in one embodiment," "in an embodiment," and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0024] Furthermore, the described features, advantages, and characteristics of the present disclosure can be combined in any suitable manner in one or more embodiments. Those skilled in the art will recognize, in light of the description herein, that the present disclosure can be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not exist in all embodiments of the present disclosure may be recognized in a particular embodiment.

[0025] Embodiments of the present disclosure relate to a system and method for split rendering for light field or immersive media by using an edge cloud and peer-to-peer based architecture. Currently, all computing devices rely on the improvement of computer capabilities. According to Moore's law, in addition to the improvement of processing speed and the reduction of chip size, as a result of adding more processing cores and bandwidth, computing power continues to increase exponentially. However, in high-performance applications, the requirements for capacity and processing requirements are increasing. To bridge this gap, an edge cloud-based rendering architecture is needed.

[0026] Cloud-based computing provides end users with more ways to access the vast computing power required for the latest computer graphics. There is currently a discussion among those skilled in the art about how much processing can be adapted in smartphones and any AR / VR device to render high-quality video frames, but by shifting to cloud rendering, games can be further enhanced. Therefore, if a sufficiently fast connection to very powerful online computing resources is available, all small devices have the potential to become supercomputers that can stream real-time video and games.

[0027] In one embodiment, a cloud-edge based split rendering architecture can be used for light field / immersive media streaming. This architecture reduces the processing requirements of the end device itself. For example, the end device may not require a GPU to provide an acceptable user experience. The task splitting between the edge and the cloud may be dynamic, i.e., the task splitting between the edge and the cloud may be based on factors such as sampling delay, computational delay including image processing and frame rendering delay, and networking delay including queuing delay and transmission delay. Referring to FIG. 1, scene 101 is divided into three components or scenes (102, 103, 104). The end device (108) based on different decision parameters streams scene 1 (107) and scene 3 (106) from the cloud (105) and streams scene 2 (109) from the edge (110). To avoid ambiguity, references to "scene" herein are merely examples, and "scene" should be understood to include any media susceptible to streaming.

[0028] In the same embodiment, the task division between the edge and the cloud may also be based on an adaptive streaming approach. As two types of adaptive streaming methods that can be adopted here, 1) adaptive streaming based on scene depth, which renders assets based on depth priority instead of rendering the entire scene at once, and 2) adaptive streaming based on asset priority, in which assets with higher priority are rendered first. Therefore, instead of allocating the entire task at once, the task is divided based on priority and allocated sequentially. For example, refer to FIG. 2 in which scene 201 is divided into three priorities (202, 203, 204). The end device (208) streams the parts with priority 1 (207) and priority 3 (206) of scene 201 from the cloud, and streams the part with priority 2 (209) of scene 201 from the edge. When a task is allocated, first, according to the available computing power (computing delay including sampling delay, image processing and frame rendering delay, networking delay including queuing delay and transmission delay), the part with higher priority is streamed from the edge or the cloud.

[0029] In another embodiment, the server / CDN may divide scene 301 into a plurality of scenes (302, 303, 304, 305) that can be stored in a plurality of edge-cloud systems (306, 310, 311, 314), thereby enabling the end user to process each scene (307, 308, 312, 313) in parallel. This also means that instead of connecting to a single edge or cloud, the end user can connect to a plurality of edges and clouds as shown in FIG. 3. This is based on the fact that when processing scenes in parallel, no extra delay occurs in the processing compared to the case of connecting to a single edge or cloud.

[0030] In another embodiment, a peer-to-peer streaming protocol may be applied to the streaming of light field / immersive media. This eliminates the need to connect each device to the cloud. Thus, for each scene, the server will have the scene file, as well as the hash of the scene file. The hash of the scene file indicates the source from which the scene file is downloaded. The end user can then decide whether to download the scene from the edge cloud or perform a peer-to-peer download, and can select an appropriate peer when the scene file is provided by multiple peers.

[0031] The method of split rendering for light field or immersive media using an edge cloud architecture and peer-to-peer streaming may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, FIG. 4 shows a computer system 400 suitable for implementing a particular embodiment of the disclosed subject matter.

[0032] The computer software can be encoded using any suitable machine code or computer language that can be subject to mechanisms such as assembly, compilation, linking, etc. to create code containing instructions that can be executed directly by, for example, a computer central processing unit (CPU), a graphics processing unit (GPU), or the like, or through interpretation, execution of microcode, etc.

[0033] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, and Internet of Things devices.

[0034] The components shown in FIG. 4 with respect to computer system 400 are essentially illustrative and are not intended to suggest any limitation as to the use or functionality of the computer software implementing the embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiments of computer system 400.

[0035] Computer system 400 may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), voice input (e.g., voice, clapping), visual input (e.g., gesture), olfactory input (not shown). The human interface device can also be used to capture certain media that is not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0036] The input human interface device may include one or more (only one of each is shown) of keyboard 401, mouse 402, trackpad 403, touch screen 410, data glove (not shown), joystick 405, microphone 406, scanner 407, camera 408.

[0037] Computer system 400 may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users via tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by touch screen 410, data glove (not shown), or joystick 405, although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speaker 409, headphones (not shown), etc.), visual output devices (such as screens 410 including CRT screens, LCD screens, plasma screens, OLED screens, regardless of whether each has a touch screen input function and regardless of whether each has a tactile feedback function, some of which may be capable of outputting two-dimensional visual output or output of three dimensions or more via means such as stereographic output, virtual reality glasses (not shown), holographic display, and smoke tank (not shown)), and may also include a printer (not shown).

[0038] Computer system 400 may also include a memory device accessible by humans, as well as optical media 420 including CD / DVD ROM / RW having a CD / DVD or similar medium 421, thumb drive 422, removable hard drive or solid state drive 423, legacy magnetic media such as tapes and floppy disks (not shown), and related media of memory devices such as dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0039] Also, those skilled in the art should understand that the term "computer-readable medium" used in connection with the subject matter of this disclosure does not include a transmission medium, carrier wave, or other transient signal.

[0040] The computer system 400 can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital networks for television including cable television, satellite television, and terrestrial television, vehicle and industrial including CANBus, etc. Certain networks generally require an external network interface adapter connected to a specific general-purpose data port or peripheral bus (449) (such as a USB port of the computer system 400, etc.), and others are generally integrated into the core of the computer system 400 by connection to the system bus as described below (such as an Ethernet interface 435 to a PC computer system or a cellular network interface 433 to a smartphone computer system). Using any of these networks, the computer system 400 can communicate with other entities. Such communication can be, for example, unidirectional, receive only (such as broadcast television), transmit only unidirectionally (such as CANbus to a specific CANbus device), or bidirectional using a local or wide area digital network to other computer systems. Specific protocols and protocol stacks can be used with each of the networks and network interfaces described above.

[0041] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core 440 of the computer system 400.

[0042] The core 440 can include one or more central processing units (CPUs) 441, a graphics processing unit (GPU) 442, a dedicated programmable processing device in the form of a field programmable gate area (FPGA), and a hardware accelerator 444 for specific tasks, etc. These devices can be connected via a system bus 448 together with a read-only memory (ROM) 445, a random access memory 446, an internal mass storage 447 such as an internal hard drive or SSD that is not accessible to the user. In some computer systems, the system bus 448 may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs and GPUs, etc. Peripheral devices can be attached directly to the system bus 448 of the core or can be attached via a peripheral bus 449. Architectures for peripheral buses include PCI, USB, etc.

[0043] The CPU 441, GPU 442, FPGA 443, and accelerator 444 can execute specific instructions that can be combined to form the aforementioned computer code. The computer code can be stored in the ROM 445 or RAM 446. Migration data can also be stored in the RAM 446, while persistent data can be stored, for example, in the internal mass storage 447. Fast storage and retrieval for any of the memory devices can be enabled using a cache memory that can be closely associated with one or more of the CPUs 441, GPUs 442, mass storage 447, ROM 445, RAM 446, etc.

[0044] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well-known and available to persons having skill in the computer software arts.

[0045] Rather than being limiting, as an example, a computer system having an architecture 400, particularly a core 440, can provide functionality as a result of software embodied on one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be user-accessible mass storage as introduced above, and media associated with specific storage of the core 440 that is non-transitory in nature, such as mass storage 447 internal to the core or ROM 445. The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core 440. The computer-readable media can include one or more memory devices or chips depending on specific requirements. The software can cause the core 440, particularly the processors (including CPU, GPU, and FPGA, etc.) therein, to define data structures stored in the RAM 446 and modify such data structures according to processes defined by the software, thereby causing the core 440 to execute specific processes or specific portions of specific processes described herein. Additionally, or alternatively, the computer system can provide functionality as a result of circuitry (e.g., accelerator 444) wired or otherwise embodied with logic that can operate instead of or in conjunction with software to execute specific processes described herein or specific portions of specific processes. References to software can, if desired, include logic, and vice versa. References to computer-readable media can, if desired, include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0046] Although the present disclosure describes several exemplary embodiments, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Thus, it will be understood by those skilled in the art that, although not explicitly illustrated or described herein, many systems and methods can be devised that embody the principles of the present disclosure and are thus within its spirit and scope.

Explanation of Signs

[0047] 101 Scene 102 Scene 103 Scene 104 Scene 105 Cloud 106 Scene 3 107 Scene 1 108 End Device 109 Scene 2 110 Edge 201 Scene 202 Priority 203 Priority 204 Priority 206 Priority 3 207 Priority 1 208 End Device 209 Priority 2 210 Edge 301 Scene 302 Scene 303 Scene 304 Scene 305 Scene 306 Edge Cloud System 307 Scene 308 Scene 310 Edge Cloud System 311 Edge Cloud System 312 Scene 313 Scene 314 Edge Cloud System 400 Computer System 401 Keyboard 402 Mouse 403 Trackpad 405 Joystick 406 Microphone 407 Scanner 408 Camera 409 Speaker 410 Touch Screen 420 Optical Media 421 Media 422 Thumb Drive 423 Removable Hard Drive or Solid State Drive 433 Cellular Network Interface 435 Ethernet Interface 440 Core 441 Central Processing Unit (CPU) 442 Graphics Processing Unit (GPU) 443 Field Programmable Gate Array (FPGA) 444 Hardware Accelerator 445 Read Only Memory (ROM) 446 Random Access Memory (RAM) 447 Internal Mass Storage 448 System Bus 449 Peripheral Bus

Claims

1. A method for light field or immersive media streaming in a network comprising a cloud architecture, an edge architecture, and an end-user device, the method being executed by at least one processor, The step of dividing tasks associated with the light field or immersive media streaming into a plurality of computing tasks based on the priorities of adaptive streaming techniques, wherein a first set of the plurality of computing tasks is executed on the cloud architecture, a second set of the plurality of computing tasks is executed on the edge architecture, the second set of the plurality of computing tasks having a different priority and non-overlapping with the first set of the plurality of computing tasks, and the adaptive streaming technique includes adaptive streaming based on depth priority or adaptive streaming based on asset priority, And the step of streaming the light field or immersive media from the cloud architecture and the edge architecture to the end-user device based on the adaptive streaming technique.

2. The method according to claim 1, wherein the network is operable to implement a peer-to-peer streaming protocol.

3. The cloud architecture includes a plurality of cloud devices, the edge architecture includes a plurality of edge devices, and the method further includes, The step of dividing the light field or immersive media into a plurality of components, The step of storing each component of the plurality of components in the plurality of cloud devices and the plurality of edge devices, And the step of delivering the plurality of components to the end-user device in parallel. The method according to claim 1.

4. A system for light field or immersive media streaming, comprising, A network, the network including a cloud architecture and an edge architecture, the network being operable to stream a light field or immersive media from the cloud architecture and the edge architecture to an end-user device, a network; A system comprising at least one processor configured to execute the method according to any one of claims 1 to 3.

5. A computer program for causing one or more processors to execute the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • System and method for distributing the processing load of realistic image formation

    JP2012511200A

  • Multidimensional 3D engine computing of virtual world or real world and dynamic load dispersion of virtualization base

    JP2021119453A

  • Edge Compute Systems and Methods

    US20200244723A1