Methods, apparatuses, systems, and storage media for streaming of light field / immersive media

CN116868107BActive Publication Date: 2026-09-25TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280015927.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-10-07
Filing Date
2022-10-14
Publication Date
2026-09-25
Estimated Expiration
2042-10-14

AI Technical Summary

Technical Problem

[0008]传统相机只能捕获在给定位置到达相机镜头的光线的二维表示

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116868107B_ABST
    Figure CN116868107B_ABST
Patent Text Reader

Abstract

A system and method for split rendering of light field or immersive media using an edge-cloud and peer-to-peer network based architecture is disclosed. The system and method includes using a combination of cloud-based devices and edge devices to provide distributed processing related to streaming of media, particularly light field or immersive media, to end user devices. The system and method also includes using multiple cloud and edge devices to provide parallel streaming of a given media data packet to end user devices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 272,654, filed October 27, 2021, and U.S. Patent Application No. 17 / 962,122, filed October 7, 2022, the disclosures of which are incorporated herein by reference. Technical Field

[0003] The disclosed topics relate to edge-cloud based architectures and systems for split rendering of immersive media. Background Technology

[0004] Immersive media is defined by immersive technologies that attempt to create or mimic the physical world through digital simulation, thereby stimulating any or all of the human sensory systems to create the feeling that the user is physically present in the scene.

[0005] Currently, different types of immersive media technologies are at work: Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), light field / holography, etc. VR refers to a digital environment that places the user in a computer-generated world using headphones, replacing the user's physical environment. AR, on the other hand, uses digital media and overlays it onto the real world around you using clear glasses or a smartphone. MR refers to the fusion of the real and digital worlds, creating an environment where technology and the physical world can coexist.

[0006] Light field / holographic technology consists of light rays in three-dimensional space, originating from every point and direction. This technology is based on the concept that everything we see is illuminated by light from any light source. This light travels through space, reaching the surfaces of objects, where it is partially absorbed and partially reflected before reaching our eyes. What kind of light reaches our eyes depends on the user's precise position within the light field, and as the user moves around, they perceive a portion of the light field and use this perceived portion to determine the location of objects.

[0007] A ray can be defined by 5D full optical operations, where each ray can be defined by three coordinates and two angles in 3D space to specify its direction in 3D space.

[0008] Traditional cameras can only capture a two-dimensional representation of the light rays arriving at the camera lens at a given location. Image sensors, on the other hand, record the sum of the brightness and color of all the light rays arriving at each pixel.

[0009] When it comes to capturing content for light field-based or holographic displays, a light field camera is needed, which can capture not only brightness and color, but also the direction of all the light rays arriving at the camera sensor. Using this information, a digital scene can be reconstructed with a precise representation of the origin of each ray, making it possible to digitally reconstruct a precisely captured scene in 3D.

[0010] Currently, two main techniques are used to capture such volumetric scenes. The first technique uses camera arrays or camera modules to capture different light / views from each direction. The second technique involves using a depth camera, which can capture 3D information in a single exposure by measuring the depth of multiple objects under controlled lighting conditions, without requiring structured lighting. Summary of the Invention

[0011] The following presents a simplified summary of one or more embodiments of this application to provide a basic understanding of these embodiments. This summary is not a broad overview of all contemplated embodiments and is intended neither to identify key or essential elements of all embodiments nor to describe the scope of any or all embodiments. Its sole purpose is to present certain concepts of one or more embodiments of this application in a simplified form as a prelude to the more detailed description that follows.

[0012] According to an exemplary embodiment, a method for streaming light field or immersive media in a network including a cloud architecture, an edge architecture, and an end-user device is provided, the method being executed by at least one processor. The method for streaming light field or immersive media includes: splitting tasks associated with the streaming of the light field or immersive media into a plurality of computational tasks based on one or more latency factors, wherein a first group of the plurality of computational tasks is executed on the cloud architecture, a second group of the plurality of computational tasks is executed on the edge architecture, and the second group of computational tasks is different from and does not overlap with the first group of computational tasks. The method further includes: streaming the light field or immersive media from the cloud architecture and the edge architecture to the end-user device.

[0013] According to an exemplary embodiment, a system for streaming light field or immersive media is provided, the system comprising: a network including a cloud architecture and an edge architecture, wherein the network is operable to stream light field or immersive media from the cloud architecture and the edge architecture to an end-user device; at least one memory for storing computer program code; and at least one processor for accessing the at least one memory and operating as instructed by the computer program code. In this system, the computer program code includes splitting code configured such that the at least one processor splits a task associated with streaming the light field or immersive media into a plurality of computational tasks based on one or more latency factors, wherein a first group of the plurality of computational tasks is executed on the cloud architecture, a second group of the plurality of computational tasks is executed on the edge architecture, and the second group of the plurality of computational tasks is different from and does not overlap with the first group of the plurality of computational tasks.

[0014] According to an exemplary embodiment, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by a processor, cause the processor to perform a method for streaming light field or immersive media. This method for streaming light field or immersive media includes: dividing a task associated with the streaming of the light field or immersive media into a plurality of computational tasks based on one or more latency factors, wherein a first group of the plurality of computational tasks is executed on a cloud architecture, a second group of the plurality of computational tasks is executed on an edge architecture, and the second group of computational tasks is different from and does not overlap with the first group of computational tasks. The method further includes: streaming the light field or immersive media from the cloud architecture and the edge architecture to an end-user device.

[0015] Additional embodiments will be set forth in the description which follows, and will be apparent in part from the description, and / or may be learned by practicing the embodiments presented in this application. Attached Figure Description

[0016] The above and other objects, features and advantages of this application will be apparent from the following detailed description taken in conjunction with the accompanying drawings, in which similar parts are given the same reference numerals, and wherein:

[0017] Figure 1 An edge-cloud architecture with scenario allocation is described.

[0018] Figure 2 An edge-cloud architecture with scenario priorities is shown.

[0019] Figure 3 A multi-layered edge-cloud architecture is shown.

[0020] Figure 4 This is a schematic diagram of a computer system.

[0021] For illustrative purposes, the images in the accompanying drawings have been simplified and are not depicted to scale. In the description of the drawings, similar elements are given similar names and reference numerals as in the previous drawings. The specific reference numerals assigned to elements are for illustrative purposes only and do not imply any limitation (structural or functional) on the invention. Detailed Implementation

[0022] The following detailed description of the exemplary embodiments is with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0023] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible based on the foregoing disclosure, or may be obtained from practice of the embodiments. Furthermore, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Moreover, in the flowcharts and descriptions of operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be switched.

[0024] It is evident that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code—it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0025] Even if specific combinations of features are listed in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features can be combined in ways not specifically listed in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible embodiments includes combinations of each dependent claim with each other claim in the claim set.

[0026] Unless explicitly stated otherwise, no element, action, or instruction used herein should be construed as critical or necessary. Furthermore, as used herein, the article “a / an” is intended to include one or more items and may be used interchangeably with “one or more.” If only one item is intended to be used, the term “one” or similar terminology is used. Additionally, as used herein, the terms “have,” “possess,” “contain,” “include,” “comprise,” or similar terms are intended to be open-ended terms. Further, unless explicitly stated otherwise, the word “based on” means “at least partially based on.” Moreover, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.

[0027] Throughout this specification, references to "an embodiment," "an embodiment," or similar language mean that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of this application. Therefore, the phrases "in an embodiment," "in one embodiment," and similar language in this specification may, but do not necessarily, refer to the same embodiment.

[0028] Furthermore, in one or more embodiments, the features, advantages, and characteristics described in this application may be combined in any suitable manner. Based on the description herein, those skilled in the art will recognize that this application may be practiced without one or more specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not be present in all embodiments of this application may be recognized in some embodiments.

[0029] Embodiments of this application relate to a system and method for split rendering of light fields or immersive media using an edge-cloud and peer-to-peer network-based architecture. Currently, all computing devices rely on increased computing power. According to Moore's Law, in addition to increased speed and smaller chip size, computing power continues to grow exponentially due to the addition of more processing cores and bandwidth. However, with the emergence of high-performance applications, the demand for capacity and processing is increasing. To bridge this gap, an edge-cloud-based rendering architecture is needed.

[0030] Cloud-based computing offers end users more ways to access the massive computing power required for modern computer graphics. Amidst the current debate about how much processing power an average technologist can perform to render high-quality video frames on a smartphone or any AR / VR device, this shift has moved towards cloud rendering, as it elevates the level of gaming. Therefore, if connected to very powerful online computing resources at sufficiently fast speeds, all small devices could potentially become high-performance supercomputers capable of streaming real-time video and games.

[0031] In one embodiment, a cloud-edge-based split rendering architecture can be used for light field / immersive media streaming. This architecture reduces the processing demands on the end device itself. For example, the end device may not require a GPU to provide an acceptable user experience. Task splitting between the edge and cloud can be dynamic; that is, task splitting between the edge and cloud can be based on factors such as sampling latency, computational latency including image processing and frame rendering latency, and networking latency including queuing and transmission latency. Reference Figure 1 Scenario 101 is divided into three components or scenarios (102, 103, 104). The terminal device (108) streams scenario 1 (107) and scenario 3 (106) from the cloud (105) and scenario 2 (109) from the edge (110) based on different decision parameters. For the avoidance of doubt, the term "scenario" as used herein is for illustrative purposes only, and "scenario" should be understood to include any media susceptible to streaming.

[0032] In the same embodiment, task splitting between the edge and the cloud can also be based on adaptive streaming technology. Two adaptive streaming methods can be used here: 1) scene depth-based adaptive streaming, where assets are rendered based on depth priority instead of rendering the entire scene at once, and 2) asset priority-based adaptive streaming, where priority assets are rendered first. Therefore, instead of allocating the entire task at once, tasks are split based on priority and allocated sequentially. For example, see [reference]. Figure 2 Scene 201 is divided into three priorities (202, 203, 204). A terminal device (208) streams portions of scene 201 with priorities 1 (207) and 3 (206) from the cloud, and a portion of scene 201 with priority 2 (209) from the edge. When allocating tasks, portions with higher priorities are first streamed from the edge or cloud based on their available computing power (sampling latency, computational latency including image processing and frame rendering latency, and networking latency including queuing and transmission latency).

[0033] In another embodiment, the server / CDN can split scenario 301 into multiple scenarios (302, 303, 304, 305) that can be stored across multiple edge-cloud systems (306, 310, 311, 314), thereby allowing end users to process each scenario in parallel (307, 308, 312, 313). This also means that end users can connect to multiple edges and clouds, rather than connecting to a single edge or cloud, as... Figure 3 As shown. This is based on the fact that parallel processing scenarios do not add additional processing latency compared to connecting to a single edge or cloud.

[0034] In another embodiment, a peer-to-peer streaming protocol can be applied to streaming light field / immersive media. This avoids the need for each device to connect to the cloud. Therefore, for each scene, the server will have a scene file along with a hash of that scene file. The hash of the scene file will indicate the source from which to download the scene file. The end user will then be able to decide whether they want to download the scene from the edge-cloud or via peer-to-peer, and select the appropriate peer when multiple peers can provide the scene file.

[0035] Techniques for split rendering of light fields or immersive media using edge-cloud architecture and peer-to-peer streaming can be implemented as computer software with computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 4 A computer system 400 suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0036] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by a computer's central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode, or similar means.

[0037] This instruction can be executed on various types of computers or their components, including personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things devices.

[0038] Figure 4 The components of the computer system 400 shown are exemplary in nature and are not intended to impose any limitation on the scope or functionality of the computer software implementing the embodiments of this application. The configuration of the components should also not be construed as having any dependency or requirement relating to any one or a combination of components shown in the exemplary embodiments of the computer system 400.

[0039] Computer system 400 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users, for example, through input such as: tactile input (e.g., keystrokes, swipes, data glove movement), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images acquired from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0040] The input human-machine interface device may include one or more of the following (only one of each is shown): keyboard 401, mouse 402, touchpad 403, touch screen 410, data glove (not depicted), joystick 405, microphone 406, scanner 407, camera 408.

[0041] Computer system 400 may also include certain human-machine interface (HMI) output devices. Such HMI output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. These HMI output devices may include tactile output devices (e.g., tactile feedback via touchscreen 410, data gloves (not depicted), or joystick 405, but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers 409, headphones (not depicted)), visual output devices (e.g., screen 410, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input, each with or without tactile feedback—some of which are capable of outputting two-dimensional or more than three-dimensional visual output via devices such as stereoscopic output), virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).

[0042] The computer system 400 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 420 with media such as CD / DVD 421, finger drives 422, removable hard disk drives or solid-state drives 423, conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0043] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0044] Computer system 400 may also include interfaces to one or more communication networks. These networks may be, for example, wireless networks, wired networks, or optical networks. Networks may also be local area networks, wide area networks, metropolitan area networks, vehicle and industrial networks, real-time networks, latency-tolerant networks, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CANBus, etc. Some networks typically require external network interface adapters (e.g., USB ports of computer system 400; other network interfaces are typically integrated into the core of computer system 400 by connecting to a system bus, such as Ethernet interface 435 connected to a PC computer system or cellular network interface 433 connected to a smartphone computer system) to connect to certain general-purpose data ports or peripheral buses (449). Computer system 400 can use any of these networks to communicate with other entities. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANbus connected to certain CANbus devices), or bidirectional, such as connecting to other computer systems using a local area network (LAN) or wide area network (WAN) digital network. As mentioned above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.

[0045] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel 440 of the computer system 400.

[0046] Core 440 may include one or more central processing units (CPUs) 441, graphics processing units (GPUs) 442, dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) 443, hardware accelerators 444 for certain tasks, etc. These devices, as well as read-only memory (ROM) 445, random access memory 446, and internal mass storage 447 such as internal non-user-accessible hard disk drives (SDs), etc., may be connected via system bus 448. In some computer systems, system bus 448 may be accessed via one or more physical connectors to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the core's system bus 448 or connected via peripheral bus 449. Peripheral bus architectures include PCI, USB, etc.

[0047] CPU 441, GPU 442, FPGA 443, and accelerator 444 can execute certain instructions, which can be combined to form the aforementioned computer code. This computer code can be stored in ROM 445 or RAM 446. Transient data can also be stored in RAM 446, while permanent data can be stored, for example, in internal mass storage 447. Fast storage and retrieval to any storage device can be achieved using a cache, which can be closely associated with one or more CPUs 441, GPU 442, mass storage 447, ROM 445, RAM 446, etc.

[0048] Computer-readable media may have computer code thereon for performing various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0049] As a non-limiting example, a computer system having architecture 400, particularly kernel 440, can provide functionality by having one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as some non-transitory memory of kernel 440, such as internal kernel mass storage 447 or ROM 445. Software implementing various embodiments of this application can be stored in such means and executed by kernel 440. Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause kernel 440, particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 446 and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system may provide functionality through hard-wired or otherwise embodied logic in circuitry (e.g., accelerator 444), which may replace or operate with the software to perform a particular process or a specific portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., integrated circuits (ICs)), circuitry embodying logic for execution, or both. This application includes any suitable combination of hardware and software.

[0050] Although several exemplary embodiments have been described in this application, modifications, substitutions, and various equivalent substitutions that fall within the scope of this application are possible. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, while not explicitly shown or described herein, embody the principles of this application and thus fall within its spirit and scope.

Claims

1. A method for streaming light field or immersive media, performed in a network including a cloud architecture, an edge architecture, and end-user devices, said method being executed by at least one processor, characterized in that, The method includes: Based on one or more latency factors and adaptive streaming techniques, the tasks associated with streaming the light field or immersive media are broken down into multiple computational tasks and assigned according to priority. The adaptive streaming techniques include depth-priority-based adaptive streaming or asset-priority-based adaptive streaming. A first group of the computational tasks is executed on the cloud architecture, and a second group of the computational tasks is executed on the edge architecture. The second group of computational tasks is different from and does not overlap with the first group of computational tasks. The light field or immersive media is streamed from the cloud architecture and the edge architecture to the end-user device.

2. The method according to claim 1, characterized in that, The task associated with streaming the light field or immersive media is broken down into multiple computational tasks, including: Analyze one or more predetermined metrics associated with the network, the one or more predetermined metrics including the latency factor; and Based on the analysis of the one or more predetermined metrics associated with the network, a given computational task among the plurality of computational tasks associated with the streaming of the light field or immersive media is assigned to the cloud architecture or the edge architecture.

3. The method according to claim 2, characterized in that, The one or more predetermined metrics include one or more of the following: sampling latency, computation latency, image processing load, frame rendering latency, and networking latency, wherein the networking latency includes queuing and transmission latency.

4. The method according to claim 2, characterized in that, The network is capable of operating to implement peer-to-peer streaming protocols.

5. The method according to claim 1, characterized in that, The cloud architecture includes multiple cloud devices, and the edge architecture includes multiple edge devices. The method further includes: The light field or immersive media is divided into multiple components; Each of the plurality of components is stored on the plurality of cloud devices and the plurality of edge devices; and The multiple components are transmitted to the end user equipment in parallel.

6. The method according to claim 5, characterized in that, The task associated with streaming the light field or immersive media is broken down into multiple computational tasks, including: Analyze one or more predetermined metrics associated with the network, including the latency factor; and Based on the analysis of the one or more predetermined metrics, a given computing task among the plurality of computing tasks associated with the streaming of the light field or immersive media is assigned to a given cloud device among the plurality of cloud devices or a given edge device among the plurality of edge devices; The one or more predetermined metrics include: sampling latency, computation latency, image processing load, frame rendering latency, and network latency, wherein the network latency includes: queuing and transmission latency.

7. A system for streaming light field or immersive media, characterized in that, include: A network comprising a cloud architecture and an edge architecture, wherein the network is operable to stream light fields or immersive media from the cloud architecture and the edge architecture to end-user devices; At least one memory for storing computer program code; and At least one processor is configured to access the at least one memory and operate according to instructions of the computer program code to perform the method as claimed in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that, It has instructions stored thereon that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Edge Compute Systems and Methods

    US20200244723A1

  • Apparatus and method for real time graphics processing using local and cloud-based graphics processing resources

    US20210097641A1