Interactive video streaming for 3D applications

By using a combination of GPU instances and WebRTC APIs in streaming video applications, dynamically adjusting 3D streaming sessions solves the problems of high cost and complexity in the prior art, and implements a high-quality, low-latency global scalable streaming solution.

CN120283395AActive Publication Date: 2025-07-08MONKEYWAY GMBH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202380069892.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-29
Publication Date
2025-07-08
Estimated Expiration
2043-09-29

AI Technical Summary

Technical Problem

Existing streaming video applications require deep technical expertise when providing high-quality, low-latency immersive experiences, making it difficult to create reliable and scalable solutions globally, and existing approaches are often expensive and complex.

Method used

The graphical processing unit (GPU) instance is used to directly host streaming services, operate in peer-to-peer communication networks through the WebRTC API, and render and encoding an interactive 3D environment using the encapsulated 3D streaming engine, dynamically adjust streaming sessions to achieve efficient video and audio streaming, and optimize the streaming process in combination with artificial intelligence systems and deep learning technology.

Benefits of technology

Provides high visual quality and low latency streaming experience, reduces setup costs, supports high-performance streaming worldwide, avoids technical lock-in, adapts to different user needs, and is affordable, easy to understand and implement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283395A_ABST
    Figure CN120283395A_ABST
Patent Text Reader

Abstract

Systems and methods for video processing may be provided that generate 3D streaming data via an API. 3D streaming sessions may be delivered at a client device with low latency and high throughput capacity. Upon verifying the client device, a server may be selected to support an encapsulated 3D streaming engine to initiate the session. A library file may be injected into an executable file of the encapsulated 3D streaming engine to generate a digital representation of the interactive 3D environment. The interactive 3D environment may be rendered, encoded, and streamed to the client device via a GPU. The 3D streaming session may be adjusted via the client device, and maintained until termination at the client device. The number of virtual compute instances of an encapsulated 3D streaming engine provided by the server may be dynamically adjusted based on metrics and machine learning predictions of connected client devices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 377,986, filed Sep. 30, 2022. The entire teachings of the above application are incorporated herein by reference. BACKGROUND OF THE INVENTION

[0003] Streaming video applications have advanced in a new dimension recently, attempting to provide users with a more immersive experience. As virtual alternatives to in-person interactions continue to evolve, opportunities are provided for users to securely and reliably connect across great geographical distances. SUMMARY OF THE INVENTION

[0004] Embodiments of the present disclosure provide techniques for streaming interactive video and audio content of interest to users. The streaming techniques provide users with a high-quality, low-latency video and audio streaming experience. Graphics Processing Unit (GPU) instances are configured to directly host services that support the streaming content, thereby facilitating the implementation of the method without middleware or plugins. Embodiments of the present invention relate to systems and methods for video processing and for streaming processed media content to client devices of these users.

[0005] In some embodiments, a video processing system includes a plurality of servers that operate in a peer-to-peer communication network via a WebRTC Application Programming Interface (API) to generate 3D stream data. The plurality of servers may receive a request for a 3D stream session at a session handler at a first client device, in part, by receiving context verification about the first client device from the session handler communicating with the first client device, the context verification including the Internet Protocol (IP) location of the first client device. The plurality of servers may respond to the context verification about the first client device by assigning a first server of the plurality to initiate a 3D stream session at the first client device based on the IP location of the first client device.

[0006] Continuing, the plurality of servers may direct a first instance of an encapsulated 3D stream engine at the first server to initiate a 3D stream session at the first client device via the WebRTC API. The plurality of servers may inject a library file of the WebRTC API into an executable file of the encapsulated 3D stream engine to generate a digital representation of an interactive 3D environment for the 3D stream session at the first client device.

[0007] Multiple servers can continue to render a digital representation of an interactive 3D environment within the buffer of the GPU, encode the rendered representation for transmitting a 3D streaming session to a first client device, and stream the encoded representation of the interactive 3D environment at the first client device. The multiple servers can dynamically adjust the 3D streaming session of the interactive 3D environment in response to input received from the first client device.

[0008] The multiple servers can maintain the digital representation of the interactive 3D environment of the 3D streaming session by iteratively performing the processes involving the above-mentioned rendering, encoding, streaming, and adjustment until the 3D streaming session terminates at the first client device.

[0009] The first server can be configured to establish a predetermined number of virtual computing instances (VCIs) of a 3D streaming session with one or more client devices, and the first server can be configured to automatically adjust the number of provided VCIs based on a change in the number of client devices. At least one of the one or more client devices can be the first client device.

[0010] At least one of the multiple servers can be configured to respond to requests for a 3D streaming session from a client device by assigning a server among the multiple servers. The first server can be one of the multiple servers.

[0011] At least one of the multiple servers can perform a computational analysis to determine an appropriate number of 3D streaming sessions of multiple client devices based on a predetermined number of VCIs. The computed analysis can include concurrent users (CCU), session time, daily active users (DAU), monthly active users (MAU), and sessions. The encapsulated 3D streaming engine can configure the 3D streaming session as a video streaming cluster such that at least one of the multiple servers performs a multicast transmission of data packets containing at least a portion of the digital representation of the interactive 3D environment to one or more client devices.

[0012] Dynamically adjusting the 3D streaming session of the interactive 3D environment in response to input received from the first client device can further include that the encapsulated 3D streaming engine is configured to capture events associated with the interactive 3D environment. The events can be detected by one or more listeners at one or more of the client devices. The dynamic adjustment can further include that the encapsulated 3D streaming engine is configured to capture metadata, consumption information, and interaction data during the 3D streaming session at one or more of the client devices. The dynamic adjustment can further include that the encapsulated 3D streaming engine is configured to respond to the captured events, metadata, consumption information, and interaction data by redrawing the interactive 3D environment at one or more of the client devices to maintain the digital representation of the interactive 3D environment.

[0013] The encapsulated 3D streaming engine can maintain a digital representation of an interactive 3D environment in the following ways: calculating composite 3D video data based on captured events, metadata, consumption information, and interaction data; defining a composite image layout based on attributes derived from a 3D streaming session at one or more client devices; configuring the 3D streaming session to provide a composite video signal according to the defined composite image layout; and transmitting the composite video signal in the form of data packets to one or more client devices via a packetizer. The encapsulated 3D streaming engine can interface with one or more client devices to facilitate a control communication interface, thereby enabling voice, text, and video transmission among one or more client devices in a data packet exchange communication system of a peer-to-peer network.

[0014] The encapsulated 3D streaming engine can be configured as a bound container containing software components. The software components can include at least one of the following: library files of the WebRTC API, an event manager, a scene manager, a resource manager, a session manager, a physics manager, and an artificial intelligence system. The WebRTC API can include one or more WebRTC API functions configured to call one or more software components. The encapsulated 3D streaming engine can not be a software plug-in.

[0015] Inputs can be received through a data channel established using the RTCDataChannel API. The encapsulated 3D streaming engine can be configured to train an interactive frame prediction model for the digital representation of an interactive 3D environment based on the encoded stream data of a 3D streaming session and events, metadata, interaction data, and consumption data from one or more client devices.

[0016] The encapsulated 3D streaming engine can be configured to encode a 3D streaming session by selectively using prediction frames generated by a training-based interactive frame prediction model, and transmit the trained interactive frame prediction model and the encoded stream data to a first client device and a second client device to create a digital representation of an interactive 3D environment. The first client device and the second client device can be configured to receive the trained interactive frame prediction model and the encoded stream data, and decode the encoded stream data based on the trained interactive frame prediction model to create a digital representation of an interactive 3D environment.

[0017] In some embodiments, a video processing method includes configuring a server computer system to respond to a session handler receiving a request for a 3D stream session at a first client device. Responding to the session handler can include receiving, from the session handler, context verification regarding the first client device, the context verification including the IP location of the first client device. Responding to the session handler can include guiding a 3D stream session at the first client device to respond to the context verification regarding the first client device by assigning a first server among a plurality of servers of the server computer system based on the IP location of the first client device.

[0018] Responding to the session handler can include instructing a first instance of an encapsulated 3D stream engine at the first server to initiate a 3D stream session at the first client device via a WebRTC API. Responding to the session handler can include injecting a library file of the WebRTC API into an executable file of the encapsulated 3D stream engine to generate a digital representation of an interactive 3D environment for the 3D stream session at the first client device.

[0019] Responding to the session handler can include rendering a digital representation of the interactive 3D environment within a buffer of a GPU, encoding the rendered representation for transmission of the 3D stream session to the first client device, and streaming the encoded representation of the interactive 3D environment at the first client device. Responding to the session handler can include dynamically adjusting the 3D stream session of the interactive 3D environment in response to input received from the first client device.

[0020] Responding to the session handler can include maintaining a digital representation of the interactive 3D environment of the 3D stream session by iteratively performing processes involving the above-mentioned rendering, encoding, streaming, and adjustment until the 3D stream session terminates at the first client device. In an embodiment of the video processing method, the method performs operations to implement any embodiment or combination of embodiments described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The foregoing will be apparent from a more particular description of example embodiments as illustrated in the accompanying drawings, in which like reference numerals refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis being placed upon illustrating the embodiments shown.

[0022] Figure 1 is a schematic block diagram of an example video processing system according to the present disclosure.

[0023] Figure 2 is a flowchart of an example video processing method according to the present disclosure.

[0024] Figure 3 is a schematic block diagram of an example embodiment of a video processing system.

[0025] Figure 4 is a schematic block diagram of an example embodiment of a video processing system.

[0026] Figure 5A is a schematic diagram and overview of an example use case of an embodiment of a video processing system.

[0027] Figure 5B depicts Figure 5A various example user interfaces of an example video processing system use case of

[0028] Figure 6 is a schematic diagram of an example computer network environment in which an embodiment of the present invention is deployed.

[0029] Figure 7 is Figure 6 a block diagram of a computer node in a network of Detailed Description

[0030] The description of the example embodiment is as follows.

[0031] A high-performance, low-cost streaming solution is described herein. This solution provides high visual quality and low latency, while being easily implementable by users without programming skills as a complete solution.

[0032] System Architecture

[0033] Figure 1 is a schematic block diagram of an example embodiment of a video processing system 100. In the example embodiment, the video processing system 100 includes a plurality of servers 105-1, 105-n, which operate in a peer-to-peer (i.e., P2P) communication network 110 via a WebRTC application programming interface (API) 115 to generate 3D stream data 120. The plurality of servers 105-1, 105-n can respond to a request 130 for a 3D stream session received at a session handler 125 at a first client device 135-1 by partially receiving context verification 140 about the first client device 135-1 from the session handler 125 communicating with the first client device 135-1, the context verification including the Internet Protocol (IP) location of the first client device 135-1. The plurality of servers 105-1, 105-n can initiate a 3D stream session 145 at the first client device 135-1 in response to the context verification 140 about the first client device 135-1 by assigning the first server 105-1 among the plurality 105-1, 105-n based on the IP location of the first client device 135-1.

[0034] In an example, the context verification 140 may include a distance metric based on the distance between the respective servers 105-1, 105-n and the first client device 135. In an embodiment, the context verification 140 regarding the first client device 135 may include the number of network hops - the time taken by intermediate routers or servers to process signals transmitted from the respective servers 105-1, 105-n to the first client device 134. Synthetic monitoring tools may be deployed from the respective servers 105-1, 105-n to run artificial calls to the first client 145 and detect any increase in latency or degradation in performance. Based on the context verification data analysis, a server that delivers the lowest latency and highest throughput capacity at the client may be selected from the servers 105-1, 105-n.

[0035] Continuing, the multiple servers 105-1, 105-n may cause a first instance of the encapsulated 3D streaming engine 150 at the first server 105-1 to initiate a 3D streaming session 145 at the first client device 135-1 via the WebRTC API 115. The multiple servers 105-1, 105-n may inject the library file 155 of the WebRTC API 115 into the executable file 160 of the encapsulated 3D streaming engine 145 to generate a digital representation 165 of an interactive 3D environment for the 3D streaming session 145 at the first client device 135-1.

[0036] The multiple servers 105-1, 105-n may continue to render the digital representation 165 of the interactive 3D environment within the buffer of the graphics processing unit (GPU) 170, encode the rendered representation for transmission of the 3D streaming session 145 to the first client device 135-1, and stream the encoded representation of the interactive 3D environment at the first client device 135-1. The multiple servers 105-1, 105-n may dynamically adjust the 3D streaming session 145 of the interactive 3D environment in response to an input 175 received from the first client device 135-1.

[0037] The multiple servers 105-1, 105-n may maintain the digital representation 165 of the interactive 3D environment of the 3D streaming session 145 by iteratively performing the processes involving the above-mentioned rendering, encoding, streaming, and adjustment until the 3D streaming session 145 terminates at the first client device 135-1.

[0038] Figure 2FIG. 0 is a flow chart of an example embodiment of a video processing method 200. In the example embodiment, the video processing method 200 includes configuring a first server 105-1 of a plurality of servers 105-1, 105-n of a server computer system, such as system 100, to respond to a request 130 for a 3D streaming session received at a first client device 135-1 by a session handler 125. In method 200, responding to the session handler 125 may include receiving 204 from the session handler a context verification 140 for the first client device 135-1, the context verification including the IP location of the first client device 135-1. In method 200, responding to the session handler 125 may include initiating 209 a 3D streaming session 145 at the first client device 135-1 by assigning, based on the IP location of the first client device 135-1, the first server 105-1 of the plurality of servers 105-1, 105-n of the server computer system in response to the context verification 140 for the first client device 135-1.

[0039] In method 200, responding to the session handler 125 may include instructing 214 a first instance of an encapsulated 3D streaming engine 150 at the first server 105-1 to initiate 3D streaming session 145 at the first client device 135-1 via a WebRTC API 115. In method 200, responding to the session handler 125 may include injecting 219 a library file 155 of the WebRTC API into an executable file 160 of the encapsulated 3D streaming engine 150 to generate a digital representation 165 of an interactive 3D environment for the 3D streaming session 145 at the first client device 135-1.

[0040] In method 200, responding to the session handler 125 may include rendering 224 the digital representation 165 of the interactive 3D environment within a buffer of a GPU 170, encoding 229 the rendered representation for transmission of the 3D streaming session 145 to the first client device 135-1, and streaming 234 the encoded representation of the interactive 3D environment at the first client device 135-1. In method 200, responding to the session handler 125 may include dynamically adjusting 239 the 3D streaming session 145 of the interactive 3D environment in response to an input 175 received from the first client device 135-1.

[0041] In the method 200, responding to the session handler 125 may include maintaining 244-1 the digital representation 165 of the interactive 3D environment of the 3D streaming session 145 by iteratively performing the processes involving the aforementioned rendering 224, encoding 229, streaming 234, and adjusting 239 until the 3D streaming session 145 is terminated 244-2 at the first client device 244. In response to terminating 224-2 the 3D streaming session 145, the method 200 may include terminating 249 streaming the digital representation 165 of the interactive 3D environment.

[0042] Return to Figure 1 As discussed above, the first server 135-1 of the system 100 may be configured to establish a predetermined number of virtual computing instances (VCIs) 180-1, 180-j for a 3D streaming session 145 with one or more client devices 135-1, 135-2, 135-i, and the first server 105-1 may be configured to automatically adjust the number of provided VCIs 180-1, 180-j based on a change in the number of client devices 135-1, 135-2, 135-i. At least one of the one or more client devices 135-1, 135-2, 135-i may be the first client device 135-1.

[0043] exist Figure 1 In the example system 100 of , a first client device 135-1, a second client device 135-2, and an i-th client device 135-i are shown; however, it should be noted that there may be only the first client device 135-1, only the first client device 135-1 and the second client device 135-2, or another number i of client devices 135-1, 135-2, 135-i. Similarly, it should be noted that although in Figure 1 In the example system 100 of , a first virtual computing instance (VCI 1) 180-1, a second virtual computing instance (VCI 2) 180-2, and a jth virtual computing instance (VCI j) are shown, but there may be only the first virtual computing instance (VCI 1) 180-1, only the first virtual computing instance (VCI 1) 180-1 and the second virtual computing instance (VCI 2) 180-2, or another number j of virtual computing instances (VCI 1, VCI 2, ..., VCI j) 180-1, 180-2, 180-j. Similarly, it should be noted that there may be any number of servers 105-1, 105-n, including one server 105-1 or another number n of servers 105-1, 105-n.

[0044] At least one of the plurality of servers 105-1, 105-n can be configured to respond to a corresponding request 130 for a 3D streaming session 145 by assigning a server among the plurality of servers 105-1, 105-n to respond to the request 130 for the 3D streaming session 145 from the client devices 135-1, 135-i. The first server 105-1 can be one of the plurality of servers 105-1, 105-n.

[0045] At least one of the plurality of servers 105-1, 105-n can perform a computational analysis to determine an appropriate number of 3D streaming sessions 145 of the plurality of client devices 135-1, 135-i adjusted based on a predetermined number of VCIs 180-1, 180-j. The computed analysis can include concurrent users (CCU), session time, daily active users (DAU), monthly active users (MAU), and sessions. The encapsulated 3D streaming engine 150 can configure the 3D streaming session 130 as a video stream cluster such that at least one of the plurality of servers 105-1, 105-n performs a multicast transmission of at least a portion of the data packets containing the digital representation 165 of the interactive 3D environment to one or more client devices 135-1, 135-i.

[0046] Dynamically adjusting the 3D streaming session 145 of the interactive 3D environment in response to the input 175 received from the first client device 135-1 can further include the encapsulated 3D streaming engine 150 being configured to capture events associated with the interactive 3D environment. The events can be detected by one or more listeners at one or more of the client devices 135-1, 135-i. The dynamic adjustment can further include the encapsulated 3D streaming engine 150 being configured to capture metadata, consumption information, and interaction data during the 3D streaming session 145 at one or more of the client devices 135-1, 135-i. The dynamic adjustment can further include the encapsulated 3D streaming engine 175 being configured to respond to the captured events, metadata, consumption information, and interaction data by redrawing the interactive 3D environment at one or more of the client devices 135-1, 135-i to maintain the digital representation 165 of the interactive 3D environment.

[0047] The encapsulated 3D streaming engine 150 can maintain the digital representation 165 of an interactive 3D environment in the following ways: calculating composite 3D video data based on the captured events, metadata, consumption information, and interaction data, defining a composite image layout based on attributes derived from 3D streaming sessions 145 at one or more client devices 135-1, 135-i, configuring the 3D streaming session to provide a composite video signal according to the defined composite image layout, and transmitting the composite video signal in the form of data packets to one or more client devices 135-1, 135-i via a packetizer. The encapsulated 3D streaming engine 150 can interface with one or more client devices 135-1, 135-i to facilitate a control communication interface, thereby enabling voice, text, and video transmission among one or more client devices 135-1, 135-i in the data packet exchange communication system of the peer-to-peer network 110.

[0048] The encapsulated 3D streaming engine 150 can be configured as a bound container containing software components. The software components can include at least one of the following: library files 155 of the WebRTC API 115, an event manager, a scene manager, a resource manager, a session manager, a physics manager, and an artificial intelligence system. The WebRTC API 115 can include one or more WebRTC API functions configured to call one or more software components. The encapsulated 3D streaming engine 150 can not be a software plug-in.

[0049] The encapsulated 3D streaming engine 150 can help improve the scalability of HTTP streaming and the scalability of the artificial intelligence system. The artificial intelligence system can include HTTP-based dynamic adaptive streaming (MPEG-DASH) and deep learning video streaming architectures, such as deep neural networks, deep belief networks, deep reinforcement learning, recurrent neural networks, and convolutional neural networks. The encapsulated 3D streaming engine 150 communicating with the artificial intelligence system can help configure the client devices 135-1, 135-I as self-learning HTTP adaptive streaming clients, such as State-Action-Reward-State-Action (SARSA) or Q-learning (model-free reinforcement learning algorithms).

[0050] Inputs 175 can be received through a data channel established using the RTCDataChannel API. The encapsulated 3D streaming engine 150 can be configured to train an interactive frame prediction model for the digital representation 165 of an interactive 3D environment based on the encoded stream data of the 3D streaming session 145 and events, metadata, interaction data, and consumption data from one or more client devices 135-1, 135-i.

[0051] Example embodiments can implement streaming via an encapsulated 3D streaming engine 150 as a service for high-end 3D applications that utilize ultra-low latency streaming technology. Any cloud provider can provide servers 105-1, 105-n for the systems and methods described herein, such as Amazon Web Services (AWS), Google Cloud Platform (GCP), Azure, and OnPremise. Additionally, the encapsulated 3D streaming engine 150 can be embedded with one or more game engines, such as the Unreal engine, Unity, etc. The encoder can include, for example, NVIDIA nvenc for encoding the 3D streaming session 145. Such an encoder can use, for example, a codec (such as an h264 codec) to encode the 3D streaming session 145. The API 115 can be configured as a WebRTC API or another API for an ultra-low latency video connection between client devices 135-1, 135-2, 135-i and GPU instances 170 provided by cloud-based servers 105-1, 105-n.

[0052] Figure 3 is a schematic block diagram of an example embodiment of a video processing system; in particular, its server 305 provides a GPU 370 to support a 3D streaming session 145. In an example, the encapsulated 3D streaming engine 150 is provided by a DirectX-based 3D game application 350. The library file 155 of the WebRTC API 115 can be injected via a DirectX injector 319. The video data containing images can be rendered 319 in a buffer of the GPU 370 after being received from the 3D game application 350. The rendered video data can be passed to a streaming server 315 (which can be based on WebRTC) and can be configured to encode the video data according to an NVIDIA nvenc codec. The WebRTC server 315 can be configured to control 310 an interface to interact with the 3D game application 350. The encoded video data can be streamed 320 by the WebRTC server 315 over the Internet.

[0053] The encapsulated 3D streaming engine 150 may include multiple execution instances, which may be distributed across multiple servers 16, 105-1, 105-n. The encapsulated 3D streaming engine 150 may be configured to encode a 3D streaming session 145 by selectively using predicted frames generated by a training-based interactive frame prediction model, and transmit the trained interactive frame prediction model and the encoded stream data to a first client device and a second client device to create a digital representation 165 of an interactive 3D environment. The first client device 135-1 and the second client device 135-2 may be configured to receive the trained interactive frame prediction model and the encoded stream data, and decode the encoded stream data based on the trained interactive frame prediction model to create a digital representation 165 of an interactive 3D environment.

[0054] In an example embodiment, the encapsulated 3D streaming engine 150 may be configured to train an interactive frame prediction model. The interactive frame prediction model may be configured and trained based on data from a streaming session, a 3D model, user input, metadata associated with a client session, and consumption data. The encapsulated 3D streaming engine 150 interfaces with an encoder and a decoder to predict frames of a 3D stream using the trained interactive frame prediction model. For example, the encoder may be configured to encode 3D stream data and transmit the trained interactive frame prediction model. The decoder may receive the 3D stream data and transmit the trained interactive frame prediction model, and decode the encoded stream data based on the trained interactive frame prediction model, and provide the reconstructed 3D stream data at one or more clients 135-i, 135-1, 135-2.

[0055] In an example embodiment, the encapsulated 3D streaming engine 150 may be configured with deep neural network (DNN)-based interactive frame prediction for video coding. For example, the interactive frame prediction may be trained and configured using DNN to help improve video encoding and decoding efficiency. The DNN may be implemented at both the encoder and the decoder. The DNN may use previously decoded frames and the trained interactive frame prediction model to predict the current frame. This may be configured as a separate interactive frame prediction mode prediction mode, which may be optimized to compete with other prediction modes. The DNN may be implemented to avoid transmitting motion vectors. In this way, the DNN may train the interactive frame prediction model to perform uni-directional and bi-directional prediction, which may provide significant coding efficiency gains relative to High Efficiency Video Coding (H.265 and MPEG-H Part 2). DNN-based frame prediction.

[0056] In an embodiment, the encapsulated 3D streaming engine 150 can be configured to generate and process a scene graph for transmitting a 3D data packet stream. The scene graph can be generated using machine learning and machine vision computing prediction systems to provide powerful scene graph generation (SGG). For example, the encapsulated 3D streaming engine 150 can perform scene graph generation (SGG), which can be enhanced with a powerful semantic representation that defines the scene. Scene graph generation (SGG) involves automatically mapping an image or video into a semantic structure scene graph. This process can include encoding position data about detected objects and their corresponding relationships, and this process can be adjusted using deep learning techniques to improve system performance.

[0057] The scene graph generation (SGG) model can be processed by the encapsulated 3D streaming engine 150 to render 3D video from a visual base scene graph. The scene graph can be a structured representation that can capture detailed semantics by explicitly modeling objects. The semantic structure of the scene graph can be processed using perceptual statistics to calculate a mapping indicating which regions of a video frame are important to the human eye. This process can incorporate motion estimation and motion vectors for inter-frame prediction, resulting in an encoded bitstream. A motion vector quality metric can be used to construct a vector map, which can be used to encode the bitstream using metrics such as block variance and block luminance. Applying the vector map and the scene graph generation (SGG) model in 3D video coding within a model-based compression framework (such as continuous block tracking (CBT)) can improve the compression quality and bitstream throughput capacity of the encapsulated 3D streaming engine 150.

[0058] Scene graph generation (SGG) can include parsing an image or a series of images in order to generate a structured representation of a visual scene. The scene graph can be encoded with objects and their corresponding relationships, as well as context data that surrounds and binds those visual relationships. In this way, machine learning predictions of objects and relationships are predicted and modeled based on their surrounding context.

[0059] In an embodiment, the encapsulated 3D streaming engine 150 executed on a 305 GPU server with a 3D game application can be configured to identify objects and scenes in a static image using a mathematical model or statistical learning, and then proceed to perform motion recognition, object tracking, action recognition, etc. in a streaming video. An object or feature detector can obtain the shape and its location, and the attributes of the object can be modeled in three-dimensional space to help further enable the detection, recognition, tracking, interaction, and prediction of the object.

[0060] In an instance, the 3D streaming engine 150 can be configured with real-time 3D rendering capabilities, thus enabling the streaming of 3D "hologram" content. In an embodiment, the 3D "hologram" content includes point clouds, multi-views, and streaming scene graphs (object-oriented representations of 3D video content), as well as dynamic animated meshes with texture streams.

[0061] Figure 4 It is a schematic block diagram of an example embodiment of a video processing system 400. In an example, a user 435 requests 130 a 3D streaming session 145 that uses a library file 145 (such as a Javascript library 455), for example, via a first client device 135-1. Session management (i.e., the session handler 125) can manage the connection to a streaming service 405 supported by a server such as a first server 105-1. The Javascript library 455 or another library can be used to process input controls such as mouse, touch, and key controls. The Javascript library 455 can define specifications specific to the current session. The publicDomainName parameter 418 can be used to connect to a WebSocket server 437 to communicate with the streaming service 405 and an associated encapsulated 3D streaming engine 150 (such as an Unreal application 450). Such a connection can be protected by encryption such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. The WebRTC API 115 can be configured as a proxy 417 to route incoming messages to the WebSocket server 437 and provide the above encryption. A traversal relay server (turnserver) 406 can mediate between the proxy 417 and the streaming service 405. Additional mouse, touch, and keyboard controls can be processed between the proxy 417 and the streaming service 405. The WebSocket server 437 can receive incoming messages from the user 435 at a front-end interface such as a website 436. Such messages can contain an identifier of the listening port of the WebSocket server 437. Incoming and outgoing messages arriving at and leaving the WebSocket server 437 can use separate communication paths. Such communication paths can be manually selected and configured. Multiple sessions of the system 400 can be configured separately. The GPU instance 370 can thus host services separately without the need to configure middleware or additional proxies between the services.

[0062] Example streaming use cases

[0063] Figure 5ASchematic diagram and overview of example use case 500a of an embodiment of a video processing system (e.g., system 100). A sales representative 535-1 can connect with a remotely located client 535-2 via a cloud-based server implementation 505 of system 100 and conduct a virtual-guided sales tour. Advantages of such a connection include enabling simulated face-to-face interactions even when the point of sale is closed or otherwise inaccessible to client 535-2 or even to sales representative 535-1. Such an implementation 505 can host multiple clients on one instance of implementation 505. Video and chat functionality can be provided by implementation 505. Additionally, implementation 505 provides a controlled environment for distributing content related to products offered by sales representative 535-1, resulting in a controlled cost for providing and maintaining such an environment.

[0064] Web-based streaming according to the present disclosure can provide a high-end 3D streaming experience for multiple users on a public website. Thus, a streaming solution can be provided for a point of sale (e.g., a dealership). Access to such a streaming solution can be limited to specific users, thereby increasing the security of the platform. Instead of purchasing and maintaining hardware such as a proprietary streaming server, the point of sale can use the cloud streaming solution described herein, thus saving resources, which can result in increased returns when applied elsewhere, such as by improving the customer experience. Use cases can be provided globally to facilitate high-quality connections between users and GPU cloud instances.

[0065] In an example embodiment, a digital sales lounge is provided for an automotive dealership. The dealership can act as a host by starting a streaming session and inviting multiple clients (e.g., up to 6 clients or more) to a separate streaming session. All participants in the streaming session can watch a common video presentation while effectively being unrestricted by their physical locations.

[0066] Figure 5B is according to Figure 5A Illustration 500b of various example frames of a video presentation configured according to use case 500a. Illustration 500b shows starting 578 a streaming session by sharing an access code with client 535-2, client 535-2 accessing 583 the streaming session using the access code, starting a joint configuration application 588 such that sales representative 535-1 and client 535-2 can work together to configure various options in a car to be sold, and the streaming session ending 593 with a customized video or an image rendered for an individual manual according to the options selected via joint configuration application 588.

[0067] Challenges faced by existing real-time streaming applications

[0068] Other real-time streaming applications that exist at the time of this disclosure require deep technical expertise to develop streaming solutions or experiences. Traditionally, setting up such a streaming environment has been a very manual process. Using existing methods, it has generally been difficult to create reliable and scalable solutions globally. Additionally, existing methods typically cannot provide a high enough visual quality with a low enough latency. For existing methods using Pixelstreaming, users are locked into using the Unreal Engine as the encapsulated 3D streaming engine. Additionally, Pixelstreaming provides users with a limited feature set. Existing methods require multiple service providers to set up a streaming business. As a result, streaming via existing methods is expensive, and calculating the actual cost can be a complex process. Existing cloud-based streaming solutions suffer from a complex IT landscape within companies that adopt such solutions, making it difficult to set up such solutions. The above challenges result in an overly long time to market for streaming real-time solutions.

[0069] Aspects of the real-time streaming solutions disclosed herein

[0070] The embodiments disclosed herein provide fully developed streaming technologies that are available from an online marketplace such as the Google Play Store. Such streaming solutions can be set up via a simple and highly automated process without the development expertise required of users. Such streaming solutions can thus be established to cater to content creators rather than software developers. Streaming can be effectively separated from the associated application engine, thus not causing a technical lock-in with respect to the specific application engine to be used. As a result, a complete streaming solution can be provided from a one-stop shop.

[0071] Additionally, the streaming solutions according to this disclosure support a large variety of advanced technologies and provide a high-quality, high-performance streaming experience. Embodiments can be made available worldwide without limitation. The setup can be flexible to meet the diversity of user needs. Embodiments use a configuration such as software as a service (SaaS) as a blueprint and are thus compatible with most IT guidelines within a company. Such streaming solutions also come with an affordable and easy-to-understand pricing structure.

[0072] Digital processing environment

[0073] An example embodiment of a multimedia system 600 for streaming selected media content to a client device 15 (e.g., client devices 135-1, 135-2, 135-i) of a user can be implemented in a software, firmware, or hardware environment. Figure 6Illustrates such an environment. One or more client devices 15 (e.g., mobile phones) and a cloud 16 (or a server computer or a cluster thereof) provide processing, storage, and input / output devices for executing applications and the like. The client device may be interchangeably referred to as a client computer herein.

[0074] The client device 15 is linked to other computing devices via a communication network 17, including other client devices / programs 15 and one or more server computers 16. The communication network 17 may be a remote access network, a global network (e.g., the Internet), an out-of-band network, a worldwide collection of computers, a local area network or a wide area network, a cloud network, and a part of gateways that communicate with each other using corresponding protocols (TCP / IP, HTTP, Bluetooth, etc.). Other electronic device / computer network architectures are suitable.

[0075] The server computer 16 (e.g., servers 105-1, 105-n) may be configured to implement a streaming media server for providing, formatting, and storing selected media content (e.g., audio, video, text, and images / pictures), and the selected media content is processed and played at the client device 15. The server computer 16 is communicatively coupled to the client device 15, and the client device implements a corresponding video encoder for capturing, encoding, loading, or otherwise providing the selected media content transmitted to the server computer 16. In one exemplary embodiment, one or more server computers 16 are scalable Java application servers such that if there is a traffic spike, the server can handle the increased load.

[0076] Figure 7 is Figure 6 a diagram of the internal structure of a computer / computing node (e.g., a client processor / device / mobile phone device / tablet 15, 135-i, 135-1, 135-2 or a server computer server 16, 105-1, 105-n) in a processing environment, which can be used to facilitate the display of such audio, video, image, or data signal information. Each computer 15, 135-i, 135-1, 135-2, 16, 105-1, 105-n includes a system bus 11, where the bus is a set of physical or virtual hardware lines for data transfer among components of a computer or a processing system. The bus 11 is essentially a shared channel that connects different elements of a computer system (e.g., a processor, a disk storage device, a memory, an input / output port, etc.), enabling data transfer between the elements. An I / O device interface 82 is attached to the system bus 11 for connecting various input and output devices (e.g., a keyboard, a mouse, a touch screen interface, a display, a printer, a speaker, etc.) to the computers 15, 16. A network interface 86 allows the computer to connect to a network attached (e.g., Figure 6each other device of the network shown at 17). The memory 24 provides volatile storage for the computer software instructions 25 and data 26 for implementing the software implementation of the present invention (e.g., capture / load, provide, format, retrieve, download, and / or store selected media content streams and user-initiated command streams).

[0077] The disk storage device 95 provides non-volatile storage for the computer software instructions 92 (equivalent to "OS program") and data 94 for implementing an embodiment of the multimedia system 600 of the present invention. The central processing unit (CPU) 84 is also attached to the system bus 11 and provides the execution of computer instructions. The processor 84 may include one or more microprocessors, microcontrollers, digital signal processors, field programmable gate arrays (FPGAs), programmable logic devices, state machines, gated logic, discrete hardware circuits to handle the load balancing of 3D streaming.

[0078] In an embodiment, the CPU 84 is a hybrid CPU / graphics processing unit (GPU) with an embedded vector packet processor (VPP)-based hardware accelerator. The hybrid CPU / GPU with an embedded VPP is optimized to improve packet processing for multiple 3D streaming sessions, including complete H.264 decoding of multiple channels. In an example embodiment, the hybrid CPU / graphics processing unit (GPU) with an embedded vector packet processor (VPP)-based hardware accelerator is configured to handle the ultra-large scale cloud workload of 3D streaming sessions, including 5G transmission processing and 5G radio access network intelligent controller (RIC) and edge inference. In a preferred example, the hybrid CPU / graphics processing unit (GPU) with an embedded vector packet processor (VPP)-based hardware accelerator includes an integrated 1 terabit switch, as well as true on-line streaming and highly programmable 3D packet processing. The VPP accelerator can provide low latency and high throughput capacity, making it suitable for deploying multiple high-speed 3D streaming sessions.

[0079] In an embodiment, a hybrid CPU / graphics processing unit (GPU) includes an embedded hardware / firmware implementation of a 3D streaming engine 150 that can generate multiple executable instances of the encapsulated 3D streaming engine 150. Each encapsulated execution instance of the 3D streaming engine 150 is preferably in a container. In this way, the generated instances of the 3D streaming engine 150, along with the corresponding libraries and dependencies, execute in lightweight executable containers that are capable of always running and can be optimized for faster and more secure deployment on a portable computing unit. Different from traditional computing methods, by executing the 3D streaming engine 150 in a container, the 3D streaming engine is more resilient to errors and inaccuracies when transferred to a new location. Containerization eliminates this problem by bundling the application code with the relevant configuration files, libraries, and dependencies required for its operation. This container can then potentially be extracted from the host system and become portable, capable of running across any platform or cloud without problems. In an embodiment, the encapsulated 3D streaming engine 150 can be implemented as a virtual machine executed from servers 16, 105-1, 105-n or as a virtual machine executed via a secure socket layer deployed at client devices 135-i, 135-1, 135-2.

[0080] In one embodiment, the processor routines 92 and data 94 of a multimedia processing system (video processing system) can be implemented as a computer program product, including a computer-readable medium capable of being stored on a storage device 95 or deployed as software as a service (SaaS), which provides at least a portion of the software instructions for the multimedia processing system 600, the encapsulated 3D streaming engine 150, the GPU server with a 3D game application 305, the context verification 140, and the session handler 125. An example of a software implementation of the multimedia processing system 600 can be implemented as a computer program product 92 and can be installed through any suitable software installation process known in the art. In another embodiment, at least a portion of the instructions of the multimedia processing system 600 can also be downloaded via a cable, communication, and / or wireless connection. In other embodiments, the software components of the multimedia processing system 600 can be implemented as a computer program propagated signal product 77 embodied in a propagated signal on a carrier medium (e.g., radio waves, infrared waves, laser waves, sound waves, or electrical waves propagated through a global network such as the Internet or one or more other networks). Such carrier media or signals provide at least a portion of the software instructions for the routines / programs 92 of the multimedia processing system 600.

[0081] In an alternative embodiment, the propagated signal is an analog carrier or digital signal carried on a propagation medium. For example, the propagated signal can be a digitized signal propagated through a global network (e.g., the Internet), a telecommunications network, an out-of-band network, or other networks. In one embodiment, the propagated signal is transmitted over a propagation medium over a period of time, such as instructions for a software application sent in packet form over a network over milliseconds, seconds, minutes, or longer. In another embodiment, the computer-readable medium of the computer program product 92 is a propagation medium that can be received and read by the computer system 15, for example, by receiving the propagation medium and identifying the propagated signal embodied in the propagation medium, as described above for the computer program propagated signal product.

[0082] The multimedia processing system 600 described herein can be configured using any known programming language, including any high-level object-oriented programming language. The client computer / device 15 of the multimedia system 600 can be implemented via a software embodiment and can operate within a browser session. The multimedia processing system 600 can be developed using HTML, JavaScript, Flash, etc. The HTML code can be configured to embed the system into a web browsing session at the client 15. JavaScript can be configured to perform clickstream and session tracking at the client 15 and store streaming media recording and editing data in a cache. In another embodiment, the system can be implemented in HTML5 for client devices 15 that do not have Flash installed and use the HTTP Live Streaming (HLS) or MPEG-DASH protocol. The system can be implemented to transmit media streams using real-time streaming protocols, such as: Real-Time Messaging Protocol (RTMP), Real-Time Streaming Protocol (RTSP), Web Real-Time Communication (WebRTC), etc. The components of the multimedia processing system 600 can be configured to create and load XML, JSON, or CSV data files or other structured metadata files (e.g., manifest files) that have information about where and how the components of the multimedia processing system 600 are stored, hosted, or formatted, such as timing information, size, footnotes, attachments, interactive components, style sheets, etc.

[0083] In an example mobile implementation, the user interface framework for the components of the multimedia processing system 600 can be based on XHP, Javelin, and WURFL. In another example mobile implementation of the OS X and iOS operating systems and their corresponding APIs, Cocoa and Cocoa Touch can be implemented using Objective-C or any other high-level programming language that adds Smalltalk-style messaging to the C programming language.

[0084] Example Advantages

[0085] The benefits provided to users through the streaming solutions disclosed herein include the following. Embodiments occur with low streaming costs for users (e.g., 70% cheaper than some current products). Additionally, no additional development costs are incurred on the client side. The implementation provides a very fast and automated process, resulting in a quick turnaround time for projects created by such implementations. Further, due to the SaaS method used, this streaming solution is not subject to limitations due to IT regulations. Embodiments provide high-performance streaming with high visual quality and low latency. Such solutions avoid technology lock-in and are thus open to future developments in the real-time engine and streaming fields. Cost control can be provided by automatically adjusting virtual computing instances based on actual demand. For example, thresholds can be employed to ensure that a minimum number of virtual computing instances are always available. Embodiments provide full transparency through a separate dashboard for each client. Such dashboards can display information including multiple runtime instances that are configured or available. The complete solution can be obtained as a service from a single provider.

[0086] This high-fidelity streaming solution can be configured to support ray tracing. 3D applications can be streamed according to the disclosed methods and systems without additional software plugins. External interfaces can be used to control the streaming application. The implementation can be agnostic to cloud providers and devices working with many such providers and user device types. Embodiments can include built-in maintenance and analysis tools.

[0087] Any game engine or streaming engine can be supported. Modern web browsers such as Chrome, Safari, Firefox, etc. can be used in various implementations. Multiple applications can be hosted by a single graphics card or GPU. External protocols such as HTTP and WebSocket can be supported. The implementation allows for analytics reports such as session time, concurrent users (CCU), and other analytics. The implementation can include an auto-shutdown feature. Cross-region support can be achieved by such streaming solutions. Website integration can be accomplished, for example, via the included Javascript library. Existing streaming applications can be easily migrated to a platform configured according to the currently disclosed methods. Security can be provided globally where source files and streaming applications are protected on the cloud server. Automatic adjustment of virtual computing instances based on multiple connected users can, for example, achieve connections to the service on one thousand users. All general application features can be supported, including animations, products in motion, etc.

[0088] Although example embodiments have been specifically shown and described, those skilled in the art will understand that various changes in form and detail can be made thereto without departing from the scope of the embodiments covered by the appended claims.

Claims

1. A video processing system, the video processing system comprising: A plurality of servers, the plurality of servers operating in a peer-to-peer communication network via a WebRTC application programming interface (API) to generate 3D stream data; The plurality of servers respond to a request for a 3D stream session received at a first client device by a session handler in the following manner: (i) Receive, from the session handler communicating with the first client device, context verification regarding the first client device, the context verification including the Internet Protocol (IP) location of the first client device; (ii) Respond to the context verification regarding the first client device by assigning a first server among the plurality to initiate a 3D stream session at the first client device based on the Internet Protocol (IP) location of the first client device; (iii) Guide a first instance of an encapsulated 3D stream engine at the first server to initiate the 3D stream session at the first client device via the WebRTC application programming interface (API); (iv) Inject a library file of the WebRTC application programming interface (API) into an executable file of the encapsulated 3D stream engine to generate a digital representation of an interactive 3D environment for the 3D stream session at the first client device; (v) Render the digital representation of the interactive 3D environment within a buffer of a graphics processing unit (GPU); (vi) An encoder encodes the rendered representation for transmitting the 3D stream session to the first client device; (vii) Stream the encoded representation of the interactive 3D environment at the first client device; (viii) Dynamically adjust the 3D stream session of the interactive 3D environment in response to an input received from the first client device; And (ix) Maintain the digital representation of the interactive 3D environment of the 3D stream session by iteratively processing (v) to (viii) until the 3D stream session terminates at the first client device.

2. The video processing system according to claim 1, wherein the first server is configured to establish a predetermined number of virtual computing instances of the 3D stream session with one or more client devices; and The first server is configured to automatically adjust the number of provided virtual computing instances based on a change in the number of the client devices, wherein at least one of the one or more client devices is the first client device.

3. The video processing system according to claim 1 or 2, wherein at least one of the plurality of servers is configured to respond to a request for a 3D stream session from a client device by assigning a server among the plurality of servers to respond to the corresponding request for the 3D stream session, wherein the first server is one of the plurality of servers.

4. The video processing system according to any one of claims 1 to 3, wherein at least one of the plurality of servers performs computational analysis to determine an appropriate number of 3D streaming sessions of the plurality of client devices adjusted based on the predetermined number of virtual computing instances.

5. The video processing system according to any one of claims 1 to 4, wherein the computational analysis includes concurrent users (CCU), session time, daily active users (DAU), monthly active users (MAU), and sessions.

6. The video processing system according to any one of claims 1 to 5, wherein the encapsulated 3D streaming engine configures the 3D streaming session as a media stream cluster such that at least one of the plurality of servers performs multicast transmission of data packets including at least a part of the digital representation of the interactive 3D environment to the one or more client devices.

7. The video processing system according to any one of claims 1 to 6, wherein dynamically adjusting the 3D streaming session of the interactive 3D environment in response to an input received from the first client device further includes: The encapsulated 3D streaming engine is configured to capture events associated with the interactive 3D environment, which are detected by one or more listeners at one or more of the client devices; The encapsulated 3D streaming engine is configured to capture metadata, consumption information, and interaction data at one or more of the client devices during the 3D streaming session; The encapsulated 3D streaming engine is configured to respond to the captured events, metadata, consumption information, and interaction data by redrawing the interactive 3D environment at one or more of the client devices to maintain the digital representation of the interactive 3D environment.

8. The video processing system according to any one of claims 1 to 7, wherein the encapsulated 3D streaming engine maintains the digital representation of the interactive 3D environment by: Calculating composite 3D media data based on the captured events, metadata, consumption information, and interaction data; Defining a composite image layout based on attributes derived from the 3D streaming session at one or more of the client devices Configuring the 3D streaming session to provide a composite media signal according to the defined composite image layout; and Transmitting the composite media signal in the form of data packets to the one or more client devices via a packetizer.

9. The video processing system according to any one of claims 1 to 8, the video processing system further includes the encapsulated 3D streaming engine interfacing with the one or more client devices to facilitate a control communication interface, thereby enabling voice, text, and video transmission among the one or more client devices in a data packet exchange communication system of a peer-to-peer network.

10. The video processing system according to any one of claims 1 to 9, wherein the encapsulated 3D streaming engine is configured as a software container that includes bindings of software components with configuration files, libraries, and dependencies required for execution.

11. The video processing system according to any one of claims 1 to 10, wherein the software component includes at least one of the following: the library file of the WebRTC application programming interface (API), an event manager, a scene manager, a resource manager, a session manager, a physical manager, and an artificial intelligence system; and wherein the WebRTC application programming interface (API) includes one or more WebRTC application programming interface (API) functions configured to call one or more of the software components.

12. The video processing system according to any one of claims 1 to 11, wherein the encapsulated 3D streaming engine is not a software plug-in.

13. The video processing system according to any one of claims 1 to 12, wherein the encapsulated 3D streaming engine is deployed as a virtual machine via Secure Sockets Layer (SSL), and the virtual machine is configured in the software container.

14. The video processing system according to any one of claims 1 to 13, wherein the input is received through a data channel established by using the RTCDataChannel application programming interface (API).

15. The video processing system according to any one of claims 1 to 14, wherein the encapsulated 3D streaming engine is configured to train an interactive frame prediction model for the digital representation of the interactive 3D environment based on the encoded stream data of the 3D stream session and the events, the metadata, the interaction data, and the consumption data from the one or more client devices.

16. The video processing system according to any one of claims 1 to 15, wherein the encoder is configured to encode the 3D stream session by selectively using prediction frames generated by a training-based interactive frame prediction model, and transmit the trained interactive frame prediction model and the encoded stream data to the first client device and the second client device to create the digital representation of the interactive 3D environment; and the first client device and the second client device are configured to receive the trained interactive frame prediction model and the encoded stream data, and decode the encoded stream data based on the trained interactive frame prediction model to create the digital representation of the interactive 3D environment.

17. A video processing method, the video processing method comprising: configuring a server computer system to respond to a request for a 3D stream session received at a first client device by a session handler by: (i) receiving, from the session handler communicating with the first client device, context verification regarding the first client device, the context verification including the Internet Protocol (IP) location of the first client device; (ii) guiding the 3D stream session at the first client device to respond to the context verification regarding the first client device by assigning a first server among a plurality of servers of the server computer system based on the Internet Protocol (IP) location of the first client device; (iii) Cause a first instance of the encapsulated 3D streaming engine at the first server to initiate the 3D streaming session at the first client device via the WebRTC application programming interface (API); (iv) Inject a library file of the WebRTC application programming interface (API) into the executable file of the encapsulated 3D streaming engine to generate a digital representation of an interactive 3D environment for the 3D streaming session at the first client device; (v) Render the digital representation of the interactive 3D environment within a buffer of a graphics processing unit (GPU); (vi) Encode the rendered representation via an encoder for transmitting the 3D streaming session to the first client device; (vii) Stream the encoded representation of the interactive 3D environment at the first client device; (viii) Dynamically adjust the 3D streaming session of the interactive 3D environment in response to an input received from the first client device; and (ix) Maintain the digital representation of the interactive 3D environment of the 3D streaming session by iteratively processing (v) to (viii) until the 3D streaming session terminates at the first client device.

18. The video processing method according to claim 17, the video processing method further comprising configuring the first server to establish a predetermined number of virtual computing instances of the 3D streaming session with one or more client devices; and configuring the first server to automatically adjust the provided number of virtual computing instances based on a change in the number of the client devices, wherein at least one of the one or more client devices is the first client device.

19. The video processing method according to claim 17 or 18, the video processing method further comprising configuring at least one of the plurality of servers to respond to a request for a 3D streaming session from the client device by assigning a server among the plurality of servers to respond to the corresponding request for the 3D streaming session, wherein the first server is one of the plurality of servers.

20. The video processing method according to any one of claims 17 to 19, the video processing method further comprising configuring at least one of the plurality of servers to perform a computational analysis to determine an appropriate number of 3D streaming sessions of the plurality of client devices adjusted based on the predetermined number of virtual computing instances.

21. The video processing method according to any one of claims 17 to 20, the video processing method further comprising configuring the computational analysis to include concurrent users (CCU), session time, daily active users (DAU), monthly active users (MAU), and sessions.

22. The video processing method according to any one of claims 17 to 21, the video processing method further comprising configuring the 3D stream session as a media stream cluster such that at least one of the plurality of servers performs multicast transmission of data packets containing at least a portion of the digital representation of the interactive 3D environment to the one or more client devices.

23. The video processing method according to any one of claims 17 to 22, wherein dynamically adjusting the 3D stream session of the interactive 3D environment in response to an input received from the first client device further comprises: Configuring the encapsulated 3D stream engine to capture events associated with the interactive 3D environment, the events being detected by one or more listeners at one or more of the client devices; Configuring the encapsulated 3D stream engine to capture metadata, consumption information, and interaction data during the 3D stream session at one or more of the client devices; Configuring the encapsulated 3D stream engine to respond to the captured events, metadata, consumption information, and interaction data by redrawing the interactive 3D environment at one or more of the client devices to maintain the digital representation of the interactive 3D environment.

24. The video processing method according to any one of claims 17 to 23, wherein the encapsulated 3D stream engine maintains the digital representation of the interactive 3D environment by: Calculating composite 3D media data based on the captured events, metadata, consumption information, and interaction data; Defining a composite image layout based on attributes derived from the 3D stream session at one or more of the client devices Configuring the 3D stream session to provide a composite media signal according to the defined composite image layout; and Transmitting the composite media signal in the form of data packets to the one or more client devices via a packetizer.

25. The video processing method according to any one of claims 17 to 24, the video processing method further comprising causing the encapsulated 3D stream engine to interface with the one or more client devices to facilitate a control communication interface, thereby enabling voice, text, and video transmission among the one or more client devices in a data packet exchange communication system of a peer-to-peer network.

26. The video processing method according to any one of claims 17 to 25, the video processing method further comprising configuring the encapsulated 3D stream engine as a bound container containing software components.

27. The video processing method according to any one of claims 17 to 26, wherein the software components include at least one of the following: library files of the WebRTC application programming interface (API), an event manager, a scene manager, a resource manager, a session manager, a physics manager, and an artificial intelligence system; and wherein the WebRTC application programming interface (API) includes one or more WebRTC application programming interface (API) functions configured to call one or more of the software components.

28. The video processing method according to any one of claims 17 to 27, wherein the encapsulated 3D streaming engine is not a software plug-in.

29. The video processing method according to any one of claims 17 to 28, the video processing method further comprising deploying the encapsulated 3D streaming engine as a virtual machine via a Secure Sockets Layer (SSL), the virtual machine being configured in the software container.

30. The video processing method according to any one of claims 17 to 29, the video processing method further comprising receiving the input using a data channel established using an RTCDataChannel application programming interface (API).

31. The video processing method according to any one of claims 17 to 30, the video processing method further comprising configuring the encapsulated 3D streaming engine to train an interactive frame prediction model for the digital representation of the interactive 3D environment based on the encoded stream data of the 3D stream session and the events, the metadata, the interaction data, and the consumption data from the one or more client devices.

32. The video processing method according to any one of claims 17 to 31, the video processing method further comprising configuring the encoder to encode the 3D stream session by selectively using prediction frames generated by a training-based interactive frame prediction model, and transmitting the trained interactive frame prediction model and the encoded stream data to the first client device and the second client device to create the digital representation of the interactive 3D environment; and configuring the first client device and the second client device to receive the trained interactive frame prediction model and the encoded stream data, and decoding the encoded stream data based on the trained interactive frame prediction model to create the digital representation of the interactive 3D environment.

33. The system or method according to any one of the preceding claims, wherein one or more of the plurality of servers comprise a hybrid CPU / graphics processing unit (GPU) having an embedded encapsulated 3D streaming engine, the embedded encapsulated 3D streaming engine having a hardware accelerator based on a vector packet processor (VPP), the hybrid CPU / graphics processing unit being optimized to improve packet processing for a plurality of the 3D stream sessions, including full H.264 decoding of multiple channels.

Citation Information

Patent Citations

  • Conversational framework

    CN110321413A

  • Building intelligent control three-dimensional model display method and related equipment

    CN112783064A

  • Unified, browser-based enterprise collaboration platform

    EP3563248A1

  • System architecture for cloud gaming

    EP3862060A1