Video processing system and video processing method
By combining GPU hosting and WebRTC API with a packaged 3D streaming engine and artificial intelligence system, the problem of high-quality, low-latency 3D streaming video transmission in existing technologies has been solved, achieving a low-cost, globally available, and efficient streaming solution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MONKEYWAY GMBH
- Filing Date
- 2023-09-29
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies struggle to provide high-quality, low-latency interactive 3D streaming video, and traditional methods require deep technical expertise and are costly, making it difficult to achieve reliable and scalable streaming solutions globally.
By configuring graphics processing unit (GPU) instances to host streaming services, 3D streaming data can be generated in peer-to-peer networks using the WebRTC API, dynamically adjusting the interactive 3D environment, and combining an encapsulated 3D streaming engine and artificial intelligence system to achieve efficient video processing and streaming.
It enables interactive 3D streaming video transmission with high visual quality and low latency, lowers the technical threshold and cost, and supports reliable and scalable streaming solutions worldwide.
Smart Images

Figure CN120283395B_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 377,986, filed September 30, 2022. The entire teachings of the above application are incorporated herein by reference. Background Technology
[0003] Streaming video applications have recently made strides in a new dimension, attempting to provide users with a more immersive experience. As virtual alternatives to face-to-face interaction continue to evolve, they offer users the opportunity for secure and reliable connections across vast geographical distances. Summary of the Invention
[0004] Embodiments of this disclosure provide techniques for streaming interactive video and audio content of interest to users. Streaming technologies provide users with a high-quality, low-latency video and audio streaming experience. Graphics processing unit (GPU) instances are configured to directly host services supporting streaming content, thereby facilitating the implementation of the methods without middleware or plugins. Embodiments of the invention relate to systems and methods for video processing and for client devices for streaming processed media content to these users.
[0005] In some embodiments, the video processing system includes multiple servers that operate in a peer-to-peer network via a WebRTC application programming interface (API) to generate 3D streaming data. The multiple servers may, in part, receive a request for a 3D streaming session at the first client device from a session handler communicating with the first client device, whereby the context verification includes the Internet Protocol (IP) location of the first client device. The multiple servers may direct the 3D streaming session at the first client device to respond to the context verification based on the IP location of the first client device by assigning a first server among multiple servers.
[0006] Continuing, multiple servers can bootstrap a first instance of the encapsulated 3D streaming engine from the first server to initiate a 3D streaming session on the first client device via the WebRTC API. Multiple servers can inject the WebRTC API library files into the executable of the encapsulated 3D streaming engine to generate a digital representation of an interactive 3D environment for the 3D streaming session on the first client device.
[0007] Multiple servers can continue to render a digital representation of the interactive 3D environment within the GPU's buffer, encode the rendered representation for use in transmitting a 3D streaming session to a first client device, and stream the encoded representation of the interactive 3D environment at the first client device. The multiple servers can dynamically adjust the 3D streaming session of the interactive 3D environment in response to input received from the first client device.
[0008] Multiple servers can maintain a digital representation of the interactive 3D environment of a 3D streaming session by iteratively performing the processes involved in rendering, encoding, streaming, and adjustment described above, until the 3D streaming session terminates at the first client device.
[0009] The first server can be configured to establish a predetermined number of virtual computing instances (VCIs) for 3D streaming sessions with one or more client devices, and the first server can be configured to automatically adjust the number of VCIs provided based on changes in the number of client devices. At least one of the one or more client devices can be the first client device.
[0010] At least one of the multiple servers can be configured to respond to requests for a 3D streaming session from a client device by assigning one of the multiple servers to the request. The first server can be one of the multiple servers.
[0011] At least one of the multiple servers can perform computational analysis to determine the appropriate number of 3D streaming sessions for multiple client devices, adjusted based on a predetermined number of VCIs. The computational analysis may include concurrent users (CCU), session time, daily active users (DAU), monthly active users (MAU), and sessions. The encapsulated 3D streaming engine can configure the 3D streaming sessions as a video streaming cluster, enabling at least one of the multiple servers to perform multicast transmission of data packets containing at least a portion of a digital representation of an interactive 3D environment to one or more client devices.
[0012] A 3D streaming session that dynamically adjusts the interactive 3D environment in response to input received from a first client device may further include an encapsulated 3D streaming engine configured to capture events associated with the interactive 3D environment. These events can be detected by one or more listeners at one or more locations on the client devices. The dynamically adjusted 3D streaming engine is further configured to capture metadata, consumption information, and interaction data at one or more locations on the client devices during the 3D streaming session. The dynamically adjusted 3D streaming engine is further configured to respond to the captured events, metadata, consumption information, and interaction data by redrawing the interactive 3D environment at one or more client devices to maintain a digital representation of the interactive 3D environment.
[0013] The encapsulated 3D streaming engine maintains the digital representation of an interactive 3D environment by: computing composite 3D video data based on captured events, metadata, consumption information, and interaction data; defining a composite image layout based on attributes derived from a 3D streaming session at one or more client devices; configuring the 3D streaming session to provide composite video signals according to the defined composite image layout; and transmitting the composite video signals as data packets to one or more client devices via a packetizer. The encapsulated 3D streaming engine can interface with one or more client devices to facilitate control communication interfaces, thereby enabling voice, text, and video transmission among one or more client devices in a peer-to-peer packet-switched communication system.
[0014] The encapsulated 3D streaming engine can be configured as a container containing bindings of software components. These software components can include at least one of the following: a WebRTC API library, an event manager, a scene manager, a resource manager, a session manager, a physics manager, and an artificial intelligence system. The WebRTC API can contain one or more WebRTC API functions configured to call one or more software components. The encapsulated 3D streaming engine does not need to be a software plugin.
[0015] Input can be received via a data channel established using the RTCDataChannel API. The encapsulated 3D streaming engine can be configured to train an interactive frame prediction model for digital representations of interactive 3D environments based on the encoded streaming data of the 3D streaming session and events, metadata, interaction data, and consumption data from one or more client devices.
[0016] The encapsulated 3D streaming engine can be configured to encode a 3D streaming session by selectively using predicted frames generated based on a trained interactive frame prediction model, and to transmit the trained interactive frame prediction model and the encoded stream data to a first client device and a second client device to create a digital representation of an interactive 3D environment. The first and second client devices can be configured to receive the trained interactive frame prediction model and the encoded stream data, and to decode the encoded stream data based on the trained interactive frame prediction model to create a digital representation of the interactive 3D environment.
[0017] In some embodiments, a video processing method includes configuring a server computer system to respond to a request for a 3D streaming session received at a first client device by a session handler. Responding to the session handler communicating with the first client device may include receiving context verification from the session handler regarding the first client device, the context verification including the IP location of the first client device. Responding to the session handler may include directing the 3D streaming session at the first client device to respond to the context verification regarding the first client device by assigning a first server from a plurality of servers on the server computer system based on the IP location of the first client device.
[0018] Responding to the session handler may include instructing a first instance of the encapsulated 3D streaming engine at the first server to initiate a 3D streaming session at the first client device via the WebRTC API. Responding to the session handler may also include injecting the WebRTC API library files into the executable file of the encapsulated 3D streaming engine to generate a digital representation of an interactive 3D environment for the 3D streaming session at the first client device.
[0019] Responding to the session handler may include rendering a digital representation of the interactive 3D environment within the GPU's buffer, encoding the rendered representation for transmitting the 3D streaming session to a first client device, and streaming the encoded representation of the interactive 3D environment at the first client device. Responding to the session handler may also include a 3D streaming session that dynamically adjusts the interactive 3D environment in response to input received from the first client device.
[0020] Responding to the session handler may involve maintaining a digital representation of the interactive 3D environment of the 3D streaming session by iteratively performing the processes involved in rendering, encoding, streaming, and adjustment described above, until the 3D streaming session terminates at the first client device. In embodiments of the video processing method, the method performs operations to implement any of the embodiments or combinations of embodiments described herein. Attached Figure Description
[0021] The foregoing will become clear from the following more specific description of exemplary embodiments, as illustrated in the accompanying drawings, in which similar reference numerals refer to the same parts in different views. The drawings are not necessarily drawn to scale, but rather the emphasis is on the illustrated embodiments.
[0022] Figure 1 This is a schematic block diagram of an example video processing system according to the present disclosure.
[0023] Figure 2 This is a flowchart of an example video processing method according to this disclosure.
[0024] Figure 3 This is a schematic block diagram of an example embodiment of a video processing system.
[0025] Figure 4 This is a schematic block diagram of an example embodiment of a video processing system.
[0026] Figure 5A This is a schematic diagram and overview of example use cases for an embodiment of a video processing system.
[0027] Figure 5B Depicting Figure 5A Various example user interfaces for example video processing system use cases.
[0028] Figure 6 This is a schematic diagram of an example computer network environment in which embodiments of the present invention are deployed.
[0029] Figure 7 yes Figure 6 A block diagram of computer nodes in a network. Detailed Implementation
[0030] The example embodiment is described below.
[0031] This paper describes a high-performance, low-cost streaming solution. This solution offers high visual quality and low latency, while being easily implemented as a complete solution by users without programming skills.
[0032] System Architecture
[0033] Figure 1 This is a schematic block diagram of an example embodiment of a video processing system 100. In the example embodiment, the video processing system 100 includes a plurality of servers 105-1, 105-n that operate in a peer-to-peer (i.e., P2P) communication network 110 via a WebRTC application programming interface (API) 115 to generate 3D streaming data 120. The plurality of servers 105-1, 105-n may respond in part to a request 130 for a 3D streaming session received at the first client device 135-1 by the session handler 125 communicating with the first client device 135-1, receiving a context verification 140 regarding the first client device 135-1, the context verification including the Internet Protocol (IP) location of the first client device 135-1. Multiple servers 105-1, 105-n can be assigned, based on the IP location of the first client device 135-1, to direct a 3D streaming session 145 at the first client device 135-1 to respond to a context verification 140 regarding the first client device 135-1.
[0034] In an example, context verification 140 may include a distance metric based on the distance between the respective servers 105-1, 105-n and the first client device 135. In an embodiment, context verification 140 regarding the first client device 135 may include the number of network hops—the time taken for intermediate routers or servers to process signals transmitted from the respective servers 105-1, 105-n to the first client device 134. Synthetic monitoring tools can be deployed from the respective servers 105-1, 105-n to run manual calls to the first client 145 and detect any increases in latency or performance degradation. Based on context verification data analysis, the server with the lowest latency and highest throughput capacity delivered to the client can be selected from the servers 105-1, 105-n.
[0035] Continuing, multiple servers 105-1 and 105-n can guide a first instance of the encapsulated 3D streaming engine 150 at the first server 105-1 to initiate a 3D streaming session 145 at the first client device 135-1 via the WebRTC API 115. The multiple servers 105-1 and 105-n can inject the WebRTC API 115 library file 155 into the executable file 160 of the encapsulated 3D streaming engine 145 to generate a digital representation 165 of the interactive 3D environment for the 3D streaming session 145 at the first client device 135-1.
[0036] Multiple servers 105-1, 105-n can continue to render a digital representation 165 of the interactive 3D environment within a buffer of the graphics processing unit (GPU) 170, encode the rendered representation for transmission of a 3D streaming session 145 to a first client device 135-1, and stream the encoded representation of the interactive 3D environment at the first client device 135-1. The multiple servers 105-1, 105-n can dynamically adjust the 3D streaming session 145 of the interactive 3D environment in response to input 175 received from the first client device 135-1.
[0037] Multiple servers 105-1, 105-n can maintain the digital representation 165 of the interactive 3D environment of the 3D streaming session 145 by iteratively performing the processes involved in the above-described rendering, encoding, streaming, and adjustment, until the 3D streaming session 145 terminates at the first client device 135-1.
[0038] Figure 2This is a flowchart of an example embodiment of video processing method 200. In the example embodiment, video processing method 200 includes configuring a first server 105-1 of a plurality of servers 105-1, 105-n of system 100 to respond to a request 130 for a 3D streaming session received at a first client device 135-1 by session handler 125. In method 200, responding to session handler 125 communicating with the first client device 135-1 may include receiving 204 context verification 140 from the session handler regarding the first client device 135-1, the context verification including the IP location of the first client device 135-1. In method 200, responding to session handler 125 may include directing a 3D streaming session 145 at the first client device 135-1 to respond to context verification 140 regarding the first client device 135-1 by assigning the first server 105-1 among a plurality of servers 105-1, 105-n of the server computer system based on the IP location of the first client device 135-1.
[0039] In method 200, responding to session handler 125 may include instructing 214 that a first instance of the encapsulated 3D streaming engine 150 at the first server 105-1 initiates a 3D streaming session 145 at the first client device 135-1 via the WebRTC API 115. In method 200, responding to session handler 125 may include injecting 219 a library file 155 of the WebRTC API into the executable file 160 of the encapsulated 3D streaming engine 150 to generate a digital representation 165 of an interactive 3D environment for the 3D streaming session 145 at the first client device 135-1.
[0040] In method 200, responding to session handler 125 may include rendering a digital representation 165 of the interactive 3D environment 224 within a buffer of GPU 170, encoding 229 the rendered representation for transmitting the 3D streaming session 145 to the first client device 135-1, and streaming 234 the encoded representation of the interactive 3D environment at the first client device 135-1. In method 200, responding to session handler 125 may include dynamically adjusting the 3D streaming session 145 of the interactive 3D environment 239 in response to input 175 received from the first client device 135-1.
[0041] In method 200, responding to session handler 125 may include maintaining the digital representation 165 of the interactive 3D environment of 244-1 3D streaming session 145 by iteratively performing processes involving the aforementioned rendering 224, encoding 229, streaming 234, and adjustment 239, until 244-2 3D streaming session 145 is terminated at the first client device 244. In response to terminating 224-2 3D streaming session 145, method 200 may include terminating 249 streaming the digital representation 165 of the interactive 3D environment.
[0042] Return to Figure 1 According to the description, the first server 135-1 of system 100 can be configured to establish a predetermined number of virtual computing instances (VCIs) 180-1, 180-j with one or more client devices 135-1, 135-2, 135-i 3D streaming sessions 145, and the first server 105-1 can be configured to automatically adjust the number of VCIs 180-1, 180-j provided based on changes in the number of client devices 135-1, 135-2, 135-i. At least one of the one or more client devices 135-1, 135-2, 135-i can be the first client device 135-1.
[0043] exist Figure 1 In the example system 100, a first client device 135-1, a second client device 135-2, and an i-th client device 135-i are shown; however, it should be noted that there may be only the first client device 135-1, only the first client device 135-1 and the second client device 135-2, or another number i client devices 135-1, 135-2, and 135-i. Similarly, it should be noted that although in Figure 1 Example system 100 shows a first virtual computing instance (VCI 1) 180-1, a second virtual computing instance (VCI 2) 180-2, and a j-th virtual computing instance (VCI j). However, there may be only the first virtual computing instance (VCI 1) 180-1, only the first virtual computing instance (VCI 1) 180-1 and the second virtual computing instance (VCI 2) 180-2, or another number of j virtual computing instances (VCI 1, VCI 2, ..., VCI j) 180-1, 180-2, 180-j. Similarly, it should be noted that any number of servers 105-1, 105-n may exist, including one server 105-1 or another number of n servers 105-1, 105-n.
[0044] At least one of the multiple servers 105-1, 105-n can be configured to respond to a corresponding request 130 to the 3D streaming session 145 by assigning one of the multiple servers 105-1, 105-n to respond to a request 130 to the 3D streaming session 145 from the client devices 135-1, 135-i. The first server 105-1 can be one of the multiple servers 105-1, 105-n.
[0045] At least one of the multiple servers 105-1, 105-n can perform computational analysis to determine the appropriate number of 3D streaming sessions 145 for multiple client devices 135-1, 135-i, adjusted based on a predetermined number of VCIs 180-1, 180-j. The computational analysis may include concurrent users (CCU), session time, daily active users (DAU), monthly active users (MAU), and sessions. The encapsulated 3D streaming engine 150 can configure the 3D streaming sessions 130 as a video streaming cluster, such that at least one of the multiple servers 105-1, 105-n performs multicast transmission of at least a portion of data packets containing a digital representation 165 of an interactive 3D environment to one or more client devices 135-1, 135-i.
[0046] A 3D streaming session 145 that dynamically adjusts the interactive 3D environment in response to input 175 received from the first client device 135-1 may further include an encapsulated 3D streaming engine 150 configured to capture events associated with the interactive 3D environment. These events can be detected by one or more listeners at one or more of the client devices 135-1, 135-i. The dynamically adjusted 3D streaming engine 150 is configured to capture metadata, consumption information, and interaction data at one or more of the client devices 135-1, 135-i during the 3D streaming session 145. The dynamically adjusted 3D streaming engine 175 is configured to respond to the captured events, metadata, consumption information, and interaction data by redrawing the interactive 3D environment at one or more client devices 135-1, 135-i, to maintain a digital representation 165 of the interactive 3D environment.
[0047] The encapsulated 3D streaming engine 150 maintains the digital representation 165 of the interactive 3D environment by: calculating composite 3D video data based on captured events, metadata, consumption information, and interaction data; defining a composite image layout based on attributes derived from a 3D streaming session 145 at one or more client devices 135-1, 135-i; configuring the 3D streaming session to provide composite video signals according to the defined composite image layout; and transmitting the composite video signals in packet form to one or more client devices 135-1, 135-i via a packetizer. The encapsulated 3D streaming engine 150 can interface with one or more client devices 135-1, 135-i to facilitate control communication interfaces, thereby enabling voice, text, and video transmission among one or more client devices 135-1, 135-i in the packet-switched communication system of the peer-to-peer network 110.
[0048] The encapsulated 3D streaming engine 150 can be configured as a container containing bound software components. These software components can include at least one of the following: a library file 155 for the WebRTC API 115, an event manager, a scene manager, a resource manager, a session manager, a physics manager, and an artificial intelligence system. The WebRTC API 115 can contain one or more WebRTC API functions configured to call one or more software components. The encapsulated 3D streaming engine 150 does not need to be a software plugin.
[0049] The encapsulated 3D streaming engine 150 can help improve the scalability of HTTP streaming and artificial intelligence systems. Artificial intelligence systems can incorporate HTTP-based dynamic adaptive streaming (MPEG-DASH) and deep learning video streaming architectures, such as deep neural networks, deep belief networks, deep reinforcement learning, recurrent neural networks, and convolutional neural networks. The encapsulated 3D streaming engine 150, communicating with the artificial intelligence system, can help configure client devices 135-1 and 135-I as self-learning HTTP adaptive streaming clients, such as those using State-Action-Reward-State-Action (SARSA) or Q-learning (model-free reinforcement learning algorithms).
[0050] Input 175 can be received via a data channel established using the RTCDataChannel API. The encapsulated 3D streaming engine 150 can be configured to train an interactive frame prediction model for a digital representation 165 for an interactive 3D environment based on the encoded streaming data of the 3D streaming session 145 and events, metadata, interaction data, and consumption data from one or more client devices 135-1, 135-i.
[0051] The example implementation enables streaming via a packaged 3D streaming engine 150 as a service for high-end 3D applications utilizing ultra-low latency streaming technology. Any cloud provider can provide servers 105-1, 105-n for the systems and methods described herein, such as Amazon Web Services (AWS), Google Cloud Platform (GCP), Azure, and OnPremise. Additionally, the packaged 3D streaming engine 150 can embed one or more game engines, such as Unreal Engine, Unity, etc. An encoder can include, for example, NVIDIA Nvenc, for encoding the 3D streaming session 145. Such an encoder can use, for example, a codec (e.g., an h264 codec) to encode the 3D streaming session 145. API 115 can be configured as a WebRTC API or another API for ultra-low latency video connections between client devices 135-1, 135-2, 135-i and GPU instances 170 provided by cloud-based servers 105-1, 105-n.
[0052] Figure 3 This is a schematic block diagram of an example embodiment of a video processing system; in particular, its server 305 provides a GPU 370 to support a 3D streaming session 145. In this example, a wrapped 3D streaming engine 150 is provided by a DirectX-based 3D game application 350. A library file 155 for the WebRTC API 115 can be injected via a DirectX injector 319. The video data described in 319 can be rendered in a buffer on the GPU 370 after receiving video data containing images from the 3D game application 350. The rendered video data can be passed to the streaming server 315 (which may be WebRTC-based) and can be configured to encode the video data according to the NVIDIA nvenc codec. The WebRTC server 315 can be configured to control an interface 310 to interact with the 3D game application 350. The encoded video data can be streamed over the Internet 320 by the WebRTC server 315.
[0053] The encapsulated 3D streaming engine 150 can contain multiple execution instances, which can be distributed across multiple servers 16, 105-1, and 105-n. The encapsulated 3D streaming engine 150 can be configured to encode a 3D streaming session 145 by selectively using predicted frames generated based on a trained interactive frame prediction model, and to transmit the trained interactive frame prediction model and the encoded streaming data to a first client device and a second client device to create a digital representation 165 of an interactive 3D environment. The first client device 135-1 and the second client device 135-2 can be configured to receive the trained interactive frame prediction model and the encoded streaming data, and to decode the encoded streaming data based on the trained interactive frame prediction model to create a digital representation 165 of the interactive 3D environment.
[0054] In an example embodiment, the encapsulated 3D streaming engine 150 can be configured to train an interactive frame prediction model. The interactive frame prediction model can be configured and trained based on data from the streaming session, a 3D model, user input, metadata associated with the client session, and consumed data. The encapsulated 3D streaming engine 150 interfaces with an encoder and a decoder to predict frames of the 3D stream using the trained interactive frame prediction model. For example, the encoder can be configured to encode 3D streaming data and transmit the trained interactive frame prediction model. The decoder can receive 3D streaming data and transmit the trained interactive frame prediction model, decode the encoded streaming data based on the trained interactive frame prediction model, and provide the reconstructed 3D streaming data at one or more clients 135-i, 135-1, 135-2.
[0055] In an example embodiment, the encapsulated 3D streaming engine 150 may be configured with interactive frame prediction based on a deep neural network (DNN) for video decoding. For example, interactive frame prediction can be trained and configured using a DNN to help improve video encoding and decoding efficiency. The DNN can be implemented at both the encoder and decoder. The DNN can use previously decoded frames and a trained interactive frame prediction model to predict the current frame. This can be configured as a separate interactive frame prediction mode, which can be optimized to compete with other prediction modes. The DNN can be implemented to avoid transmitting motion vectors. In this way, the DNN can be trained to perform unidirectional and bidirectional prediction, which can provide significant coding efficiency gains relative to high-efficiency video coding (H.265 and MPEG-H Part 2). (DNN-based frame prediction.)
[0056] In an embodiment, the encapsulated 3D streaming engine 150 can be configured to generate and process scene graphs for transmitting 3D data packet streams. The scene graph can be generated using machine learning and machine vision computational prediction systems to provide robust Scene Graph Generation (SGG). For example, the encapsulated 3D streaming engine 150 can perform Scene Graph Generation (SGG), which can be enhanced with a strong semantic representation defining the scene. Scene Graph Generation (SGG) involves automatically mapping images or videos to a semantically structured scene graph. This process can include encoding location data about detected objects and their corresponding relationships, and this process can be tuned using deep learning techniques to improve system performance.
[0057] Scene graph generation (SGG) models can be processed by the encapsulated 3D streaming engine 150 to render 3D video from a visually based scene graph. The scene graph can be a structured representation that captures detailed semantics by explicitly modeling objects. The semantic structure of the scene graph can be processed using perceptual statistics to compute mappings indicating which regions of the video frames are important to the human eye. This process can be incorporated into motion estimation and motion vectors used for inter-frame prediction, resulting in the encoded bitstream. Motion vector quality metrics can be used to construct vector maps, which can then be used to encode the bitstream using metrics such as block variance and block lumen. Applying vector maps and scene graph generation (SGG) models in 3D video coding within a model-based compression framework (e.g., Continuous Block Tracking (CBT)) can improve the compression quality and bitstream throughput of the encapsulated 3D streaming engine 150.
[0058] Scene graph generation (SGG) can involve parsing an image or a series of images to generate a structured representation of a visual scene. Scene graphs can be encoded using objects and their corresponding relationships, as well as contextual data surrounding and binding those visual relationships. In this way, machine learning predictions of objects and relationships are predicted and modeled based on their surrounding context.
[0059] In one embodiment, a packaged 3D streaming engine 150, running on a 305 GPU server with a 3D game application, can be configured to use mathematical models or statistical learning to identify objects and scenes in static images, and then to perform motion recognition, object tracking, action recognition, etc., in streaming video. Object or feature detectors can obtain shapes and their positions, and the attributes of objects can be modeled in three-dimensional space to facilitate further detection, recognition, tracking, interaction, and prediction of objects.
[0060] In this example, the 3D streaming engine 150 can be configured with real-time 3D rendering capabilities to enable the streaming of 3D "hologram" content. In this embodiment, the 3D "hologram" content includes point clouds, multiple views, and a streaming scene graph (an object-oriented representation of the 3D video content), as well as a dynamically animated mesh with texture streaming.
[0061] Figure 4 This is a schematic block diagram of an example embodiment of the video processing system 400. In this example, user 435 requests a 3D streaming session 145 using library file 145 (e.g., JavaScript library 455) via a first client device 135-1. Session management (i.e., session handler 125) can manage the connection to the streaming service 405 supported by a server, such as a first server 105-1. JavaScript library 455 or another library can be used to handle input controls, such as mouse, touch, and key controls. JavaScript library 455 can define specifications specific to the current session. The publicDomainName parameter 418 can be used to connect to a WebSocket server 437 to communicate with the streaming service 405 and the associated encapsulated 3D streaming engine 150 (e.g., Unreal application 450). This connection can be protected by encryption such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. The WebRTC API 115 can be configured as a proxy 417 to route incoming messages to the WebSocket server 437 and provide the aforementioned encryption. A turnserver 406 can intermediate between the proxy 417 and the streaming service 405. Additional mouse, touch, and keyboard controls can be handled between the proxy 417 and the streaming service 405. The WebSocket server 437 can receive incoming messages from the user 435 at a front-end interface, such as that of a website 436. These messages can contain an identifier of the listening port of the WebSocket server 437. Incoming and outgoing messages arriving at and departing from the WebSocket server 437 can use separate communication paths. These communication paths can be manually selected and configured. Multiple sessions of system 400 can be configured individually. The GPU instance 370 can therefore host services independently without requiring middleware or additional proxies to be configured between services.
[0062] Example flow use case
[0063] Figure 5AThis is a schematic diagram and overview of an example use case 500a of an embodiment of a video processing system (e.g., system 100). A sales representative 535-1 can connect to a remotely located client 535-2 and conduct a virtually guided sales journey via a cloud-based server implementation 505 of system 100. The advantages of this type of connection include the ability to simulate face-to-face interaction even if the point of sale is closed or otherwise inaccessible to client 535-2 or even sales representative 535-1. This implementation 505 can host multiple clients on a single instance of implementation 505. Video and chat functionality can be provided by implementation 505. Additionally, implementation 505 provides a controlled environment for distributing content related to the products offered by sales representative 535-1, thereby incurring controlled costs for providing and maintaining said environment.
[0064] The website-based streaming disclosed herein can provide a high-end 3D streaming experience to multiple users on a public website. Therefore, streaming solutions for points of sale (e.g., resellers) are available. Access to such streaming solutions can be restricted to specific users, thereby increasing platform security. Instead of purchasing and maintaining hardware such as proprietary streaming servers, points of sale can use the cloud streaming solutions described herein, saving resources that can yield increased returns when applied elsewhere, such as by improving customer experience. Use cases can be provided globally to facilitate high-quality connections between users and GPU cloud instances.
[0065] In an example embodiment, a digital sales lounge is provided for car dealerships. Dealers can act as hosts by initiating a streaming session and inviting multiple clients (e.g., up to six or more) to a separate streaming session. All participants in the streaming session can watch a public video presentation while being virtually unrestricted by their physical location.
[0066] Figure 5B It is based on Figure 5A Illustration 500b shows various example frames of the video presentation configured for use case 500a. Illustration 500b shows the start of a streaming session 578 by sharing an access code with client 535-2, client 535-2 accessing the streaming session 583 using the access code, launching the joint configuration application 588 so that sales representative 535-1 and client 535-2 can work together to configure various options in the car to be sold, and the end of the streaming session 593 with a custom video or an image rendered for an individual manual, depending on the option selected via the joint configuration application 588.
[0067] Challenges faced by existing real-time streaming applications
[0068] Other real-time streaming applications existing at the time of this disclosure require deep technical expertise to develop streaming solutions or experiences. Traditionally, setting up such streaming environments has been a highly manual process. Using existing methods, it is often difficult to create reliable, scalable solutions globally. Furthermore, existing methods often fail to deliver sufficiently high visual quality with sufficiently low latency. For existing methods using Pixelstreaming, users are locked into using Unreal Engine as a wrapper for their 3D streaming engine. Additionally, Pixelstreaming offers users a limited feature set. Existing methods require multiple service providers to build the streaming business. Therefore, streaming via existing methods is expensive, and calculating the actual cost can be a complex process. Existing cloud-based streaming solutions suffer from complex IT landscapes within companies adopting such solutions, making them difficult to deploy. These challenges result in excessively long time-to-market for real-time streaming solutions.
[0069] The aspects of the real-time streaming solutions disclosed in this article
[0070] The embodiments disclosed herein provide fully developed streaming technologies available from online marketplaces such as Google Play. These streaming solutions can be set up via a simple and highly automated process, requiring no development expertise from the user. Such streaming solutions can therefore be tailored to cater to content creators rather than software developers. Streaming can be effectively decoupled from the associated application engine, thus avoiding technical lock-in, such as regarding the specific application engine to be used. Therefore, complete streaming solutions can be provided from a one-stop shop.
[0071] Furthermore, the streaming solutions disclosed herein support a wide range of advanced technologies and deliver a high-quality, high-performance streaming experience. Implementations are available worldwide without limitations. Setup can be flexible to meet diverse user needs. Implementations use configurations such as Software as a Service (SaaS) as a blueprint and are therefore compatible with most IT guidelines within a company. These streaming solutions also come with affordable and easy-to-understand pricing structures.
[0072] Digital processing environment
[0073] Example implementations of a multimedia system 600 for streaming selected media content to a user’s client device 15 (e.g., client devices 135-1, 135-2, 135-i) may be implemented in a software, firmware, or hardware environment. Figure 6An environment like this is illustrated. One or more client devices 15 (e.g., mobile phones) and cloud 16 (or server computers or clusters thereof) provide processing, storage, and input / output means for executing applications, etc. Client devices may be interchangeably referred to herein as client computers.
[0074] Client device 15 is linked to other computing devices via communication network 17, including other client devices / programs 15 and one or more server computers 16. Communication network 17 can be a remote access network, a global network (e.g., the Internet), an out-of-band network, a worldwide collection of computers, a local area network or wide area network, a cloud network, or part of a gateway that currently communicates with each other using appropriate protocols (TCP / IP, HTTP, Bluetooth, etc.). Other electronic device / computer network architectures are suitable.
[0075] Server computer 16 (e.g., server 105-1, 105-n) may be configured to implement a streaming media server for serving, formatting, and storing selected media content (e.g., audio, video, text, and images / pictures) processed and played at client device 15. Server computer 16 is communicatively coupled to client device 15, which implements a corresponding video encoder for capturing, encoding, loading, or otherwise serving the selected media content transmitted to server computer 16. In one example embodiment, one or more server computers 16 are scalable Java application servers, allowing the server to handle increased load in the event of traffic spikes.
[0076] Figure 7 yes Figure 6 A diagram of the internal structure of a computer / computing node (e.g., a client processor / device / mobile phone device / tablet 15, 135-i, 135-1, 135-2 or a server computer / server 16, 105-1, 105-n) in a processing environment, which can be used to facilitate the display of such audio, video, image, or data signal information. Each computer 15, 135-i, 135-1, 135-2, 16, 105-1, 105-n contains a system bus 11, where a bus is a set of physical or virtual hardware lines for data transfer within components of a computer or processing system. Bus 11 is essentially a shared channel that connects different components of the computer system (e.g., processor, disk storage device, memory, input / output ports, etc.) to enable data transfer between components. I / O device interface 82 is attached to system bus 11 for connecting various input and output devices (e.g., keyboard, mouse, touchscreen interface, monitor, printer, speakers, etc.) to computer 15, 16. Network interface 86 allows the computer to connect to a network (e.g., ...). Figure 6Various other devices of the network shown in the 17 locations. Memory 24 provides volatile storage (e.g., capturing / loading, providing, formatting, retrieving, downloading and / or storing selected media content streams and user-initiated command streams) for implementing software embodiments of the present invention, including computer software instructions 25 and data 26.
[0077] The disk storage device 95 provides non-volatile storage for computer software instructions 92 (equivalent to an "OS program") and data 94 for implementing embodiments of the multimedia system 600 of the present invention. A central processing unit (CPU) 84 is also attached to the system bus 11 and provides execution of computer instructions. The processor 84 may include one or more microprocessors, microcontrollers, digital signal processors, field-programmable gate arrays (FPGAs), programmable logic devices, state machines, gated logic, or discrete hardware circuitry to handle load balancing for 3D streaming.
[0078] In this embodiment, CPU 84 is a hybrid CPU / GPU unit with an embedded vector packet processor (VPP)-based hardware accelerator. The hybrid CPU / GPU with embedded VPP is optimized for improved packet processing for multiple 3D streaming sessions, including full H.264 decoding across multiple channels. In an example embodiment, the hybrid CPU / GPU unit with embedded vector packet processor (VPP)-based hardware accelerator is configured to handle hyperscale cloud workloads of 3D streaming sessions, including 5G transport processing and 5G RAN Intelligent Controller (RIC) and edge inference. In a preferred instance, the hybrid CPU / GPU unit with embedded vector packet processor (VPP)-based hardware accelerator includes an integrated 1 terabit switch, as well as true in-line streaming and highly programmable 3D packet processing. The VPP accelerator can provide low latency and high throughput, making it suitable for deploying multiple high-speed 3D streaming sessions.
[0079] In one embodiment, the hybrid CPU / GPU unit includes an embedded hardware / firmware implementation of the 3D Streaming Engine 150, which can generate multiple executable instances of the packaged 3D Streaming Engine 150. Each packaged execution instance of the 3D Streaming Engine 150 is preferably executed within a container. In this way, the generated instances of the 3D Streaming Engine 150, along with their corresponding libraries and dependencies, are executed in a lightweight executable container that is always running and can be optimized for faster and more secure deployment in portable computing units. Unlike traditional computing methods, by executing the 3D Streaming Engine 150 in a container, the 3D Streaming Engine is more resilient to errors and inconsistencies when moved to a new location. Containerization eliminates this problem by binding application code together with the relevant configuration files, libraries, and dependencies required for its operation. This container can then potentially be extracted from the host system and become portable, capable of running flawlessly across any platform or cloud. In an embodiment, the encapsulated 3D streaming engine 150 may be implemented as a virtual machine executed from servers 16, 105-1, 105-n or via a secure socket layer deployed at client devices 135-i, 135-1, 135-2.
[0080] In one embodiment, the processor routines 92 and data 94 of the multimedia processing system (video processing system) may be implemented as a computer program product comprising computer-readable media capable of being stored on storage device 95 or deployed as Software as a Service (SaaS), providing at least a portion of software instructions for the multimedia processing system 600, the packaged 3D streaming engine 150, the GPU server 305 with 3D game applications, context verification 140, and session handler 125. Examples of software embodiments of the multimedia processing system 600 may be implemented as computer program product 92 and can be installed using any suitable software installation process known in the art. In another embodiment, at least a portion of the instructions for the multimedia processing system 600 may also be downloaded via cable, communication, and / or wireless connections. In other embodiments, the software components of the multimedia processing system 600 may be implemented as a computer program propagation signal product 77 embodying signals propagated on a propagation medium (e.g., radio waves, infrared waves, laser waves, sound waves, or radio waves propagating through a global network such as the Internet or one or more other networks). Such carrier media or signals provide at least a portion of the software instructions for the multimedia processing system 600 routines / programs 92.
[0081] In an alternative embodiment, the propagating signal is an analog carrier wave or digital signal carried on a propagating medium. For example, the propagating signal may be a digitized signal propagated over a global network (e.g., the Internet), a telecommunications network, an out-of-band network, or other networks. In one embodiment, the propagating signal is transmitted over a propagating medium for a period of time, such as instructions for a software application sent over a network in the form of data packets over milliseconds, seconds, minutes, or longer. In another embodiment, the computer-readable medium of the computer program product 92 is a propagating medium that the computer system 15 can receive and read, for example, by receiving the propagating medium and recognizing the propagating signal embodied in the propagating medium, as described above for a computer program propagating signal product.
[0082] The multimedia processing system 600 described herein can be configured using any known programming language, including any high-level object-oriented programming language. The client computer / device 15 of the multimedia system 600 can be implemented via a software embodiment and can operate within a browser session. The multimedia processing system 600 can be developed using HTML, JavaScript, Flash, etc. HTML code can be configured to embed the system into a web browsing session at the client 15. JavaScript can be configured to perform clickstreaming and session tracking at the client 15 and store streaming media recordings and editing data in a cache. In another embodiment, the system can be implemented in HTML5 for client devices 15 that do not have Flash installed and use HTTP Live Streaming (HLS) or MPEG-DASH protocols. The system can be implemented to transmit media streams using real-time streaming protocols, such as Real-Time Messaging Protocol (RTMP), Real-Time Streaming Protocol (RTSP), Web Real-Time Communication (WebRTC), etc. The components of the multimedia processing system 600 can be configured to create and load XML, JSON, or CSV data files or other structured metadata files (such as manifest files) containing information about where and how the components of the multimedia processing system 600 are stored, hosted, or formatted, such as timing information, size, footnotes, attachments, interactive components, style sheets, etc.
[0083] In an example mobile implementation, the user interface framework for the components of the multimedia processing system 600 can be based on XHP, Javelin, and WURFL. In another example mobile implementation of the OS X and iOS operating systems and their corresponding APIs, Cocoa and Cocoa Touch can be implemented using Objective-C or any other high-level programming language that adds Smalltalk-style messaging to the C programming language.
[0084] Example Advantages
[0085] The benefits offered to users by the streaming solutions disclosed herein include the following: Implementations are accompanied by low streaming costs for users (e.g., 70% cheaper than some current products). Furthermore, no additional development costs are incurred on the client side. The implementation provides a very fast and automated process, resulting in rapid turnaround times for projects created by such implementations. Additionally, due to the SaaS approach used, this streaming solution is not subject to IT regulatory restrictions. Implementations provide high-performance streaming with high visual quality and low latency. Such solutions avoid technology lock-in and are therefore open to future developments in the real-time engine and streaming fields. Cost control can be provided by automatically adjusting virtual computing instances based on actual needs. For example, thresholds can be used to ensure that a minimum number of virtual computing instances are always available. Implementations provide complete transparency through a separate dashboard for each client. Such dashboards can display information including multiple runtime instances that are configured or available. The complete solution can be obtained from a provider as a service.
[0086] This high-fidelity streaming solution can be configured to support ray tracing. 3D applications can stream according to the disclosed methods and systems without additional software plugins. External interfaces can be used to control the streaming applications. The implementation can be agnostic to cloud providers and devices that work with many such providers and user device types. Implementations can include built-in maintenance and analysis tools.
[0087] It supports any game engine or streaming engine. Modern web browsers such as Chrome, Safari, and Firefox can be used in various implementations. Multiple applications can be hosted on a single graphics card or GPU. External protocols such as HTTP and WebSocket are supported. Implementations allow for analytics reporting, such as session time, concurrent users (CCU), and other analyses. Implementations can include automatic shutdown features. Cross-region support can be achieved by this type of streaming solution. Website integration can be achieved, for example, via the included JavaScript libraries. Existing streaming applications can be easily migrated to platforms configured according to currently publicly available methods. Global security can be provided, with source files and streaming applications protected on cloud servers. Automatic scaling of virtual computing instances based on multiple connected users enables connections to the service, for example, up to one thousand users. All common application features can be supported, including animations, moving products, etc.
[0088] Although exemplary embodiments have been specifically shown and described, those skilled in the art will understand that various changes in form and detail may be made thereto without departing from the scope of the embodiments covered by the appended claims.
Claims
1. A video processing system, the video processing system comprising: Multiple servers, which operate in a peer-to-peer communication network via the WebRTC application programming interface API, to generate 3D streaming data; The multiple servers respond to the request for a 3D streaming session received at the first client device by the session handler in the following manner: (i) Receive context verification about the first client device from the session handler communicating with the first client device, the context verification including the Internet Protocol IP location of the first client device; (ii) Responding to the context verification regarding the first client device by assigning a first server among the plurality of servers based on the Internet Protocol IP location of the first client device to guide the 3D streaming session at the first client device; (iii) Guide the first instance of the encapsulated 3D streaming engine at the first server to initiate the 3D streaming session at the first client device via the WebRTC application programming interface API; (iv) Injecting the library files of the WebRTC application programming interface API into the executable file of the encapsulated 3D streaming engine to generate a digital representation of the interactive 3D environment for the 3D streaming session at the first client device; (v) Render the digital representation of the interactive 3D environment within the buffer of the graphics processing unit (GPU); (vi) The encoder encodes the rendered representation for use in transmitting the 3D streaming session to the first client device; (vii) Streaming the encoded representation of the interactive 3D environment at the first client device; (viii) The 3D streaming session of the interactive 3D environment is dynamically adjusted in response to input received from the first client device; and (ix) The digital representation of the interactive 3D environment of the 3D streaming session is maintained by iteratively processing (v) to (viii) until the 3D streaming session terminates at the first client device.
2. The video processing system of claim 1, wherein the first server is configured to establish a predetermined number of virtual computing instances with one or more client devices for the 3D streaming session; and The first server is configured to automatically adjust the number of virtual computing instances provided based on changes in the number of the one or more client devices, at least one of which is the first client device.
3. The video processing system of claim 2, wherein at least one of the plurality of servers is configured to respond to a request for a 3D streaming session from the one or more client devices by assigning a server among the plurality of servers to respond to a corresponding request for a 3D streaming session. The first server is one of the plurality of servers.
4. The video processing system of claim 3, wherein at least one of the plurality of servers performs computational analysis to determine an appropriate number of 3D streaming sessions for the plurality of client devices, adjusted based on the predetermined number of virtual computing instances.
5. The video processing system according to claim 4, wherein the calculated analysis includes concurrent user CCU, session time, daily active user DAU, monthly active user MAU, and session.
6. The video processing system of claim 3, wherein the encapsulated 3D streaming engine configures the 3D streaming session as a media streaming cluster, such that at least one of the plurality of servers performs multicast transmission of data packets containing at least a portion of the digital representation of the interactive 3D environment to the one or more client devices.
7. The video processing system of claim 6, wherein the 3D streaming session that dynamically adjusts the interactive 3D environment in response to input received from the first client device further comprises: The encapsulated 3D streaming engine is configured to capture events associated with the interactive 3D environment, which are detected by one or more listeners at one or more of the one or more client devices. The encapsulated 3D streaming engine is configured to capture metadata, consumption information, and interaction data during the 3D streaming session in one or more of the one or more client devices. The encapsulated 3D streaming engine is configured to respond to captured events, metadata, consumed information, and interactive data by redrawing the interactive 3D environment at one or more client devices, in order to maintain the digital representation of the interactive 3D environment.
8. The video processing system of claim 7, wherein the encapsulated 3D streaming engine maintains the digital representation of the interactive 3D environment in the following manner: Composite 3D media data is calculated based on the captured events, metadata, consumption information, and interaction data; The composite image layout is defined based on attributes derived from the 3D streaming session at the one or more client devices. The 3D streaming session is configured to provide composite media signals based on the defined composite image layout; and The composite media signal is transmitted to the one or more client devices in the form of data packets via a packetizer.
9. The video processing system of claim 6, further comprising the encapsulated 3D streaming engine interfacing with the one or more client devices to facilitate a control communication interface, thereby enabling voice, text, and video transmission among the one or more client devices in a peer-to-peer packet-switched communication system.
10. The video processing system according to claim 7, wherein, The encapsulated 3D streaming engine is configured to train an interactive frame prediction model for the digital representation of the interactive 3D environment based on the encoded streaming data of the 3D streaming session and the captured events, metadata, interaction data, and consumption information from the one or more client devices.
11. The video processing system of claim 10, wherein the encoder is configured to encode the 3D streaming session by selectively using predicted frames generated based on a trained interactive frame prediction model, and to transmit the trained interactive frame prediction model and the encoded streaming data to the first client device and the second client device to create the digital representation of the interactive 3D environment; and The first client device and the second client device are configured to receive the trained interactive frame prediction model and the encoded stream data, and decode the encoded stream data based on the trained interactive frame prediction model to create the digital representation of the interactive 3D environment.
12. The video processing system of claim 1, wherein the encapsulated 3D streaming engine is configured as a software container, the software container containing bindings of software components having configuration files, libraries and dependencies required for execution, wherein the encapsulated 3D streaming engine is not a software plugin.
13. The video processing system of claim 12, wherein the software components comprise at least one of the following: the library files of the WebRTC application programming interface API, an event manager, a scene manager, a resource manager, a session manager, a physical manager, and an artificial intelligence system; and The WebRTC application programming interface (API) includes one or more WebRTC API functions configured to call one or more of the software components.
14. The video processing system of claim 12, wherein the encapsulated 3D streaming engine is deployed as a virtual machine via Secure Sockets Layer (SSL), and the virtual machine is configured within the software container; The input is received via a data channel established using the Real-Time Communication Data Channel (RTCDataChannel) application programming interface (API).
15. A video processing method, the video processing method comprising: Configure the server computer system to respond to requests for 3D streaming sessions received at the first client device by the session handler in the following manner: (i) Receive context verification about the first client device from the session handler communicating with the first client device, the context verification including the Internet Protocol IP location of the first client device; (ii) A 3D streaming session at the first client device is directed to respond to the context verification regarding the first client device by assigning a first server among a plurality of servers of the server computer system based on the Internet Protocol IP location of the first client device, wherein the plurality of servers operate in a peer-to-peer communication network via the WebRTC application programming interface API to generate 3D streaming data. (iii) Guide the first instance of the encapsulated 3D streaming engine at the first server to initiate the 3D streaming session at the first client device via the WebRTC application programming interface API; (iv) Injecting the library files of the WebRTC application programming interface API into the executable file of the encapsulated 3D streaming engine to generate a digital representation of the interactive 3D environment for the 3D streaming session at the first client device; (v) Render the digital representation of the interactive 3D environment within the buffer of the graphics processing unit (GPU); (vi) The rendered representation is encoded via an encoder for use in transmitting the 3D streaming session to the first client device; (vii) Streaming the encoded representation of the interactive 3D environment at the first client device; (viii) The 3D streaming session of the interactive 3D environment is dynamically adjusted in response to input received from the first client device; and (ix) The digital representation of the interactive 3D environment of the 3D streaming session is maintained by iteratively processing (v) to (viii) until the 3D streaming session terminates at the first client device.
16. The video processing method of claim 15, further comprising configuring the first server to establish a predetermined number of virtual computing instances with one or more client devices for the 3D streaming session; and The first server is configured to automatically adjust the number of virtual computing instances provided based on changes in the number of the one or more client devices, at least one of which is the first client device.
17. The video processing method of claim 16, further comprising configuring at least one of the plurality of servers to respond to a request for a 3D streaming session from one or more client devices by assigning one of the plurality of servers to respond to a corresponding request for a 3D streaming session, wherein the first server is one of the plurality of servers.
18. The video processing method of claim 17, further comprising configuring at least one of the plurality of servers to perform computational analysis to determine an appropriate number of 3D streaming sessions of the plurality of client devices, adjusted based on the predetermined number of virtual computing instances.
19. The video processing method according to claim 18, further comprising configuring the calculated analysis to include concurrent user CCU, session time, daily active user DAU, monthly active user MAU, and sessions.
20. The video processing method of claim 17, further comprising configuring the 3D streaming session as a media streaming cluster such that at least one of the plurality of servers performs multicast transmission of data packets containing at least a portion of the digital representation of the interactive 3D environment to the one or more client devices.
21. The video processing method of claim 20, wherein the 3D streaming session that dynamically adjusts the interactive 3D environment in response to input received from the first client device further comprises: The encapsulated 3D streaming engine is configured to capture events associated with the interactive 3D environment, which are detected by one or more listeners at one or more of the one or more client devices. Configure the encapsulated 3D streaming engine to capture metadata, consumption information, and interaction data during the 3D streaming session in one or more of the one or more client devices; The encapsulated 3D streaming engine is configured to respond to captured events, metadata, consumed information, and interactive data by redrawing the interactive 3D environment at one or more client devices, thereby maintaining the digital representation of the interactive 3D environment.
22. The video processing method of claim 21, wherein the encapsulated 3D streaming engine maintains the digital representation of the interactive 3D environment in the following manner: Composite 3D media data is calculated based on the captured events, metadata, consumption information, and interaction data; The composite image layout is defined based on attributes derived from the 3D streaming session at the one or more client devices. The 3D streaming session is configured to provide composite media signals based on the defined composite image layout; and The composite media signal is transmitted to the one or more client devices in the form of data packets via a packetizer.
23. The video processing method of claim 20, further comprising interfacing the encapsulated 3D streaming engine with the one or more client devices to facilitate a control communication interface, thereby enabling voice, text, and video transmission among the one or more client devices in a peer-to-peer packet-switched communication system.
24. The video processing method of claim 21, further comprising configuring the encapsulated 3D streaming engine to train an interactive frame prediction model for the digital representation of the interactive 3D environment based on the encoded streaming data of the 3D streaming session and the captured events, the metadata, the interaction data, and the consumption information from the one or more client devices; One or more of the plurality of servers include a hybrid CPU / GPU unit with an embedded packaged 3D streaming engine, the embedded packaged 3D streaming engine having a hardware accelerator based on a vector packet processor (VPP), the hybrid CPU / GPU unit being optimized to improve packet processing for the plurality of the 3D streaming sessions, including full H.264 decoding of multiple channels.
25. The video processing method of claim 15, further comprising configuring the encapsulated 3D streaming engine as a bound container containing software components having configuration files, libraries, and dependencies required for execution; wherein the encapsulated 3D streaming engine is not a software plugin.
26. The video processing method of claim 25, wherein the software component comprises at least one of the following: the library file of the WebRTC application programming interface API, the event manager, the scene manager, the resource manager, the session manager, the physical manager, and the artificial intelligence system; and The WebRTC application programming interface (API) includes one or more WebRTC API functions configured to call one or more of the software components.
27. The video processing method of claim 15, further comprising deploying the encapsulated 3D streaming engine as a virtual machine via Secure Sockets Layer (SSL), the virtual machine being configured in a software container; The input is received via a data channel established using the Real-Time Communication Data Channel (RTCDataChannel) application programming interface API.
28. The video processing method of claim 27, further comprising configuring the encoder to encode the 3D streaming session by selectively using predicted frames generated based on a trained interactive frame prediction model, and transmitting the trained interactive frame prediction model and the encoded streaming data to the first client device and the second client device to create the digital representation of the interactive 3D environment; and The first client device and the second client device are configured to receive the trained interactive frame prediction model and the encoded stream data, and to decode the encoded stream data based on the trained interactive frame prediction model to create the digital representation of the interactive 3D environment.
Citation Information
Patent Citations
Conversational framework
CN110321413A
Building intelligent control three-dimensional model display method and related equipment
CN112783064A