Spatial Perception Multimedia Router System and Method

By adopting spatially aware multimedia router system and method in the video conferencing system, analyzing and processing multimedia streams, optimizing the forwarding of multimedia streams based on user priority and spatial relationships, the scalability and resource utilization problems of multi-user video conferencing systems in the priorities are solved, and efficient user experience quality is achieved.

CN114205548BActive Publication Date: 2025-07-01TMRW FOUNDATION IP SARL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110973122.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-28
Filing Date
2021-08-24
Publication Date
2025-07-01
Estimated Expiration
2041-08-24

AI Technical Summary

Technical Problem

The existing video conferencing system has limited scalability in multi-user scenarios, resulting in inefficient use of network bandwidth and computing resources, affecting the quality of user experience.

Method used

A space-aware multimedia router system and method is adopted, which includes a media server and a virtual environment, optimizes the forwarding of multimedia streams based on user priority and spatial relationships, and improves bandwidth and computing resource utilization by analyzing and processing multimedia streams from client devices.

Benefits of technology

It realizes the optimization of network bandwidth and computing resources in multi-user video conferencing scenarios, improves user experience quality, and enhances the scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114205548B_ABST
    Figure CN114205548B_ABST
Patent Text Reader

Abstract

A spatial perception-based multimedia routing system, comprising: at least one media server computer for receiving and analyzing incoming data, the incoming data including incoming multimedia streams from client devices; adjusting outgoing multimedia streams for respective client devices based on the incoming data received from the client devices. The incoming multimedia streams include elements from a virtual environment. Adjusting the outgoing multimedia streams for respective client devices based on user priority data and spatial orientation data, the spatial orientation data describing a spatial relationship between a corresponding user graphical representation in the virtual environment and a source of the incoming multimedia stream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application is related to co - filed U.S. Patent Application 17 / 005,767, which was filed concurrently herewith and is titled "Spatial Awareness Multimedia Router System and Method", and which is incorporated herein by reference. Technical Field

[0003] This application generally relates to computer systems, and more particularly, to multimedia router systems and methods. Background Art

[0004] Video conferencing enables remote communication among multiple users and is becoming a relatively low - cost and fast communication tool for people in multiple locations. Due to the widespread deployment of broadband networks, the advancement of video compression technology, and the increasing number of network - based video communication implementation methods with lower infrastructure requirements and costs, video conferencing has recently become popular.

[0005] For example, one method of implementing video conferencing is a mesh (peer - to - peer) infrastructure, where each client device sends multimedia streams to all other client devices. This represents a low - cost solution that does not require any intermediate infrastructure, but due to rapid bandwidth overload and limited processing power of client devices, it results in low scalability.

[0006] Another example method is a Multipoint Control Unit (MCU), which is implemented in a central media server, receives all multimedia streams from client devices, decodes and combines all the streams into one stream, and then re - encodes and sends it to all client devices, reducing the latency and bandwidth issues associated with the peer - to - peer model. However, MCU implementations are often complex and require a large amount of computing resources from the media server.

[0007] Another example method is the Selective Forwarding Unit (SFU), which is used in the Web Real-Time Communication (WebRTC) video conferencing standard. The WebRTC standard generally supports browser-to-browser applications, such as voice calls, video chats, and peer-to-peer (P2P) file sharing applications, while avoiding the need to use plugins to connect video communication endpoints. The SFU can be implemented in a central media server computer. The SFU includes a software program for routing video packets in a video stream to multiple participant devices without performing intensive media processing (such as decoding and re-encoding) on the server side. Thus, the SFU receives all encoded media streams from client devices over the network and then selectively forwards these streams to the client devices of the corresponding participants for subsequent decoding and display. The selectivity of the SFU forwarding is based on multiple parameters, which can be used to optimize the bandwidth associated with multimedia stream forwarding, thereby resulting in a higher Quality of Experience (QoE). For example, the SFU can identify the participant who is speaking from the received multimedia stream and forward the high-bitrate multimedia stream to the participants who are listening. On the other hand, the SFU sends the low-bitrate multimedia streams of the listening participants to other participants, achieving a certain degree of improvement in bandwidth and QoE.

[0008] Given the limitations of the bandwidth and processing capabilities of the central media or routing server or the participating client devices, the limitation of typical video conferencing tools (such as through the above methods) lies in limited scalability. Therefore, there is a need for new methods to further optimize network bandwidth and computing resources during multimedia routing and forwarding operations while maintaining a high QoE for the relevant participants. Summary of the Invention

[0009] This summary is provided to introduce some concepts in a simplified form that will be further described in the detailed description below. This summary is not intended to identify key features of the subject matter to be protected, nor is it intended to assist in determining the scope of the subject matter to be protected.

[0010] In one aspect of the present application, a spatially aware multimedia router system is provided. The spatially aware multimedia router system includes at least one processor and a memory that stores instructions for a data exchange management module that manages data exchange between client devices. In one embodiment, the system further includes one or more computing devices that implement at least one virtual environment that is connected to at least one media server computer and is thus capable of accessing one or more graphical representations (also referred to as user graphical representations) of users of multiple client devices. In view of the virtual elements in the virtual environment and the input data from at least one client device, multiple multimedia streams (e.g., 2D video streams, 3D video streams, audio streams, or combinations of these or other media streams) can be generated in the virtual environment. Thus, the input data is received and combined in the virtual environment, which includes multiple virtual elements and at least one graphical representation of the corresponding users of the client devices. The multiple client devices are connected to the at least one media server via a network, and the multiple client devices are configured to send data including multimedia streams to the at least one media server computer.

[0011] The at least one media server is configured to: receive and analyze incoming data from client devices, the incoming data including incoming multimedia streams from the client devices; and adjust the outgoing multimedia streams for each client device based on the incoming data. The incoming multimedia streams include elements from within at least one virtual environment. The outgoing multimedia streams are adjusted for each client device based on, for example, user priority data and spatial orientation data, the spatial orientation data describing, for example, the spatial relationship between the corresponding user graphical representation in at least one virtual environment and the source of the incoming multimedia stream.

[0012] In one embodiment, the at least one media server computer performs data exchange management, including analyzing and processing incoming data that includes multimedia streams from client devices, and the data exchange management further includes evaluating and optimizing the forwarding of the outgoing multimedia streams based on the incoming data received from the multiple client devices, the incoming data including elements from within the at least one virtual environment. The incoming data is associated with user priority data and the spatial relationship between the corresponding user graphical representation and the incoming multimedia streams.

[0013] In some embodiments, the at least one virtual environment is hosted on at least one dedicated server computer, and the at least one dedicated server computer is connected to the at least one media server computer via a network. In other embodiments, the at least one virtual environment is hosted in a peer-to-peer infrastructure and relayed through the at least one media server computer. The virtual environment can be used to host real-time video communications where users can interact with each other, and can also be used for meetings, work, education, shopping, entertainment, and services, etc. In some embodiments, the virtual environment is a virtual copy of a real-world location, where the real-world location includes multiple sensors that provide further data to the virtual copy of the real-world location.

[0014] In some embodiments, the at least one media server computer uses a routing topology. In other embodiments, the at least one media server computer uses a media processing topology. In other embodiments, the at least one media server computer uses a forwarding server topology. In other embodiments, the at least one media server computer uses other suitable multimedia server routing topologies, or media processing and forwarding server topologies, or other suitable server topologies.

[0015] In one embodiment, the at least one media server computer uses a routing topology, the at least one media server computer uses a Selective Forwarding Unit (SFU) topology, or uses a Traversal Using Relay NAT (TURN) topology, or a Spatial Analysis Media Server topology (SAMS), or some other multimedia server routing topology.

[0016] In one embodiment, the at least one media server computer uses a media processing topology, and the at least one media server computer is used to perform one or more operations on the incoming data, including: compressing, encrypting, re-encrypting, decrypting, decoding, combining, improving, mixing, enhancing, calculating, manipulating, or encoding, or a combination thereof. In a further embodiment, the combination of the incoming data is performed in the form of a mosaic, and the mosaic includes individual tiles, where the respective multimedia streams of the user graphical representation are streamed.

[0017] In one embodiment, the at least one media server computer uses a forwarding server topology, and the at least one media server computer serves as a Multipoint Control Unit (MCU), or a cloud media mixer, or a cloud 3D renderer.

[0018] In one embodiment, the at least one media server computer is used for SAMS, and the at least one media server computer is used to analyze and process the incoming data of each client device, and the incoming data is associated with user priorities and spatial relationships (e.g., the spatial relationship between the corresponding user graphical representation and the source of the incoming multimedia stream). In this embodiment, the at least one media server computer can also be used to determine user priorities and / or spatial relationships based on the above data. In some embodiments, the incoming data includes one or more of the following: metadata, priority data, data classes, spatial structure data, scene graphs, three-dimensional positions, orientation or motion information, speaker or listener status data, availability status data, image data or video based on a scalable video codec, or a combination thereof. In a further embodiment, adjusting the outgoing multimedia stream (e.g., implemented by at least one media server computer implementing SAMS) includes optimizing the bandwidth and computing resource utilization of one or more receiving client devices. Adjusting the outgoing multimedia stream also includes adjusting temporal characteristics, spatial characteristics, quality, or color characteristics, or a combination thereof. In yet a further embodiment, SAMS optimizes the forwarding of the outgoing data stream to each receiving client device by modifying, magnifying, or reducing the temporal characteristics, spatial characteristics, quality, and color characteristics of the media.

[0019] On the other hand, the present application provides a method for a space-aware multimedia router. The method includes: providing, in the memory of at least one media server computer, data and instructions for implementing a data exchange management module for client devices, the module managing data exchange among a plurality of client devices; the at least one media server computer receiving incoming data, the incoming data including incoming multimedia streams from a plurality of clients, wherein the incoming data is associated with user priority data and spatial orientation data. For example, the spatial orientation data can describe, for example, the spatial relationship between the corresponding user graphical representation and one or more sources of the incoming multimedia stream. Subsequently, in the method, the data exchange management module performs data exchange management. In one embodiment, the data exchange management includes: analyzing and / or processing the incoming data from a plurality of client devices, the incoming data including graphical elements from within the virtual environment; and adjusting the outgoing multimedia stream based on the incoming data received from the plurality of client devices. Finally, in the method, the adjusted outgoing multimedia stream is forwarded to one or more receiving client devices, wherein the adjusted outgoing multimedia stream is for display at the receiving client device (e.g., displayed to a user represented by a user graphical representation).

[0020] In some embodiments, the method further includes: when forwarding an outgoing multimedia stream, optimizing a routing topology, or a media processing topology, or a forwarding server topology, or other suitable multimedia server routing topologies, or media processing and forwarding server topologies, or other suitable server topologies.

[0021] In some embodiments, in the routing topology, the at least one media server computer uses a Selective Forwarding Unit (SFU) topology, uses a Traversal Using Relay NAT (TURN) topology, or a Spatial Analysis Media Server topology (SAMS), or other multimedia server routing topologies.

[0022] In some embodiments, in the media processing topology, the at least one media server computer is used to perform one or more media processing operations on the incoming data, including: compressing, encrypting, re-encrypting, decrypting, decoding, combining, improving, mixing, enhancing, calculating, manipulating, or encoding, or combinations thereof.

[0023] In some embodiments, the method further includes: when utilizing a forwarding router topology, utilizing a Multipoint Control Unit (MCU), a cloud media mixer, or a cloud 3D renderer.

[0024] In some embodiments, using a SAMS configuration, the method further includes: analyzing and processing the incoming data of each client device related to user priorities and spatial relationships (such as, the distance relationship or other spatial relationships between the user graphical representation and the incoming multimedia stream). In this embodiment, the method further includes: determining user priorities and / or spatial relationships based on the above data. The above incoming data includes one or more of the following: metadata, priority data, data classes, spatial structure data, scene graphs, three-dimensional positions, orientations, or motion information, speaker or listener status data, availability status data, image data, or video based on a scalable video codec, or combinations thereof.

[0025] In some embodiments, adjusting the outgoing multimedia stream (e.g., implemented by at least one media server computer implementing SAMS) includes optimizing the bandwidth and computing resource utilization of one or more receiving client devices. Adjusting the outgoing multimedia stream also includes adjusting temporal characteristics, spatial characteristics, quality, or color characteristics, or combinations thereof. In a further embodiment, SAMS optimizes the forwarding of the outgoing data stream to each receiving client device by modifying, amplifying, or reducing the temporal characteristics, spatial characteristics, quality, and color characteristics of the media.

[0026] In another aspect of the present application, instructions are stored on a computer-readable medium for causing one or more computing devices to perform any of the techniques described herein. In one embodiment, instructions are stored on at least one computer-readable medium for causing at least one media server computer to perform the following steps, the at least one media server computer including a processor and a memory: the at least one media server computer receives incoming data, the incoming data including incoming multimedia streams from a plurality of client devices, wherein the incoming data is associated with spatial orientation data that describes the spatial relationship between one or more user graphical representations and at least one element of at least one virtual environment; analyzes the incoming data from the plurality of client devices; adjusts an outgoing multimedia stream based on the incoming data received from the plurality of client devices; and forwards the adjusted outgoing multimedia stream to a receiving client device, wherein the adjusted outgoing multimedia stream is configured to be displayed at the receiving client.

[0027] The above summary does not include a detailed listing of all aspects of the present application. The present application includes all systems and methods that can be practiced from the various aspects summarized above, as well as those disclosed in the following detailed description, particularly those set forth in the claims filed herewith. These combinations have advantages not specifically recited in the above summary. Other features and advantages of the present application are apparent in the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The specific features, aspects, and advantages of the present application will be better understood with reference to the following embodiments and drawings, in which:

[0029] Figure 1 is a schematic diagram of a routing topology of a conventional Selective Forwarding Unit (SFU).

[0030] Figure 2 is a schematic diagram of a spatial awareness multimedia router system according to an embodiment.

[0031] Figure 3 is a schematic diagram of a system for SAMS including at least one media server computer according to an embodiment.

[0032] Figure 4A is a schematic diagram of a virtual environment.

[0033] Figure 4B is a schematic diagram of the forwarding of an outgoing media stream from a speaking user within a virtual environment of a SAMS employing the present application.

[0034] Figures 5A - 5BSchematic diagram of a usage scenario according to an embodiment, where SAMS combines media streams of multiple client devices.

[0035] Figure 6 Block diagram of the spatial awareness multimedia router method of the present application according to an embodiment. Detailed implementation

[0036] In the following embodiments, the embodiments are described with reference to the accompanying drawings. In addition, several embodiments are described by the following examples. It should be understood that these embodiments may include design and structural changes without departing from the scope of the claimed subject matter.

[0037] The present application provides a spatial awareness multimedia router system and method for receiving input data from multiple client devices and performing data exchange management on the input data. The input data is received and combined in a virtual environment that includes multiple virtual elements and at least one graphical representation of the corresponding user of the client device. The virtual environment can be used to host real-time video communications in which users can interact with each other and can be used for meetings, work, education, shopping, entertainment, and services, etc. Data exchange management includes analyzing and processing incoming data, which at least includes multimedia streams from client devices (such as, 2D video streams, 3D video streams, audio streams, or combinations of these streams or other media streams), and also includes evaluating and optimizing the forwarding of outgoing multimedia streams based on the incoming data received from multiple client devices, the incoming data including elements from within at least one virtual environment. The incoming data is associated with user priority data and the spatial relationship between the corresponding user graphical representation and the incoming multimedia stream. Therefore, the system and method of the present application enable the routing of multimedia streams received from client devices to occur within a virtual environment, while optimizing the forwarding of input data and outgoing multimedia streams to perform an optimal selection of receiving client devices, while ensuring the efficiency of bandwidth and computing resources. These efficiencies can make the spatial awareness multimedia router system and method of the present application a viable and effective option for handling multi-user video conferences that include a large number (such as, hundreds or thousands) of users accessing the virtual environment.

[0038] Figure 1 Shows a schematic diagram of a conventional Selective Forwarding Unit (SFU) routing topology 100.

[0039] This example's conventional SFU routing topology 100 includes at least one media server computer 102, which has at least one processor 104 and a memory 106. The memory 106 stores a computer program that implements an SFU 108, which is used to provide video forwarding in a real-time video application. Multiple client devices 110 can communicate in real time through the SFU 108, where the SFU 108 forwards an outgoing media stream to the client devices 110 based on one or more parameters. The outgoing media stream includes Real-time Transport Protocol (RTP) video packets. For example, client device B can represent the current speaker in a real-time communication with client device A. Client device B sends two or more media streams to client device A. Among them, for example, one media stream is sent at a high resolution (B) 112, and one media stream is sent at a low resolution (B) 114. Additionally, in this example, client device A also sends two or more media streams to client device B, such as a media stream with a high resolution (A) 116 and a media stream with a low resolution (A) 118. Client device B receives the low resolution (A) 118 media stream from client device A because client device A can be used by a passive user who is only listening to the user of client device B at that time. On the other hand, client device A can receive the high resolution (B) 112 media stream from client device B because the user of client device B can be an active user who is currently speaking. At least one media server computer 102 and two or more client devices 110 are connected through a network, which can be, for example, one or more wired or wireless communication networks (such as local area network (LAN), wide area network (WAN), the Internet, paths, links, and any intermediate node network hardware such as routers, gateways, firewalls, switches, etc.). This video application can utilize, for example, the WebRTC standard to implement a multi-party video conferencing application.

[0040] The SFU routing topology 100 can include the following limitations: regardless of the state of the client device 110, each media stream must be sent to each client device 110; media modification cannot be enabled, such as data operations on the media stream (such as enhancement, improvement, combination, etc.); and only limited parameters are considered for prioritizing and optimizing the outgoing media stream, resulting in sub-optimal network bandwidth and resource optimization of the system, and ultimately limiting the number of users who can simultaneously interact with and view the multimedia stream. Finally, the conventional SFU routing topology 100 is not optimal for video communication and interaction in a virtual environment, such as social interaction.

[0041] Figure 2 A schematic diagram of a spatial awareness multimedia router system 200 according to an embodiment is shown.

[0042] The spatial awareness multimedia router system 200 includes at least one media server computer 202 having at least one processor 204 and a memory 206 that stores instructions implementing a data exchange management module 208 that manages data exchange between client devices 210 connected to the at least one media server computer 202 via a network 212. The spatial awareness multimedia router system 200 also includes at least one virtual environment server 220 connected to the at least one media server computer 202 to access one or more user graphical representations of a user 216 in a virtual environment 214 among a plurality of client devices 210. A plurality of multimedia streams are generated in the virtual environment 214, including multimedia streams obtained by obtaining real-time feed data from one or more users 206 via a camera 218 and graphical elements in the virtual environment.

[0043] The at least one media server computer 202 is a server computer device that includes resources (e.g., at least one processor 204 and a memory 206 having network access capabilities) for performing the techniques disclosed herein, such as storing, processing, routing, and forwarding input data received from a plurality of client devices 210. The at least one media server computer 202 performs client data exchange management via the data exchange management module 208, including: analyzing and processing incoming data, and adjusting outgoing multimedia streams, where the incoming data includes multimedia streams from the client devices 210 that include graphical elements within the virtual environment 214. In one embodiment, it includes evaluating and optimizing the forwarding of outgoing multimedia streams based on incoming data received from a plurality of client devices 210. The outgoing multimedia streams are adjusted for each client device based on incoming data, such as user priority data and spatial orientation data that describes a spatial relationship, for example, between a corresponding user graphical representation and the source of an incoming multimedia stream in at least one virtual environment. In one embodiment, the incoming data is associated with a spatial relationship between one or more user graphical representations of at least one virtual environment 214 and at least one element.

[0044] In some embodiments, the at least one virtual environment 214 is a virtual copy of a real-world location. The real-world location may include a plurality of sensors that provide real-world data to the virtual copy of the real-world location via the virtual environment 214. The sensors send captured data to the virtual environment server 220 via the network 212, and the at least one media server 202 may utilize the captured data to update, augment, and synchronize corresponding virtual copies of real-world elements in at least one virtual environment. Additionally, one or more media servers 202 may also be used to merge real-world data captured by the sensors and virtual data in the virtual environment 214 into the virtual environment 214 to enhance the real-world data with virtual data.

[0045] In the present application, the term "augmentation" is used to describe the act of providing more attributes to a virtual replica based on multi-source data. An augmented virtual replica is considered a special form of updating a virtual replica by using one or more new forms of data that did not previously exist in the virtual replica. For example, augmenting a virtual replica can refer to providing real-world data captured from the sensing mechanisms of multiple devices. Further real-world data can include, for example, video data, temperature data, real-time energy consumption data, real-time water consumption data, speed or acceleration data, etc.

[0046] In one embodiment, at least one virtual environment 214 is hosted on at least one virtual environment server computer 220, which includes at least one processor 222 and a memory 224 that implement the virtual environment 214 described above. In another embodiment, at least one virtual environment 214 is hosted in a Peer-to-Peer (P2P) infrastructure and is relayed to multiple client devices 210 through at least one media server computer 202.

[0047] The arrangement of at least one virtual environment 214 can be associated with one or more themes, such as for meetings (e.g., as a virtual meeting room), work (e.g., as a virtual office space), education (e.g., as a virtual classroom), shopping (e.g., as a virtual store), entertainment (e.g., karaoke, event hall or arena, theater, nightclub, sports field or stadium, museum, cruise ship, video game, etc.), and services (e.g., hotel, travel agency or restaurant reservation or food ordering, government agency services, etc.). Combinations of virtual environments 214 from the same and / or different themes can form a virtual environment cluster, which includes hundreds or even thousands of virtual environments (e.g., multiple virtual classrooms are part of a virtual school). The virtual environment 214 can be a 2D or 3D virtual environment, including a physical arrangement or visual appearance associated with the theme of the virtual environment 214, which can be customized by the user according to their preferences or needs. The user can access the virtual environment 214 through a graphical representation, insert the graphical representation into the virtual environment 214, and combine it graphically with the two-dimensional or three-dimensional virtual environment 214.

[0048] At least one virtual environment server computer 220 or P2P infrastructure may provide corresponding resources (e.g., memory, network, and computing capabilities) to each virtual environment 214. One or more users 216 may access at least one virtual environment 214 via a client device 210 through a graphical user interface. The graphical user interface may be included in a downloadable client application or a web browsing application, which uses, e.g., the WebRTC standard, to provide the application data and instructions required to execute the selected virtual environment 214 and to implement various interactions therein. In addition, each virtual environment 214 may include one or more human or artificial intelligence (AI) hosts or assistants, which may provide the required data and / or services through their corresponding user graphical representations to assist the users within the virtual environment 214. For example, an artificial or AI bank service staff may assist the users of a virtual bank by providing the required information in the form of demonstrations, tables, lists, etc. according to the users' requirements.

[0049] In some embodiments, the user graphical representation is a user 3D virtual cutout, which may be composed of one or more input images, such as a background-removed photo uploaded by the user or from a third-party source; or a background-removed user real-time 3D virtual cutout, which is generated based on input data, such as real-time 2D, stereoscopic, or depth image or video data, or 3D video data in a real-time video stream data feed obtained from a camera, including the user's real-time video stream, or a video without a removed background, or a video with a removed background. In some embodiments, a polygon structure may be utilized to render and display the user graphical representation. Such a polygon structure may be a quadrilateral structure or a more complex 3D structure, which serves as a virtual framework to support the video. In other embodiments, one or more user graphical representations are inserted into three-dimensional coordinates within the virtual environment 214 and graphically combined therein.

[0050] In this application, the term "user 3D virtual cutout" refers to a virtual copy of a user constructed based on 2D photos uploaded by the user or from third-party sources. Using the 2D photos uploaded by the user or from third-party sources as input data, a 3D mesh or 3D point cloud of the user with the background removed is generated, and the user 3D virtual cutout is created through a 3D virtual reconstruction process via machine vision technology. In this application, the term "user real-time 3D virtual cutout" refers to a virtual copy of a user obtained by removing the user's background based on real-time 2D or 3D live video stream data feeds obtained from a camera. Using the user real-time data feeds as input data, a 3D mesh or 3D point cloud of the user with the background removed is generated, and the user real-time 3D virtual cutout is created through a 3D virtual reconstruction process via machine vision technology. In this application, the term "background-removed video" refers to a video streamed to a client device, where a background removal process has been performed on the video so that only the user is visible, and then the user is displayed on the receiving client device using a polygon structure. In this application, an "un-background-removed video" refers to a video streamed to a client device, where the video truthfully represents what the camera captures so that the user and their background are visible, and then the user and their background are displayed on the receiving client device using a polygon structure.

[0051] The P2P infrastructure can use a suitable P2P communication protocol to enable real-time communication between client devices 210 in the virtual environment 214 through a suitable application programming interface (API), thereby achieving real-time interaction and synchronization. An example of such a suitable P2P communication protocol is the WebRTC communication protocol, which is a collection of standards, protocols, and JavaScript APIs that, when combined, enable P2P audio, video, and data sharing between peer client devices 210. Client devices 210 using the P2P infrastructure use, for example, one or more rendering engines to perform real-time 3D rendering of real-time sessions. An exemplary rendering engine can be a WebGL-based 3D engine. WebGL is a JavaScript API for rendering 2D and 3D graphics in any compatible web browser without using a plugin, thereby allowing for accelerated use and effects of physical and image processing through one or more processors (such as one or more graphics processing units (GPUs)) of at least one client device 210. In addition, client devices 210 using the P2P infrastructure can perform image and video processing as well as machine learning computer vision techniques through one or more suitable computer vision libraries. An example of a suitable computer vision library is OpenCV, which is a programming function library mainly used for real-time computer vision tasks.

[0052] In some embodiments, at least one media server computer 202 uses a routing topology. In another embodiment, at least one media server computer 202 uses a media processing topology. In another embodiment, at least one media server computer 202 uses a forwarding server topology. In another embodiment, at least one media server computer 202 uses other suitable multimedia server routing topologies, or media processing and forwarding server topologies, or other suitable server topologies. The topology used by at least one media server computer 202 can depend on the processing capabilities of the client devices and / or at least one media processor computer, and also on the capabilities of the network infrastructure used.

[0053] In some embodiments, in a media processing topology, at least one media server computer 202 is used to perform one or more media processing operations on incoming data, including compression, encryption, re-encryption, decryption, decoding, combining, improving, mixing, enhancing, computing, manipulating, encoding, or a combination thereof. Thus, in a media processing topology, at least one media server computer 202 is used not only to route and forward incoming data, but also to perform multiple media processing operations that can enhance or otherwise modify the outgoing multimedia stream to the client device 210.

[0054] In some embodiments, in a forwarding server topology, at least one media server computer 202 acts as a multipoint control unit (MCU), or a cloud media mixer, or a cloud 3D renderer. As an MCU, at least one media server computer 202 is used to receive all multimedia streams from client devices, and decode and combine all media streams into one stream, which is re-encoded and sent to all client devices 210. As a cloud media mixer, at least one media server computer 202 is used to select between different multimedia sources (e.g., audio and video) from, for example, multiple client devices 210 and at least one virtual environment 214, and to mix the input data multimedia streams and add materials and / or special effects in order to create a processed output multimedia stream for the client devices 210. The range of visual effects can be, for example, from simple mixing and wiping to complex effects. As a cloud 3D renderer, at least one media server computer 202 is used to compute 3D scenes from a virtual environment through a large number of computer calculations to generate a final animated multimedia stream, which is sent back to the client devices 210.

[0055] In a routing topology, at least one media server computer 202 is used to determine where to send a multimedia stream (e.g., to which one of one or more client devices 210), which can be performed through an Internet Protocol (IP) routing table to select the interface for the best route. Such routing decisions in the routing table are based on rules stored in the data exchange management module 208, which take into account priority data and the spatial relationship between the user graphical representation and the incoming multimedia stream to make an optimal selection of the receiving client device 210. In some embodiments, as the routing topology, at least one media server computer uses a Selective Forwarding Unit (SFU) topology, or a Traversal Using Relay NAT (TURN) topology, or a Spatial Analysis Media Server (SAMS) topology, or some other multimedia server routing topology.

[0056] As Figure 1 shown, the SFU is used for the WebRTC video conferencing standard, which generally supports browser-to-browser applications such as voice calls, video chats, P2P file sharing applications, without the need for a plugin to connect video communication endpoints. The SFU includes software programs that are used to route and forward video data packets in a video stream to multiple participant devices without performing intensive media processing operations such as decoding, re-encoding, receiving all the encoded media streams from the client devices, and then selectively forwarding the above streams to their respective participants for decoding and presentation.

[0057] A TURN topology applicable to multiple scenarios when at least one media server computer 202 cannot establish a connection between client devices 210 is an extension of the Session Traversal Utilities for NAT (STUN), where NAT stands for Network Address Translation. NAT is a method of remapping the Internet Protocol (IP) address space to another address space by modifying the network address information in the IP header of a data packet when the packet is transmitted through a traffic routing device. Thus, NAT can give a private IP address for accessing a network such as the Internet and enable a single device, such as a routing device, to act as a proxy between that Internet and a private network. NAT can be symmetric or asymmetric. A framework called Interactive Connectivity Establishment (ICE) can determine whether a symmetric or asymmetric NAT is needed, and this framework is used to find the best path for connecting client devices. A symmetric NAT is responsible not only for translating IP addresses from private to public or from public to private, but also for translating ports. On the other hand, an asymmetric NAT uses a STUN server to enable clients to discover their public IP addresses and the type of NAT behind them, which can be used to establish a connection. In many cases, STUN can be used only during connection establishment, and once the session is established, data can start flowing between client devices. TURN can be used in the case of a symmetric NAT, and TURN remains in the media path after connection establishment while relaying processed and / or unprocessed data between client devices.

[0058] Figure 3 FIG. shows a schematic diagram of a system 300 according to an embodiment, the system including at least one media server computer, which serves as SAMS 302. Figure 3 Some of the elements in Figure 2 refer to similar elements in

[0059] SAMS 302 is used to analyze and process the incoming data 304 of each client device. In some embodiments, the incoming data 304 may be related to user priorities, as well as the distance relationship between the corresponding user graphical representation and the multimedia stream in the virtual environment. The incoming data 302 includes one or more of the following data: metadata 306, priority data 308, data class 310, spatial structure data 312, a scene graph (not shown), three-dimensional data 314 including position, orientation, or motion data, user availability status data (e.g., active or passive status) 316, image data 318, media 320, and video data 322 based on a Scalable Video Codec (SVC), or a combination thereof. The SVC-based video data 322 enables the client device to send data that contains different resolutions without sending two or more streams, one for each resolution.

[0060] The incoming data 304 is sent by the client device. In the context of the application running on the client device executing in the virtual environment, the client device generates the incoming data 304, and SAMS 302 uses this incoming data 304 to perform incoming data operations 324 and data forwarding optimization 326. Thus, SAMS 302 can avoid storing information related to the virtual environment, the distance relationship between user graphical representations, the availability status, etc., because for the application running in the virtual environment, this data is already included in the incoming data sent by the client, thereby generating processing efficiency for SAMS 302. Because SAMS 302 can focus its resources only on data operations, routing, and forwarding before sending the multimedia stream to the client device, this efficiency can make SAMS 302 a viable and effective option for handling multi-user video conferences that include a large number (e.g., hundreds or thousands) of users accessing the virtual environment. Optionally, the incoming data 304 may include preprocessed spatial forwarding guidance in which the data forwarding operations have been performed. In this case, SAMS 302 can use only this guidance to send the multimedia stream to the client device without performing additional operations.

[0061] In some embodiments, the incoming data operations 326 implemented by at least one media server computer implementing SAMS may include compressing, encrypting, re-encrypting, decrypting, improving, mixing, enhancing, calculating, manipulating, encoding, or combining the incoming data. Depending on the priority and the spatial relationship (e.g., distance relationship) between the user graphical representation of the source having the multimedia stream and the remaining user graphical representations, these incoming data operations 326 may be performed in each client instance.

[0062] In some embodiments, data forwarding optimization 326 implemented by at least one media server computer implementing SAMS includes optimizing the bandwidth and computing resource utilization of one or more receiving client devices. In further embodiments, SAMS optimizes the forwarding of the outgoing data stream to each receiving client device by modifying, amplifying, or reducing the temporal characteristics, spatial characteristics, quality, and color characteristics of the media. Modifying, amplifying, or reducing the input data for temporal characteristics includes, for example, changing the frame rate; spatial characteristics may refer to, for example, image size; quality refers to, for example, different compression- or encoding-based qualities; color characteristics refer to, for example, color resolution and range. These operations can be performed based on the spatial, three-dimensional orientation, distance, and priority relationships of the specific receiving client user associated with the incoming data as described above, which helps optimize bandwidth and computing resources.

[0063] Priority data is related to, for example, speaker or listener status data, where one or more multimedia streams from a speaker have a higher priority score compared to the multimedia streams of the listeners. Spatial relationships include directly correlating the distance and orientation of the user graphical representation relative to the source of the virtual multimedia stream and the graphical representations of the remaining users. Thus, this spatial relationship correlates the multimedia stream that provides a higher resolution or a higher degree of enhancement for the user graphical representation and the multimedia stream that provides a lower resolution or a lower degree of enhancement for the user graphical representation, with the former being closer to and facing the source of the virtual multimedia stream and the latter being farther away from and partially facing or not facing the source of the virtual multimedia stream. Any combination between the two can also apply, for example, with any degree of facial and head orientation, any distance from the source of the multimedia stream, the user graphical representation partially faces the source of the multimedia stream, which has a direct impact on the quality of the multimedia stream received by the user.

[0064] The source of the multimedia stream can be, for example, a user speaking in a virtual video conference occurring in a virtual environment, a group of speakers participating in a debate or a conference in a virtual environment, a webinar, an entertainment event, a performance, etc., where at least one user is a speaker. In an example where a user is speaking (such as giving a speech, a webinar, a conference, etc.), multiple users can be located within the virtual environment listening to the speaker. Some users face, partially face, or do not face the speaker, which affects the priority of each user and thus the quality of the received multimedia stream. However, in other embodiments, the multimedia stream may not be from other user graphical representations but from other multimedia sources, such as virtual animations, augmented reality virtual objects, pre-recorded or live videos from an event or a location, application graphical representations, video games, etc., where data operations are performed based on the spatial, three-dimensional orientation, distance, and priority relationships of the specific receiving client device user associated with the multimedia stream.

[0065] Figure 4AA schematic diagram of a virtual environment 400 in an embodiment is shown, which adopts the SAMS topology of the present application.

[0066] The virtual environment 400 includes five user graphical representations 402, namely user graphical representations A - E, where user graphical representation A represents the speaker, and user graphical representations B - E represent four listeners, each of which is located at a different 3D coordinate position in the virtual environment 400, and they have different facial and head orientations, as well as different Points of View (PoV). Each user graphical representation 402 is associated with a corresponding user who interacts in the virtual environment 404 through a client device, and this client device is connected to the virtual environment 404 and at least one media server using the SAMS topology of the present application. This SAMS topology is, for example, Figure 3 SAMS 302 in

[0067] When user graphical representation A is speaking, user graphical representations B - E face each other user graphical representation with their respective orientations (e.g., the same or different orientations), corresponding PoV, and different 3D coordinates. For example, user graphical representation B is located closest to user graphical representation A and looks directly at user graphical representation A; user graphical representation C is located slightly farther from user graphical representation A than user graphical representation B and looks partially in the direction of user graphical representation A; user graphical representation D is located farthest from user graphical representation A and looks partially in the direction of user graphical representation A; and user graphical representation E is as close to user graphical representation A as user graphical representation B but looks in a different direction from user graphical representation A.

[0068] Therefore, SAMS captures incoming data from each of the five user graphical representations, performs input data operations and data forwarding optimizations, and selectively sends the resulting multimedia streams to the five client devices. Thus, each client device sends its own input data and receives a corresponding one or more multimedia streams (e.g., four multimedia streams, Figure 4A one in each of the other user graphical representations), where each received multimedia stream is adjusted (e.g., managed and optimized) for the corresponding user graphical representation based on the spatial, three - dimensional orientation, distance, and priority relationships of the specific receiving client device user associated with the multimedia stream to achieve optimal bandwidth and computing resources.

[0069] Figure 4B A schematic diagram of the forwarding of the adjusted media streams outgoing from five user graphical representations is shown, and these user graphical representations are located in Figure 4A the virtual environment 404 of Figure 4Athe multimedia stream of the client device 408 with a user graphical representation, and the SAMS 406 evaluates and optimizes the outgoing multimedia stream based on the incoming data received from multiple client devices 408, where the incoming data includes elements from Figure 4A at least one virtual environment 404. The data operations and optimizations can be performed by the data exchange management module 410 of the SAMS 406.

[0070] Each client device 408 sends its own input data to the SAMS 402 and accordingly receives four multimedia streams from other client devices 408. Thus, client device A, since it is being used by the speaker, sends an incoming media stream with high-priority data to the SAMS 402 and receives four lower-priority multimedia streams B - E from the corresponding four client devices 408.

[0071] Since user graphical representation B is closest to user graphical representation A and directly looks at user graphical representation A, client device B receives the multimedia stream from client A at the highest resolution compared to the rest of the users, receives the remaining multimedia streams C - E, and sends its own corresponding multimedia stream B. The resolution of each multimedia stream C - E received by client device B is based on its spatial relationship with user graphical representations C - E.

[0072] Since user graphical representation C is located slightly farther away from user graphical representation B and looks partially in the direction of user graphical representation A, client device C receives the multimedia stream from client device A at the third-highest resolution, receives the remaining multimedia streams B - E, and sends its own corresponding multimedia stream B. The resolution of each multimedia stream B - E received by client device C is based on its spatial relationship with user graphical representations B - E.

[0073] Since the corresponding user graphical representation D is located farthest from user graphical representation A and looks partially in the direction of user graphical representation A, client device D receives the multimedia stream from client device A at the lowest resolution, receives the remaining multimedia streams B, C, and E, and sends its own corresponding multimedia stream D. The resolution of each multimedia stream B, C, and E received by client device D is based on its spatial relationship with user graphical representations B, C, and E.

[0074] Since the user graphical representation E is as close to the user graphical representation A as the user graphical representation B and looks in a direction different from the user graphical representation A, the client device E receives the second highest resolution multimedia stream, receives the remaining multimedia streams B - D, and sends its corresponding multimedia stream E. The resolution of each multimedia stream B - D received by the client device E is based on its spatial relationship with the user graphical representations B - D. However, depending on the configuration of the SAMS 402, even though the user graphical representation E looks slightly away from the user graphical representation A, the client device E can also receive the multimedia stream with the same resolution as the client device B. This is because even though the user graphical representation E does not directly look at the user graphical representation A, the user graphical representation E may suddenly turn its perspective in the virtual world 404 to look at the user graphical representation A, and if there is no multimedia stream from the user graphical representation A, or if the resolution of the multimedia stream from the user graphical representation A is low despite being close to each other, then the quality of experience of the corresponding client device E may be disturbing or not optimal. Therefore, in this embodiment, it may be more effective to consider that the client device E and the client device B receive multimedia streams with the same resolution.

[0075] As can be understood from the specification, for the spatial, three - dimensional orientation, distance, and priority relationships of the multimedia streams corresponding to other user graphical representations with the user graphical representation, each user graphical representation 402 also manages and optimizes the individual multimedia streams received from the other four user graphical representations 402 separately. Therefore, if the individual multimedia streams are related to the corresponding client devices 408, each of the client devices 408 receives separate multimedia streams from the other four client devices.

[0076] In some embodiments, if the user graphical representation 402 is too far from the multimedia source, such as Figure 4A User A as described, the SAMS 406 can extract the multimedia streams received by the user graphical representation 402 from multiple media sources. If the SAMS 406 is configured accordingly, this can apply to the user graphical representation D.

[0077] In some embodiments, the multimedia stream may not be from other user graphical representations but from other multimedia sources, such as virtual animations, augmented reality virtual objects, pre - recorded or live videos from events or locations, application graphical representations, video games, etc. In these embodiments, the multimedia stream is still adjustable and can be managed and optimized separately, for example, for the spatial, three - dimensional orientation, distance, and priority relationships between the multimedia stream and the source of the multimedia stream and the corresponding user graphical representation. In other embodiments, the multimedia stream comes from a combination of user graphical representations and other multimedia sources.

[0078] Figures 5A - 5BFIG. 0 shows a schematic diagram of a SAMS combining media streams of multiple client devices. In embodiments where SAMS can be used to combine media streams of multiple client devices, SAMS combines the streams in the form of a mosaic. The mosaic can be a virtual frame that includes individual virtual tiles, where each multimedia stream of the user graphical representation is streamed. The mosaic can be adjusted per client based on the distance relationship between the user graphical representation of the client device and the source of the multimedia stream and the remaining user graphical representations.

[0079] In Figure 5A FIG. 5, seven user graphical representations 502 interact within a virtual environment 504, where user graphical representation A is the speaker and the remaining user graphical representations B - G are listening to user graphical representation A. User graphical representations G and F are relatively close to user graphical representation A; user graphical representations E and F are located at a relatively far position from user graphical representation A; and user graphical representations D and C are located at the farthest position from user graphical representation A.

[0080] Figure 5B FIG. 9 shows a combined multimedia stream in the form of a mosaic 506 that includes individual virtual tiles 508, namely virtual tiles A - F, where each individual multimedia stream is streamed from the corresponding user graphical representation 502 in the virtual environment 504 from the perspective of user graphical representation G. From the perspective of user graphical representation G, since user graphical representation A is the speaker and its outgoing media stream has a higher priority, virtual tile A is larger and has the highest resolution; since user graphical representation G and user graphical representation F are also close, virtual tile F is the second largest and has the second highest resolution; user graphical representations B - E are equally small and have the same or similar relatively low resolution among themselves. In some embodiments, SAMS 406 sends the same mosaic 506 to all client devices, and these client devices continue by excising the unnecessary tiles, which are not relevant to the position and orientation of the client device in the virtual environment.

[0081] Figure 6 FIG. 13 shows a block diagram of the topology of a spatial awareness multimedia routing method 600 according to an embodiment. Method 600 can be implemented in a system such as the systems disclosed in Figures 2 - 3 systems 200 and 300.

[0082] Method 600 starts at step 602, where in the memory of at least one media server computer, data and instructions are provided that implement a client device data exchange management module, which manages data exchange between multiple client devices.

[0083] Then, in step 604, method 600 continues, and at least one media server computer receives incoming data, which includes at least multimedia streams from multiple client devices. The incoming data is associated with user priority data and the spatial relationship between the corresponding user graphical representation and the incoming multimedia streams.

[0084] In step 606, method 600 continues, and the data exchange management module performs client device data exchange management. The data exchange management may include analyzing and optimizing the incoming data from multiple client devices, where the incoming data includes graphical elements within a virtual environment; and evaluating and optimizing the forwarding of the outgoing multimedia streams based on the incoming data received from multiple client devices.

[0085] Finally, in step 608, method 600 may continue, and based on the data exchange management, the corresponding multimedia streams are forwarded to the client devices, where the media streams are displayed to the user graphical representation of at least one client device.

[0086] In some embodiments, method 600 further includes: when forwarding the outgoing multimedia streams, using a routing topology, or a media processing topology, or a forwarding server topology, or other suitable multimedia server routing topologies, or media processing and forwarding server topologies, or other suitable server topologies.

[0087] In a further embodiment, as the routing topology, at least one media server computer uses a Selective Forwarding Unit (SFU) topology, a Traversal Using Relay NAT (TURN) topology, a Spatial Analysis Media Server (SAMS) topology, or a multimedia server routing topology.

[0088] In a further embodiment, as the media processing topology, at least one media server computer is used to perform one or more operations on the incoming data, including: compressing, encrypting, re - encrypting, decrypting, decoding, combining, improving, mixing, enhancing, calculating, manipulating, or encoding, or a combination thereof.

[0089] In a further embodiment, when using the forwarding server topology, method 600 further includes using one or more of a MCU, a cloud media mixer, and a cloud 3D renderer.

[0090] In some embodiments, as the SAMS, at least one media server computer is used to analyze and process the incoming data of each client device, and the incoming data is associated with user priorities and the distance relationship between the corresponding user graphical representation and the incoming multimedia stream. The incoming data includes metadata, priority data, data classes, spatial structure data, three-dimensional position, orientation or motion information, status data of speakers or listeners, availability status data, images, media, scalable video codec-based video, or combinations thereof. In some embodiments, the forwarding of the optimized outgoing multimedia stream implemented by at least one media server computer implementing the SAMS includes optimizing the bandwidth and computing resource utilization of one or more receiving client devices.

[0091] In some embodiments, the SAMS optimizes the forwarding of the outgoing data stream to each receiving client device by modifying, amplifying, or reducing the temporal, spatial, quality, and color characteristics of the media.

[0092] A computer-readable medium storing instructions for causing one or more computers to perform any of the methods described herein is also described herein. The term "computer-readable medium" as used herein includes volatile, non-volatile, removable, and non-removable media implemented in any method or technology capable of storing information, such as computer-readable instructions, data structures, program modules, or other data. Generally, the functions of the computing devices described herein can be implemented in computing logic embodied in hardware or software instructions, which can be written in programming languages such as C, C++, COBOL, JAVA TM 、PHP, Perl, Python, Ruby, HTML, CSS, JavaScript, VBScript, ASPX, Microsoft.NET such as C# TM language, and / or similar languages. The computing logic can be compiled into an executable program or written in an interpreted programming language. Generally, the functions described herein can be implemented as logical modules, which can be replicated to provide greater processing power, combined with other modules, or divided into sub-modules. The computing logic can be stored on any type of computer-readable medium (e.g., non-transitory media such as memory or storage media) or computer storage device and stored on and executed by one or more general-purpose or special-purpose processors. In this way, a special-purpose computing device is created to provide the functions described herein.

[0093] While certain embodiments have been described and illustrated in the accompanying drawings, it is to be understood that these embodiments are merely illustrative of the broad invention and not limiting thereof, and that the invention is not limited to the specific constructions and configurations shown and described, since various other modifications may occur to those skilled in the art. Accordingly, this embodiment is to be regarded as illustrative rather than restrictive.

Claims

1. A multimedia router system, characterized in that, Comprising: At least one media server computer, the at least one media server computer including at least one processor and a memory, wherein the at least one media server computer is configured as a Spatial Analysis Media Server (SAMS) and receives and analyzes incoming data, the incoming data including incoming multimedia streams, user priority data, and spatial orientation data from client devices associated with respective users; and Based on the incoming data received from the client devices, adjusting the outgoing multimedia streams for the respective client devices, wherein the incoming multimedia streams include elements from at least one virtual environment, and wherein the outgoing multimedia streams are adjusted for the respective client devices based on the user priority data and the spatial orientation data, the spatial orientation data describing the spatial relationship between the corresponding user graphical representation of the user in the at least one virtual environment and the source of the incoming multimedia streams in the at least one virtual environment.

2. The system according to claim 1, wherein The at least one virtual environment is hosted on at least one dedicated server computer, the at least one dedicated server computer being connected to the at least one media server computer via a network, or the at least one virtual environment is hosted in a peer-to-peer infrastructure and relayed through the at least one media server computer.

3. The system according to claim 1, wherein The at least one virtual environment includes a virtual copy of a real-world location, wherein the real-world location includes a plurality of sensors that provide further data to the virtual copy of the real-world location.

4. The system according to claim 1, wherein The at least one media server computer is further configured to combine the incoming data in the form of a mosaic, the mosaic including individual tiles, wherein respective multimedia streams of the user graphical representation are streamed.

5. The system according to claim 1, characterized in that, The at least one media server computer is configured as a Multipoint Control Unit (MCU), or a cloud media mixer, or a cloud 3D renderer.

6. The system according to claim 1, characterized in that The at least one media server computer is configured to analyze and process the incoming data of each client device and determine the user priority and the spatial relationship between the corresponding user graphical representation and the source of the incoming multimedia streams.

7. The system according to claim 1, characterized in that, Adjusting the outgoing multimedia streams includes optimizing the bandwidth and computing resource utilization of the one or more receiving client devices.

8. The system according to claim 1, characterized in that, Adjusting the outgoing multimedia streams includes adjusting temporal characteristics, spatial characteristics, quality, or color characteristics, or a combination thereof.

9. A multimedia routing method, characterized in that, Comprising: At least one media server computer configured as a Spatial Analysis Media Server (SAMS) receives incoming data, the incoming data including incoming multimedia streams, user priority data, and spatial orientation data from a plurality of client devices associated with respective users, the spatial orientation data describing the spatial relationship between the corresponding user graphical representation of the user in at least one virtual environment and the source of the incoming multimedia streams in the at least one virtual environment; Analyzing the incoming data from the plurality of client devices, the incoming data including graphical elements within the virtual environment; Adjust the outgoing multimedia stream based on the user priority data and the spatial orientation data of the incoming data received from the multiple client devices; Forward the adjusted outgoing multimedia stream to one or more receiving client devices, wherein the adjusted outgoing multimedia stream is configured to be displayed to a user of the one or more receiving client devices.

10. The method according to claim 9, characterized in that, Further include using one or more of a multi-point control unit (MCU), a cloud media mixer, or a cloud 3D renderer.

11. The method according to claim 9, wherein Further include combining the incoming data in the form of a mosaic, the mosaic including individual tiles, wherein respective multimedia streams of user graphical representations are streamed.

12. The method according to claim 9, characterized in that, Further include determining user priorities and the spatial relationship between the corresponding user graphical representations and the sources of the incoming multimedia streams.

13. The method according to claim 9, characterized in that, Adjusting the outgoing multimedia stream includes optimizing the bandwidth and computing resource utilization of the one or more receiving client devices.

14. The method according to claim 9, wherein Adjusting the outgoing multimedia stream includes adjusting temporal characteristics, spatial characteristics, quality, or color characteristics, or a combination thereof.

15. A computer-readable medium, characterized in that, Instructions are stored on the computer-readable medium, the instructions being configured to cause at least one media server computer including a processor and a memory to perform the following steps: The at least one media server computer configured as a spatial analysis media server (SAMS) receives incoming data, the incoming data including incoming multimedia streams, user priority data, and spatial orientation data from multiple client devices associated with corresponding users, the spatial orientation data describing the spatial relationship between one or more user graphical representations of the user in at least one virtual environment and at least one element in the at least one virtual environment; Analyze the incoming data from the multiple client devices; Adjust the outgoing multimedia stream based on the user priority data and the spatial orientation data of the incoming data received from the multiple client devices; Forward the adjusted outgoing multimedia stream to a receiving client device, wherein the adjusted outgoing multimedia stream is configured to be displayed on the receiving client device.

16. The computer-readable medium according to claim 15, wherein The steps further include determining user priorities and the spatial relationship based on the incoming data.

17. The computer-readable medium according to claim 15, wherein Adjusting the outgoing multimedia stream includes optimizing the bandwidth and computing resource utilization of the one or more receiving client devices.

18. The computer-readable medium according to claim 15, wherein Adjusting the outgoing multimedia stream includes adjusting temporal characteristics, spatial characteristics, quality, or color characteristics, or a combination thereof.

Citation Information

Patent Citations

  • Augmented reality computing environments - mobile device join and load

    US20190310757A1