Spatially aware multimedia router system and method
Patent Information
- Application Number
- KR1020210113899
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-08-28
- Filing Date
- 2021-08-27
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-08-27
Smart Images

Figure 112021099322642-PAT00002_ABST
Abstract
Description
Technology Field
[0001] (Cross-reference to related applications)
[0002] This application claims the benefits of concurrently filed U.S. patent application No. XX / XXX,ZZZ under the heading “System and method for enabling interaction with a virtual presence in a virtual environment,” which is incorporated herein by reference.
[0003] (Technology field)
[0004] The present disclosure generally relates to computer systems, and more specifically to multimedia router systems and methods. Background Technology
[0005] Video conferencing enables remote communication among multiple users and is becoming a relatively inexpensive and fast communication tool for people in multiple locations. Video conferencing has recently gained popularity due to the widespread deployment of broadband networks, advancements in video compression technology, and the increased availability of approaches to implement web-based video communication with lower infrastructure requirements and lower costs.
[0006] For example, one such approach that enables video conferencing is a mesh (peer-to-peer) infrastructure, where each client device sends multimedia streams to all other client devices. This does not require intermediate infrastructure but results in a low-cost solution with low scalability due to the rapid bandwidth overload and limited processing capacity of client devices.
[0007] Another exemplary approach is a Multipoint Control Unit (MCU) implemented in a central media server, which receives all multimedia streams from client devices, decodes and re-encodes them, and combines them into a single stream transmitted to all client devices, thereby reducing latency and bandwidth issues compared to the P2P model. However, MCU implementations tend to be complex and require a significant amount of computing resources from the media server.
[0008] Another exemplary approach is the Selective Forwarding Unit (SFU) used in the WebRTC (Web Real-Time Communication) video conferencing standard. The WebRTC standard supports cross-browser applications, such as voice calls, video chat, and peer-to-peer (P2P) file sharing applications, while avoiding the need for plugins to connect video communication endpoints. An SFU, which can be implemented on a central media server computer, includes a software program configured to route video packets of a video stream to multiple participant devices without performing intensive media processing (e.g., decoding and re-encoding) on the server. Thus, the SFU receives all encoded media streams from client devices over the network and then selectively forwards the stream to each participant's client device for subsequent decoding and display. The selectivity of the SFU's forwarding can be based on multiple parameters that can be used to optimize bandwidth associated with the forwarding of multimedia streams, resulting in a higher Quality of Experience (QoE). For example, the SFU can identify the speaking participant in the received multimedia stream and forward the high-bit-rate multimedia stream to the listening participant. On the other hand, the SFU can transmit the listening participant's low-bit-rate multimedia stream to other participants to achieve some degree of bandwidth and QoE improvement. The problem to be solved
[0009] One limitation of general video conferencing tools, such as those utilizing the aforementioned approach, is their limited scalability, which takes into account the bandwidth and processing capacity limitations of the central media or routing server or any of the participating client devices. Therefore, a new approach is needed that can further optimize network bandwidth and computing resources during multimedia routing and forwarding operations while maintaining a high QoE for the relevant participants. means of solving the problem
[0010] This summary is provided to introduce the selection of concepts in a simplified form, which is further explained in the detailed description below. This summary is not intended to identify the key features of the claimed subject matter, nor is it intended to aid in determining the scope of the claimed subject matter.
[0011] In one aspect of the present disclosure, a spatially aware multimedia router system is provided. The spatially aware multimedia router system includes at least one media server computer comprising at least one processor and a memory storing instructions that implement a data exchange management module for managing data exchange between client devices. In one embodiment, the system further includes one or more computing devices that implement at least one virtual environment connected to at least one media server computer, which enables access to one or more graphic representations (also referred to as user graphic representations) of users of a plurality of client devices. A plurality of multimedia streams (e.g., 2D video streams, 3D video streams, audio streams, or combinations of such streams or other media streams) are generated from the virtual environment while considering virtual elements within the virtual environment and input data from at least one client device. Accordingly, the input data is received and combined within the virtual environment, which includes a plurality of virtual elements and at least one graphic representation of a user corresponding to the client device. The plurality of client devices are connected to the at least one media server computer via a network and configured to transmit data including multimedia streams to the at least one media server computer.
[0012] The at least one media server is configured to receive and analyze incoming data including an incoming multimedia stream from a client device, and to coordinate an outbound multimedia stream for an individual client device based on the incoming data. The incoming multimedia stream includes an element within at least one virtual environment. The outbound multimedia stream is coordinated to the individual client device based, for example, on spatial orientation data and user priority data describing the spatial relationship between a corresponding user graphic representation within the at least one virtual environment and the source of the incoming multimedia stream.
[0013] In one embodiment, the at least one media server computer performs data exchange management including analyzing and processing incoming data having a multimedia stream from the client device, and evaluating and optimizing the forwarding of the outbound multimedia stream based on the incoming data received from the plurality of client devices having elements from the at least one virtual environment. The incoming data is associated with user priority data and the spatial relationship between the corresponding user graphic representation and the incoming multimedia stream.
[0014] In some embodiments, the at least one virtual environment is hosted on at least one dedicated server computer connected to the at least one media server computer via a network. In other embodiments, the at least one virtual environment is hosted on a peer-to-peer infrastructure and relayed through the at least one media server computer. The virtual environment may be used to host real-time video communication in which users can interact with each other, and may be used, among other things, for meetings, work, education, shopping, entertainment, and services. In some embodiments, the virtual environment is a virtual replica of a real-world location, and the real-world location includes a plurality of sensors that provide additional data to the virtual replica of the real-world location.
[0015] In some embodiments, the at least one media server computer uses a routing topology. In other embodiments, the at least one media server computer uses a media processing topology. In other embodiments, the at least one media server computer uses a forwarding server topology. In other embodiments, the at least one media server computer uses another suitable multimedia server routing topology, or a media processing and forwarding server topology, or another suitable server topology.
[0016] In some embodiments where the at least one media server computer uses a routing topology, the at least one media server computer uses a Selective Forwarding Unit (SFU) topology, or a Traversal Using Relay NAT (TURN), a spatially analyzed media server topology (SAMS), or some other multimedia server routing topology.
[0017] In an embodiment in which the at least one media server computer uses a media processing topology, the at least one media server computer is configured to perform one or more operations on incoming data, including compression, encryption, re-encryption, decryption, decoding, combination, improvement, mixing, enhancement, augmentation, computing, manipulation, or encoding, or a combination thereof. In a further embodiment, combining the incoming data is performed in the form of a mosaic comprising individual tiles on which individual multimedia streams of user graphic representations are streamed.
[0018] In an embodiment where the at least one media server computer uses a forwarding server topology, the at least one media server computer is configured as a multipoint control unit (MCU), a cloud media mixer, or a cloud 3D renderer.
[0019] In an embodiment of the at least one media server computer configured as a SAMS, the at least one media server computer is configured to analyze and process the incoming data of each client device related to user priority and spatial relationships (e.g., spatial relationships between a corresponding user graphic representation and the source of the incoming multimedia stream). In such an embodiment, the at least one media server computer may be further configured to determine user priority and / or spatial relationships based on such data. In a specific embodiment, the incoming data comprises one or more of: metadata, priority data, data classes, spatial structure data, scene graphs, 3D position, orientation or locomotion information, speaker or listener state data, availability state data, image data, video based on a scalable video codec, or a combination thereof. In a further embodiment, coordinating the outbound multimedia stream (e.g., as implemented by the at least one media server computer implementing the SAMS) includes optimizing bandwidth and computing resource utilization for the one or more receiving client devices. The adjustment of the outbound multimedia stream may also include adjusting temporal characteristics, spatial characteristics, quality, or color characteristics, or a combination thereof. In another embodiment, the SAMS optimizes the forwarding of the outbound data stream to each receiving client device by modifying, upscaling, or downscaling the media with respect to temporal characteristics, spatial characteristics, quality, and color characteristics.
[0020] In another aspect of the present disclosure, a spatially aware multimedia router method is provided. The spatially aware multimedia router method comprises the step of providing a command to implement a client device data exchange management module that manages data and data exchange between a plurality of client devices in the memory of at least one media server computer. The method proceeds by receiving incoming data, which includes incoming multimedia streams from the plurality of client devices, by the at least one media server computer, wherein the incoming data is associated with user priority data and spatial orientation data. For example, the spatial orientation data may describe, for example, a spatial relationship between a corresponding user graphic representation and one or more sources of the incoming multimedia stream. The method then continues by performing data exchange management by the data exchange management module. In one embodiment, the data exchange management comprises the step of analyzing and / or processing the incoming data from the plurality of client devices, which includes graphic elements within the virtual environment; and the step of coordinating the outbound multimedia stream based on the incoming data received from the plurality of client devices. The above method is terminated by forwarding the adjusted outbound multimedia stream to one or more receiving client devices, wherein the adjusted outbound multimedia stream is configured to be displayed on the receiving client device(s) (e.g., to a user displayed as a user graphic representation).
[0021] In some embodiments, the method further includes the step of using a routing topology, or a media processing topology, or a forwarding server topology, or another suitable multimedia server routing topology, or a media processing and forwarding server topology, or other suitable server topology when forwarding the outbound multimedia stream.
[0022] In some embodiments, in the routing topology, the at least one media server computer uses an SFU (Selective Forwarding Unit) topology, TURN (Traversal Using Relay NAT), SAMS (spatially Analyzed Media Server Topology), or other multimedia server routing topology.
[0023] In some embodiments, in a media processing topology, the at least one media server computer is configured to perform one or more media processing operations on the incoming data, including compression, encryption, re-encryption, decryption, decoding, combination, improvement, mixing, enhancement, augmentation, computing, manipulation, or encoding, or a combination thereof.
[0024] In some embodiments, the method further includes the step of utilizing one or more of a Multipoint Control Unit (MCU), a cloud media mixer, and a cloud 3D renderer when using a forwarding server topology.
[0025] In some embodiments, using a SAMS configuration, the method further comprises the step of analyzing and processing the incoming data of each client device related to user priority and spatial relationships (e.g., distance relationships or other spatial relationships between the corresponding user graphic representation and the source of the incoming multimedia stream). In such embodiments, the method may further comprise the step of determining user priority and / or spatial relationships based on such data. The incoming data includes one or more of: metadata, priority data, data classes, spatial structure data, scene graphs, 3D location, orientation or movement information, speaker or listener status data, availability status data, image data and scalable video codec-based video, or a combination thereof.
[0026] In some embodiments, adjusting the outbound multimedia stream (e.g., implemented by the at least one media server computer implementing SAMS) includes the steps of optimizing bandwidth and calculating resource utilization for the one or more receiving client devices. Adjusting the outbound multimedia stream may also include the step of adjusting temporal features, spatial features, quality, or color features, or a combination thereof. In additional embodiments, the SAMS optimizes the step of forwarding the outbound data stream to each receiving client device by modifying, upscaling, or downscaling the media with respect to temporal features, spatial features, quality, and color features.
[0027] In another aspect of the present disclosure, a computer-readable medium stores thereon instructions configured to enable one or more computing devices to perform any of the techniques described herein. In one embodiment, at least one computer-readable medium stores instructions configured to enable at least one media server computer, having a processor and memory, to perform: receiving incoming data including incoming multimedia streams from a plurality of client devices by the at least one media server computer, wherein the incoming data is associated with spatial orientation data describing a spatial relationship between one or more user graphic representations and at least one element of at least one virtual environment; analyzing the incoming data from the plurality of client devices; adjusting the outbound multimedia stream based on the incoming data received from the plurality of client devices; and forwarding the adjusted outbound multimedia stream to a receiving client device, wherein the adjusted outbound multimedia stream is configured to be displayed on the receiving client device.
[0028] The foregoing summary does not constitute a complete list of all embodiments of the present disclosure. The present disclosure is to be considered to include all systems and methods that can be practiced from all suitable combinations of the various embodiments summarized above, as well as those disclosed in the detailed description below and, in particular, those indicated in the claims filed with the application. Such combinations have special advantages not specifically mentioned in the foregoing summary. Other features and advantages of the present disclosure will become apparent from the accompanying drawings and the detailed description below. Brief explanation of the drawing
[0029] Specific features, aspects, and advantages of the present disclosure will be better understood in conjunction with the following description and the accompanying drawings. Figure 1 illustrates a schematic diagram of a conventional Selective Forwarding Unit (SFU) routing topology. FIG. 2 illustrates a schematic diagram of a spatial recognition multimedia router system according to one embodiment. FIG. 3 illustrates a schematic diagram of a system comprising at least one media server computer configured as a SAMS according to one embodiment. Figure 4a illustrates a schematic diagram of a virtual environment. FIG. 4b illustrates a schematic diagram of the forwarding of an outbound media stream from a speaking user within a virtual environment using the SAMS topology of the present disclosure, according to one embodiment. FIGS. 5a-5b illustrates a schematic diagram of a usage scenario in which SAMS combines media streams from a plurality of client devices according to one embodiment. FIG. 6 illustrates a block diagram of a spatial recognition multimedia router method of the present disclosure according to one embodiment. Specific details for implementing the invention
[0030] In the following description, reference is made to drawings illustrating various embodiments. Additionally, various embodiments are described below with reference to a number of examples. It should be understood that the embodiments may include changes in design and structure without departing from the scope of the claimed subject matter.
[0031] The present disclosure provides a spatially aware multimedia router system and method configured to receive input data from a plurality of client devices and to implement data exchange management for said input data. The input data is received and combined within a virtual environment comprising a plurality of virtual elements and at least one graphic representation of a corresponding user of said client device. The virtual environment may be used to host real-time video communication in which users can interact with each other, and may be used, among other things, for meetings, work, education, shopping, entertainment, and services. The data exchange management comprises the steps of analyzing and processing incoming data from said client device comprising at least a multimedia stream (e.g., a 2D video stream, a 3D video stream, an audio stream, or a combination of such streams or other media streams), and optimizing the forwarding and evaluation of said outbound multimedia streams based on the incoming data received from said client device comprising elements from said at least one virtual environment. The incoming data is associated with user priority data and the spatial relationship between said corresponding user graphic representation and said incoming multimedia stream. Accordingly, the system and method of the present invention can optimize the forwarding of input data and outbound multimedia streams so that routing of the multimedia stream received by the client device occurs within the virtual environment while simultaneously performing an optimal selection of the receiving client device, thereby enabling simultaneous efficiency of bandwidth and computing resources. Such efficiency discloses the spatial awareness multimedia router system and method of the present invention as a viable and effective option for processing multi-user video conferencing involving a large number (e.g., hundreds or thousands) of users accessing the virtual environment.
[0032] FIG. 1 illustrates a schematic diagram of a conventional SFU (Selective Forwarding Unit) routing topology (100).
[0033] An exemplary conventional SFU routing topology (100) includes at least one media server computer (102) having at least one processor (104) and a memory (106) for storing a computer program that implements an SFU (108) for providing video forwarding in a real-time video application. A plurality of client devices (110) can communicate in real-time through the SFU (108), wherein the SFU (108) delivers an outbound media stream containing a Real-time Transmission Protocol (RTP) video packet to the client devices (110) based on one or more parameters. For example, client device B may indicate the current speaker while communicating in real-time with client device A. Client device B transmits two or more media streams to client device A, wherein, for example, one media stream is transmitted at high resolution (B) (112) and one media stream is transmitted at low resolution (B) (114). Additionally, in this example, client device A may also transmit two or more media streams to client device (B), such as one media stream at high resolution (A) (116) and one media stream at low resolution (A) (118). Client device B may receive the low resolution (A) (118) media stream from client device A because client device A may be used by the only passive user listening at that moment for the user of client device B. Meanwhile, client device A may receive the high resolution (B) (112) media stream from client device B because the user of client device B may be an active user currently speaking.At least one media server computer (102) and two or more client devices (110) can be connected via one or more wired or wireless communication networks (e.g., a Local Area Network (LAN), a Wide Area Network (WAN), the Internet, paths, links, and any intermediate node network hardware such as routers, gateways, firewalls, switches, etc.). Video applications can utilize the WebRTC standard to implement, for example, multi-party video conferencing applications.
[0034] The SFU routing topology (100) may include limitations such as forwarding each media stream of each client device (110) regardless of the state of the client device (110); not enabling media modifications such as data operations on the media streams (e.g., augmentation, enhancement, combination, etc.); and considering limited parameters to prioritize and optimize outbound media streams, thereby bringing about suboptimal network bandwidth and resource optimization for the system and ultimately limiting the number of users who can simultaneously interact with and view multimedia streams. Finally, the existing SFU routing topology (100) is not optimal for enabling interactions such as video communication and social interaction in a virtual environment.
[0035] FIG. 2 illustrates a schematic diagram of a spatial recognition multimedia router system (200) according to one embodiment.
[0036] A spatial awareness multimedia router system (200) includes at least one media server computer (202) comprising at least one processor (204) and a memory (206) that stores instructions for implementing a data exchange management module (208) that manages data exchange between client devices (210) connected to at least one media server computer via a network (212). The spatial awareness multimedia router system (200) may further include at least one virtual environment server (220) connected to at least one media server computer (202) that enables access to one or more user graphic representations of users (216) of a plurality of client devices (210) in a virtual environment (214). A plurality of multimedia streams are generated in the virtual environment (214), including live feed data from one or more users (216) and multimedia streams acquired by a camera (218) that acquires graphic elements within the virtual environment.
[0037] 적어도 하나의 미디어 서버 컴퓨터(202)는 복수의 클라이언트 장치(210)로부터 수신된 입력 데이터를 저장, 처리, 라우팅 및 포워딩하는 것과 같은 본 명세서에 개시된 기술을 수행하기 위한 리소스(예를 들어, 네트워크 액세스 능력과 함께 적어도 하나의 프로세서(204) 및 메모리(206))를 포함하는 서버 컴퓨팅 장치이다. 적어도 하나의 미디어 서버 컴퓨터(202)는 가상 환경(214) 내로부터의 그래픽 엘리먼트를 포함하는 클라이언트 장치(210)로부터의 멀티미디어 스트림을 포함하는 인입 데이터를 분석 및 처리하는 단계, 및 아웃바운드 멀티미디어 스트림을 조정하는 단계를 포함하는 데이터 교환 관리 모듈(208)을 통해 클라이언트 장치 데이터 교환 관리를 수행한다. 일 실시 예에서, 이것은 복수의 클라이언트 장치(210)로부터 수신된 인입 데이터에 기초하여 아웃바운드 멀티미디어 스트림들의 포워딩을 평가하고 최적화하는 단계를 포함한다. 아웃바운드 멀티미디어 스트림들은 인입 데이터, 예를 들어 적어도 하나의 가상 환경 내에서 인입 멀티미디어 스트림의 소스 및 대응하는 사용자 그래픽 표현 사이의 공간 관계를 설명하는 예를 들어 사용자 우선순위 데이터 및 공간 방향 데이터에 기초하여 개벌 클라이언트 장치에 대해 조정된다. In one embodiment, the input data is associated with a spatial relationship between one or more user graphic representations and at least one element of a virtual environment (214).
[0038] At least one virtual environment (214) is a virtual replica of a real-world location in some embodiments. The real-world location may include a plurality of sensors that provide real-world data to the virtual replica of the real-world location through the virtual environment (214). The sensors may send the captured data to a virtual environment server (220) via a network (212), which may be used by at least one media server (202) to update, enhance, and synchronize the corresponding virtual replica of the real-world element in at least one virtual environment. Additionally, one or more media servers (202) may be further configured to augment real-world data into the virtual environment by merging the real-world data captured by the sensors with the virtual data of the virtual environment (214) into the virtual environment (214).
[0039] In this disclosure, the term “enrichment” is used to describe the act of providing additional attributes to a virtual replica based on multi-source data. Enriching a virtual replica may be considered a special form of updating the virtual replica with one or more new forms of data that may not have previously existed in the virtual replica. For example, enriching a virtual replica may mean providing real-world data captured from sensing mechanisms in multiple devices. The additional real-world data may include, for example, video data, temperature data, real-time energy consumption data, real-time water consumption data, speed or acceleration data, etc.
[0040] In one embodiment, at least one virtual environment (214) is hosted on at least one virtual environment server computer (220) comprising at least one processor (222) and memory (224) implementing the virtual environment (214). In another embodiment, at least one virtual environment (214) is hosted on a Peer-to-Peer (P2P) infrastructure and relayed to a plurality of client devices (210) through at least one media server computer (202).
[0041] An array of at least one virtual environment (214) may be associated with one or more themes, such as for meetings (e.g., virtual conference rooms), work (e.g., virtual office spaces), education (e.g., virtual classrooms), shopping (e.g., virtual shops), entertainment (e.g., karaoke, event halls or stadiums, theaters, nightclubs, sports arenas or stadiums, museums, cruise ships, video games, etc.), and services (e.g., hotels, travel agencies or restaurant reservations or orders, government agency services, etc.). A combination of virtual environments (214) from the same and / or different themes may form a virtual environment cluster, which may include hundreds or thousands of virtual environments (e.g., multiple virtual classrooms may be part of a virtual school). A virtual environment (214) may be a 2D or 3D virtual environment that includes a physical layout and visual appearance associated with the theme of the virtual environment (214), which may be customized by the user according to the user's preferences or needs. The user can access the virtual environment (214) through a graphic representation that can be inserted into the virtual environment (214) and can be graphically combined with the two-dimensional or three-dimensional virtual environment (214).
[0042] Each virtual environment (214) may be provided with corresponding resources (e.g., memory, network, and computing power) by at least one virtual environment server computer (220) or P2P infrastructure. At least one virtual environment (214) may be accessed by one or more users (216) through a graphical user interface via a client device (210). For example, a graphical user interface that utilizes WebRTC standards and provides application data and commands necessary to run a selected virtual environment (214) and enable multiple interactions within it may be included in a downloadable client application or web browser application. Furthermore, each virtual environment (214) may include one or more human or artificial intelligence (AI) hosts or assistants capable of assisting the user within the virtual environment (214) by providing necessary data and / or services through a corresponding user graphical representation. For example, a human or AI bank service clerk may assist the virtual bank user by providing necessary information in the form of presentations, forms, lists, etc., as requested by the user.
[0043] In some embodiments, the user graphic representation may be a user 3D virtual cutout that can be constructed from one or more input images, such as a user upload or a third-party source photo with the background removed; or a user real-time 3D virtual cutout with the background removed that can be generated based on input data, such as real-time 2D, stereo or depth images or video data, or 3D video data from a live video stream data feed acquired from a camera including a user's real-time video stream, video without background removal, or video with the background removed. In some embodiments, the user graphic representation may be rendered and displayed using a polygonal structure. Such a polygonal structure may be a quad structure or a more complex 3D structure used as a virtual frame to support the video. In another embodiment, one or more user graphic representations are inserted into three-dimensional coordinates within a virtual environment (214) and graphically combined therein.
[0044] In this disclosure, the term “user 3D virtual cutout” refers to a virtual replica of a user constructed from a user-uploaded or third-party source 2D photo. The user 3D virtual cutout is generated through a 3D virtual reconstruction process using machine vision technology with the user-uploaded or third-party source 2D photo as input data to produce a 3D mesh or 3D point cloud of the user with the background removed. In this disclosure, the term “user real-time 3D virtual cutout” refers to a virtual replica of a user after the user background has been removed based on a real-time 2D or 3D live video stream data feed acquired from a camera. The user real-time 3D virtual cutout is generated through a 3D virtual reconstruction process using machine vision technology with the user live data feed as input data to produce a 3D mesh or 3D point cloud of the user with the background removed. In the present disclosure, the term “background-removed video” refers to a video streamed to a client device, on which a background removal process is performed so that only the user can view it, and which is then displayed on the receiving client device using a polygonal structure of the screen. In the present disclosure, the term “background-unremoved video” refers to a video streamed to a client device, wherein the video faithfully represents the camera capture so that the user and their background can be seen, and which is then displayed on the receiving client device using a polygonal structure.
[0045] The P2P infrastructure can enable real-time interaction and its synchronization by using a suitable P2P communication protocol that enables real-time communication between client devices (210) in a virtual environment (214) through a suitable API (application programming interface). An example of a suitable P2P communication protocol may be the WebRTC communication protocol, which is a collection of standards, protocols, and JavaScript APIs, which together enable P2P audio, video, and data sharing between peer client devices (210). Client devices (210) using the P2P infrastructure may perform real-time 3D rendering of a live session, for example, using one or more rendering engines. An exemplary rendering engine may be a WebGL-based 3D engine, which is a JavaScript API for rendering 2D and 3D graphics within any compatible web browser without using plugins, which allows for accelerated use of physics and image processing and effects by one or more processors of at least one client device (210) (e.g., one or more graphics processing units (GPUs)). Additionally, a client device (210) using a P2P infrastructure can perform image and video processing and machine learning computer vision techniques through one or more suitable computer vision libraries. An example of a suitable computer vision library is OpenCV, which is a library of programming functions configured primarily for real-time computer vision tasks.
[0046] In some embodiments, at least one media server computer (202) uses a routing topology. In other embodiments, at least one media server computer (202) uses a media processing topology. In other embodiments, at least one media server computer (202) uses a forwarding server topology. In other embodiments, at least one media server computer (202) uses another suitable multimedia server routing topology, or a media processing and forwarding server topology, or another suitable server topology. The topology used by at least one media server computer (202) may depend on the processing capacity of the client device and / or at least one media server computer, as well as the capacity of the network infrastructure being utilized.
[0047] In some embodiments, in a media processing topology, at least one media server computer (202) is configured to perform one or more media processing operations on incoming data, including compression, encryption, re-encryption, decryption, decoding, combination, improvement, mixing, enhancement, augmentation, computing, manipulation, encoding, or a combination thereof. Accordingly, in a media processing topology, at least one media server computer (202) is configured to perform a plurality of media processing operations, including routing and forwarding incoming data, as well as extending and otherwise modulating outbound multimedia streams to client devices (210).
[0048] In some embodiments, in a forwarding server topology, at least one media server computer (202) is configured as a multipoint control unit (MCU), a cloud media mixer, or a cloud 3D renderer. As an MCU, at least one media server computer (202) is configured to receive all multimedia streams from client devices, decode and re-encode all streams, and combine all streams into a single stream to be transmitted to all client devices (210). As a cloud media mixer, at least one media server computer (202) is configured to select from various client devices (210) and other multimedia sources (e.g., audio and video), such as at least one virtual environment (214), and to mix the input data multimedia streams with footage and / or special effort to generate an output multimedia stream processed for the client devices (210). Visual effects can vary, for example, from simple mixing and wiping to sophisticated effects. As a cloud 3D renderer, at least one media server computer (202) is configured to compute a 3D scene from a virtual environment through multiple computer computations to generate a final animated multimedia stream that is transmitted back to a client device (210).
[0049] In a routing topology, at least one media server computer (202) is configured to determine where to send a multimedia stream (e.g., one or more client devices (210)) that can be performed via an Internet Protocol (IP) routing table in order to select the interface that is the best path. Such routing decisions in the routing table are based on rules stored in a data exchange management module (208) that take into account priority data and the spatial relationship between the corresponding user graphic representation and the incoming multimedia stream so that the optimal selection of the receiving client device (210) is performed. In some embodiments, as a routing topology, at least one media server computer uses a Selective Forwarding Unit (SFU) topology, a Traversal Using Relay NAT (TURN) topology, or a Spatially Analyzed Media Server Topology (SAMS), or some other multimedia server routing topology.
[0050] The SFU described with reference to FIG. 1 is used in the WebRTC video conferencing standard, which generally supports cross-browser applications such as voice calls, video chat, and P2P file sharing applications, while eliminating the need for plugins to connect video communication endpoints. The SFU includes a software program configured to route and forward video packets of a video stream to multiple participant devices without performing intensive media processing tasks, such as decoding and re-encoding, receiving all encoded media streams from client devices, and then selectively forwarding the stream to each participant for decoding and display.
[0051] A TURN topology suitable for situations where at least one media server computer (202) cannot establish a connection between client devices (210) is an extension of STUN (Session Traversal Utilities for NAT), where NAT represents Network Address Translation. NAT is a method of remapping an Internet Protocol (IP) address space to another address space by modifying network address information in the IP header of a packet while the packet is being transmitted through a traffic routing device. Thus, NAT enables private IP addresses to access networks such as the Internet and allows a single device, such as a routing device, to act as an agent between the Internet and the private network. NAT can be symmetric or asymmetric. A framework called ICE (Interactive Connectivity Establishment), configured to find the optimal path to connect client devices, can determine whether symmetric or asymmetric NAT is required. Symmetric NAT performs the task of converting IP addresses from private to public or vice versa, as well as converting ports. On the other hand, asymmetric NAT uses a STUN server to allow clients to discover public IP addresses and the NAT type behind them, which can be used to establish a connection. In many cases, STUN can be used only during connection setup, and once the session is established, data can begin to flow between client devices. TURN can be used in the case of symmetric NAT and may remain in the media path after the connection is established while processed and / or unprocessed data is relayed between client devices.
[0052] FIG. 3 illustrates a schematic diagram of a system (300) comprising at least one media server computer configured as a SAMS (302) according to one embodiment. Some elements of FIG. 3 may refer to elements similar to FIG. 2, and thus the same reference numbers may be used.
[0053] SAMS (302) is configured to analyze and process incoming data (304) from each client device. Incoming data (304) may, in some embodiments, relate to user priority and distance relationships between the corresponding user graphic representation in a virtual environment and the multimedia stream. Incoming data (302) includes one or more of metadata (306), priority data (308), data classes (310), spatial structure data (312), scene graphs (not shown), three-dimensional data (314) including location, orientation, or movement data, user availability status data (e.g., active or passive status) (316), image data (318), media (320), and scalable video codec (SVC)-based video data (322), or a combination thereof. SVC-based video data (322) may enable the client device to transmit data containing different resolutions without the need to transmit two or more streams, each stream per resolution.
[0054] Incoming data (304) is transmitted by a client device, which generates incoming data (304) used by SAMS (302) to perform incoming data operations (324) and data forwarding optimization (326) in the context of an application running on a client device running in a virtual environment. Therefore, SAMS (302) does not need to store information regarding the distance relationship between the virtual environment and user graphic representations, availability status, etc., because the incoming data transmitted by the client already contains data regarding the application running in the virtual environment and generating processing efficiency for SAMS (302). Due to this efficiency, SAMS (302) can concentrate resources solely on data operations, routing, and forwarding optimization before transmitting the multimedia stream to the client device, thus making SAMS a viable and effective choice for processing multi-user video conferencing involving a large number (e.g., hundreds or thousands) of users accessing the virtual environment. Alternatively, the incoming data (304) may include pre-processed spatial forwarding instructions for which data forwarding calculations have already been performed. In this situation, SAMS (302) can simply use these instructions to transmit a multimedia stream to a client device without performing additional calculations.
[0055] In some embodiments, the incoming data operation (326) implemented by at least one media server computer implementing SAMS may include compression, encryption, re-encryption, decryption, improvement, mixing, enhancement, augmentation, computing, manipulation, encoding, or a combination thereof for the incoming data. This incoming data operation (326) may be performed for each client instance according to the priority and spatial relationship (e.g., distance relationship) of the user graphic representation with respect to the source of the multimedia stream and the rest of the user graphic representation.
[0056] In some embodiments, data forwarding optimization (326) implemented by at least one media server computer implementing SAMS includes optimizing bandwidth and calculating resource utilization for one or more receiving client devices. In additional embodiments, SAMS optimizes the forwarding of outbound data streams to each receiving client device by modifying, upscaling, or downscaling the media regarding temporal features, spatial features, quality, and color features. Such modification, upscaling, or downscaling of incoming data regarding temporal features may include, for example, changing the frame rate; spatial features may refer, for example, to image size; quality may refer, for example, to different compression or encoding-based quality; and color may refer, for example, to color resolution and range. These operations are performed based on the spatial, three-dimensional orientation, distance from such incoming data, and priority relationships of a specific receiving client device user and may contribute to the optimization of bandwidth and computing resources.
[0057] Priority data is associated, for example, with speaker or listener state data, and one or more multimedia streams from the speaker have higher priority scores than multimedia streams from the listener. Spatial relationships involve associating a direct correlation between the orientation and distance of a user graphic representation to the source of a virtual multimedia stream and the rest of the user graphic representations. Thus, spatial relationships are associated with providing higher resolution or enhanced multimedia streams to user graphic representations that are closer to and facing the source of the virtual multimedia stream, and providing lower resolution or less enhanced multimedia streams to user graphic representations that are further away from the source of the multimedia stream and partially facing or not facing it. Any combination between these may apply, such as user graphic representations that partially face the source of the multimedia stream at any facing angle and head orientation and at any distance from the multimedia stream source, and which directly affect the quality of the multimedia stream received by the user.
[0058] The source of the multimedia stream may be, for example, a user speaking in a virtual video conference occurring within a virtual environment, a panel of speakers participating in a discussion or meeting within a virtual environment, a webinar, an entertainment event, a show, etc., where at least one user is the speaker. In the example of what the user is speaking (e.g., a speech, a webinar, a meeting, etc.), multiple users may be located within the virtual environment listening to the speaker. Some users may be face-to-face, partially face-to-face, or not face-to-face with the speaker, which affects the priority of each user and the quality of the received multimedia stream accordingly. However, in other embodiments, the multimedia stream does not come from other user graphic representations but from virtual animations, augmented reality virtual objects, pre-recorded or live video of events or places, application graphic representations, video games, etc., where data operation is performed based on the spatial, three-dimensional orientation, distance, and priority relationships of a specific receiving client device user to these multimedia streams.
[0059] FIG. 4a illustrates a schematic diagram of a virtual environment (400) in which the SAMS topology of the present disclosure can be used, according to one embodiment.
[0060] The virtual environment (400) includes five user graphic representations (402), namely user graphic representation AE, wherein user graphic representation A represents a speaker and user graphic representation BE represents four listeners, each located at a different 3D coordinate position in the virtual environment (400) and having a different face and head direction along with a different perspective (PoV). Each user graphic representation (402) is associated with a corresponding user interacting in the virtual environment (404) through a client device connected to at least one media server computer using the SAMS topology of the present disclosure, such as the SAMS (302) disclosed with reference to FIG. 3.
[0061] While User Graphic Representation A is speaking, User Graphic Representation BE may face each other at different 3D coordinates, individual directions (e.g., the same or different directions), and corresponding PoVs. For example, User Graphic Representation B is located closest to User Graphic Representation A and is looking directly at User Graphic Representation A. User Graphic Representation C is located slightly further away than User Graphic Representation B and is looking partially in the direction of User Graphic Representation A. User Graphic Representation D is located furthest from User Graphic Representation A and may be looking partially in the direction of User Graphic Representation A. User Graphic Representation E is as close to User Graphic Representation A as User Graphic Representation B but may be looking in a different direction from User Graphic Representation A.
[0062] Accordingly, SAMS captures incoming data from each of the five user graphic representations, performs input data processing and data forwarding optimization, and selectively sends the resulting multimedia streams to five client devices. Accordingly, each client device transmits its own input data and receives one or more multimedia streams (e.g., four multimedia streams, one of each user graphic representation in FIG. 4a) in correspondence, wherein each received multimedia stream is individually adjusted (e.g., managed and optimized) for the corresponding user graphic representation based on the spatial, three-dimensional orientation, distance, and priority relationships for these incoming data of the specific receiving client device user to achieve optimal bandwidth and computing resources.
[0063] FIG. 4b illustrates a schematic diagram of the forwarding of outbound media streams adapted from five user graphic representations of the virtual environment (404) of FIG. 4a using the SAMS (406) topology of the present disclosure according to one embodiment. The SAMS (406) can analyze and optimize incoming data including multimedia streams from a client device (408) corresponding to the user graphic representation of FIG. 4a, and evaluate and optimize the forwarding of outbound multimedia streams based on incoming data received from a plurality of client devices (408) including elements from at least one virtual environment (404) of FIG. 4a. The data operation and optimization can be performed by the data exchange management module (410) of the SAMS (406).
[0064] Each client device (408) sends its own input data to SAMS (402) and accordingly receives four multimedia streams from other client devices (408). Thus, client device (A) sends an incoming media stream with high priority data to SAMS (402) because it is used by a speaker, and receives four low priority multimedia streams (BE) from the corresponding four client devices (408).
[0065] Since user graphic representation B is located closest to user graphic representation A and is directly viewing user graphic representation A, client device B receives the multimedia stream from client A at the highest resolution compared to the rest of the users, receives the remaining multimedia streams CE at their respective resolutions based on their spatial relationship with user graphic representation CE, and transmits its corresponding multimedia stream B.
[0066] Client device C receives the third highest resolution from client device A because user graphic representation C is slightly further away than user graphic representation B and partially faces the direction of user graphic representation A, receives the remaining multimedia stream BE at their respective resolutions based on their spatial relationship with user graphic representation BE, and sends its corresponding multimedia stream C.
[0067] Client device D receives the lowest resolution from client device A because the corresponding user graphic representation D is furthest from user graphic representation A and partially faces the direction of user graphic representation A, receives the remaining multimedia streams B, C, and E at their respective resolutions based on their spatial relationships with user graphic representations B, C, and E, and transmits its corresponding multimedia stream D.
[0068] Client device E receives the second highest resolution because User Graphic Representation E is as close to User Graphic Representation A as User Graphic Representation B but may be looking in a different direction from User Graphic Representation A, receives the remaining multimedia stream BD at their respective resolutions based on their spatial relationship with User Graphic Representation BD, and transmits its corresponding multimedia stream E. However, depending on the configuration of SAMS (402), Client device E may also receive the same resolution as Client device B even though User Graphic Representation E is looking slightly away from User Graphic Representation A. This is because User Graphic Representation E may suddenly turn its direction within the virtual world (404) to look at User Graphic Representation A even if it is not directly looking at User Graphic Representation A, and if there is no multimedia stream from User Graphic Representation A or if the multimedia stream from User Graphic Representation A is at a low resolution despite being close to each other, the experience quality of the corresponding Client device E may be disruptive or not optimal. Therefore, in this embodiment, it may be more efficient to consider the same resolution as that received by Client device B.
[0069] As can be seen from the description, the multimedia streams received by each of the remaining four user graphic representations (402) are also individually managed and optimized for the spatial, three-dimensional orientation, distance, and priority relationships of the corresponding user graphic representations to the multimedia streams of other user graphic representations. Accordingly, each client device (408) receives individual multimedia streams from the remaining four client devices when the multimedia stream is associated with the corresponding client device (408).
[0070] In some embodiments, if the user graphic representation (402) is located too far from the multimedia source as in the example of FIG. 4a, the user graphic representation (402) from user A may be spared by SAMS (406) from receiving the multimedia stream from the multimedia source. This may apply, for example, to the user graphic representation D if SAMS (406) is configured accordingly.
[0071] In some embodiments, the multimedia stream may not come from another user graphic representation, but may come from other multimedia sources such as virtual animations, augmented reality virtual objects, pre-recorded or live video of events or places, application graphic representations, video games, etc. In these embodiments, the multimedia stream is still adjustable and may be individually managed and optimized, for example, with respect to the spatial, three-dimensional orientation, distance, and priority relationships of the corresponding user graphic representation to the multimedia stream of the received multimedia stream source. In other embodiments, the multimedia stream comes from a combination of a user graphic representation and other multimedia sources.
[0072] FIGS. 5a-5b illustrate a schematic diagram of a SAMS combining media streams from a plurality of client devices according to one embodiment. In an embodiment in which the SAMS may be configured to combine media streams from a plurality of client devices, the SAMS may combine the streams in a mosaic form. The mosaic may be a virtual frame comprising individual virtual tiles on which individual multimedia streams of user graphic representations are streamed. The mosaic may be adjusted per client device according to the distance relationship between the multimedia stream source and the user graphic representation of the client device and the remaining user graphic representations.
[0073] In FIG. 5a, seven user graphic representations (502) are interacting within a virtual environment (504), where user graphic representation A is the speaker and the remaining user graphic representations BG are listening to user graphic representation A. User graphic representations G and F are found to be relatively close to user graphic representation A; user graphic representations E and F may be located relatively further away from user graphic representation A; and user graphic representations D and C may be located furthest away from user graphic representation A.
[0074] FIG. 5b illustrates a combined multimedia stream in the form of a mosaic (506) containing individual virtual tiles (508) and virtual tiles (AF) that stream each individual multimedia stream from a corresponding user graphic representation (502) of a virtual environment (504) in terms of a user graphic representation G. In terms of a user graphic representation G, virtual tile A may be a larger tile provided at the highest resolution because user graphic representation A has a higher priority for its outbound media stream as it is the speaker; virtual tile F may be a second largest tile provided at the second highest resolution because user graphic representation G is very close to user graphic representation F; and user graphic representation BE may be equally small and may have a relatively lower resolution that is the same or similar to each other. In some embodiments, SAMS (406) may send the same mosaic (506) to all client devices, and the client devices may proceed by cutting out unnecessary tiles that may not be related to their position and orientation in the virtual environment.
[0075] FIG. 6 illustrates a block diagram of the topology of the spatial awareness multimedia router method (600) of the present disclosure according to one embodiment. The method (600) may be implemented in a system such as that disclosed with reference to the systems (200 and 300) of FIG. 2 and 3.
[0076] The method (600) begins in step (602) by providing a command to memory that implements a client device data exchange management module that manages data exchange between at least one media server computer data and a plurality of client devices.
[0077] Then, in step (604), the method (600) proceeds by receiving incoming data containing at least a multimedia stream from a plurality of client devices by at least one media server computer. The incoming data is associated with user priority data and a spatial relationship between the corresponding user graphic representation and the incoming multimedia stream.
[0078] In step (606), the method (600) proceeds by performing client device data exchange management by a data exchange management module. Data exchange management includes the steps of: analyzing and processing incoming data from a plurality of client devices including graphic elements within a virtual environment; and evaluating and optimizing the forwarding of outbound multimedia streams based on incoming data received from the plurality of client devices.
[0079] Finally, in step (608), the method (600) can proceed by forwarding a corresponding multimedia stream to a client device based on data exchange management, wherein the multimedia stream is displayed on a user graphic representation of a user of at least one client device.
[0080] In some embodiments, the method (600) further includes the step of using a routing topology, or a media processing topology, or a forwarding server topology, or other suitable multimedia server routing topology, or a media processing and forwarding server topology, or other suitable server topology when forwarding an outbound multimedia stream.
[0081] In an additional embodiment, as a routing topology, at least one media server computer uses a Selective Forwarding Unit (SFU) topology, a Traversal Using Relay NAT (TURN), a Spatially Analyzed Media Server Topology (SAMS), or a multimedia server routing topology.
[0082] In an additional embodiment, as a media processing topology, at least one media server computer is configured to perform one or more operations on incoming data including compression, encryption, re-encryption, decryption, decoding, combining, improvement, mixing, enhancement, augmentation, computing, manipulation, encoding, or a combination thereof.
[0083] In an additional embodiment, when using a forwarding server topology, the method (600) further includes the step of using one or more of an MCU, a cloud media mixer, and a cloud 3D renderer.
[0084] In some embodiments, as a SAMS, at least one media server computer is configured to analyze and process incoming data from each client device related to the distance relationship between the corresponding user graphic representation and the multimedia stream and user priority. The incoming data includes one or more of metadata, priority data, data classes, spatial structure data, three-dimensional location, orientation or movement information, speaker or listener state data, availability state data, images, media, and video based on an extensible video codec, or a combination thereof. In some embodiments, optimizing the forwarding of an outbound multimedia stream implemented by at least one media server computer implementing the SAMS includes the steps of optimizing bandwidth and calculating resource utilization for one or more receiving client devices.
[0085] In some embodiments, SAMS optimizes forwarding outbound data streams to each receiving client device by modifying, upscaling, or downscaling the media with respect to temporal, spatial, quality, and color characteristics.
[0086] A computer-readable medium storing instructions configured to cause one or more computers to perform any method described herein is also described. As used herein, the term “computer-readable medium” includes volatile and non-volatile, and removable and non-removable media implemented by any method or technique capable of storing information such as computer-readable instructions, data structures, program modules, or other data. Generally, the function of the computing device described herein may be implemented as computing logic implemented as hardware or software instructions that may be written in programming languages such as C, C++, COBOL, JAVA™, PHP, Perl, Python, Ruby, HTML, CSS, JavaScript, VBScript, ASPX, C#, and Microsoft .NET™ languages and / or others. The computing logic may be compiled into an executable program or written in an interpreted programming language. Generally, the function described herein may be implemented as a logic module that may be duplicated, merged with other modules, or divided into submodules to provide greater processing power. Computing logic may be stored in any type of computer-readable medium (e.g., non-transient media such as memory or storage media) or computer storage device, and may be stored in and executed by one or more general-purpose or dedicated processors to create a dedicated computing device configured to provide the functions described herein.
[0087] Although specific embodiments have been described and illustrated in the accompanying drawings, such embodiments are merely illustrative and do not limit the extensive disclosure. It should be understood by those skilled in the art that the present disclosure is not limited to the specific configurations and arrangements illustrated and described, as various other variations may occur. Accordingly, the description should be regarded as illustrative rather than restrictive.
Claims
Claim 1 A multimedia router system comprising at least one media server computer having at least one processor and memory, wherein the at least one media server computer is configured to receive and analyze incoming data including an incoming multimedia stream from a client device and to coordinate an outbound multimedia stream for an individual client device based on the incoming data received from the client device, wherein the incoming multimedia stream includes an element from at least one virtual environment, and the outbound multimedia stream is coordinated for the individual client device based on spatial orientation data and user priority data describing the spatial relationship between the source of the incoming multimedia stream and the corresponding user graphic representation within the at least one virtual environment, and wherein the spatial orientation data serving as the basis for the coordination of the outbound multimedia stream includes the degree to which the corresponding user graphic representation faces the source of the incoming multimedia stream. Claim 2 A multimedia router system according to claim 1, wherein the at least one virtual environment is hosted on at least one dedicated server computer connected to the at least one media server computer via a network, or is hosted on a peer-to-peer infrastructure and relayed through the at least one media server computer. Claim 3 A multimedia router system according to claim 1, wherein at least one virtual environment includes a virtual replica of a real-world location, and the real-world location includes a plurality of sensors that provide additional data to the virtual replica of the real-world location. Claim 4 A multimedia router system according to claim 1, characterized in that at least one media server computer uses an SFU (Selective Forwarding Unit) topology, a TURN (Traversal Using Relay NAT) topology, or a SAMS (Spatially Analyzed Media Server Topology). Claim 5 A multimedia router system according to claim 1, wherein the at least one media server computer is further configured to combine the incoming data in a mosaic form comprising separate tiles on which individual multimedia streams of user graphic representations are streamed. Claim 6 A multimedia router system according to claim 1, characterized in that at least one media server computer is configured as a multipoint control unit (MCU), a cloud media mixer, or a cloud 3D renderer. Claim 7 delete Claim 8 A multimedia router system according to claim 1, wherein the step of adjusting the outbound multimedia stream comprises a step of optimizing bandwidth and a step of calculating resource utilization for one or more receiving client devices. Claim 9 A multimedia router system according to claim 1, wherein the step of adjusting the outbound multimedia stream includes the step of adjusting temporal features, spatial features, quality, or color features, or a combination thereof. Claim 10 A multimedia routing method comprising: a step of receiving incoming data comprising incoming multimedia streams from a plurality of client devices by at least one media server computer, wherein the incoming data is associated with spatial direction data and user priority data describing the spatial relationship between a corresponding user graphic representation and a source of the incoming multimedia stream within a virtual environment; a step of analyzing the incoming data from the plurality of client devices including a graphic element from the virtual environment; a step of adjusting an outbound multimedia stream based on the incoming data and spatial direction data received from the plurality of client devices, wherein the spatial direction data forming the basis for adjusting the outbound multimedia stream includes the degree to which the corresponding user graphic representation faces the source of the incoming multimedia stream; and a step of forwarding the adjusted outbound multimedia stream to one or more receiving client devices, wherein the adjusted outbound multimedia stream is configured to be displayed to a user of the one or more receiving client devices. Claim 11 A multimedia routing method according to claim 10, characterized in that at least one media server computer uses an SFU (Selective Forwarding Unit) topology, a TURN (Traversal Using Relay NAT) topology, or a SAMS (Spatially Analyzed Media Server Topology). Claim 12 A multimedia routing method according to claim 10, further comprising the step of utilizing one or more of a multipoint control unit (MCU), a cloud media mixer, or a cloud 3D renderer. Claim 13 A multimedia routing method according to claim 10, further comprising the step of combining the incoming data into a mosaic form including separate tiles in which individual multimedia streams of user graphic representations are streamed. Claim 14 delete Claim 15 A multimedia routing method according to claim 10, wherein the step of coordinating the outbound multimedia stream comprises a step of optimizing bandwidth and a step of calculating resource utilization for one or more receiving client devices. Claim 16 A multimedia routing method according to claim 10, wherein the step of adjusting the outbound multimedia stream includes the step of adjusting temporal features, spatial features, quality, or color features, or a combination thereof. Claim 17 A computer-readable medium characterized by storing instructions configured to enable at least one media server computer, comprising a processor and memory, to perform steps including: receiving incoming data comprising an incoming multimedia stream from a plurality of client devices, wherein the incoming data is associated with spatial orientation data describing a spatial relationship between one or more user graphic representations and at least one element of at least one virtual environment; analyzing the incoming data from the plurality of client devices; adjusting an outbound multimedia stream based on the incoming data and spatial orientation data received from the plurality of client devices, wherein the spatial orientation data forming the basis for adjusting the outbound multimedia stream includes the degree to which the corresponding user graphic representation faces the at least one element of the at least one virtual environment; and forwarding the adjusted outbound multimedia stream to a receiving client device, wherein the adjusted outbound multimedia stream is configured to be displayed on the receiving client device. Claim 18 A computer-readable medium according to claim 17, wherein the steps further include a step of determining user priority and spatial relationship based on the input data. Claim 19 A computer-readable medium according to claim 17, characterized in that the step of coordinating the outbound multimedia stream comprises the step of optimizing bandwidth and the step of calculating resource utilization for one or more receiving client devices. Claim 20 A computer-readable medium according to claim 17, wherein the step of adjusting the outbound multimedia stream comprises adjusting temporal features, spatial features, quality, or color features, or a combination thereof. Claim 21 A computer-readable medium according to claim 17, wherein the spatial orientation data forming the basis for the coordination of the outbound multimedia stream further includes the distance between the corresponding user graphic representation and the at least one element of the at least one virtual environment. Claim 22 A multimedia router system according to claim 1, characterized in that the user priority data forming the basis for the coordination of the outbound multimedia stream includes an indication of whether the corresponding user graphic representation is associated with a speaker or with a listener. Claim 23 A multimedia router system according to claim 1, wherein the spatial direction data forming the basis for the coordination of the outbound multimedia stream further includes the distance between the source of the incoming multimedia stream and the corresponding user graphic representation within the at least one virtual environment. Claim 24 A multimedia routing method according to claim 10, wherein the coordination of the outbound multimedia stream is further based on user priority data, and the user priority data forming the basis of the coordination of the outbound multimedia stream includes an indication of whether the corresponding user graphic representation is associated with a speaker or with a listener. Claim 25 A multimedia routing method characterized in that, in claim 10, the spatial direction data forming the basis for the coordination of the outbound multimedia stream further includes the distance between the source of the incoming multimedia stream and the corresponding user graphic representation within the at least one virtual environment.
Citation Information
Patent Citations
Shared virtual area communication environment based apparatus and methods
US20090254843A1
Method and apparatus for multi-experience metadata translation of media content with metadata
US20130024756A1
Methodology for negotiating video camera and display capabilities in a multi-camera / multi-display video conferencing environment
US20150092008A1
System and method for multi-user digital interactive experience
US20180352303A1
Augmented reality computing environments - mobile device join and load
US20190310757A1