A system and method for virtual broadcasting from within a virtual environment
By inserting user graphical representations into cloud server computers and utilizing image processing technology, the problem of insufficient user presence and interactive experience in virtual environments is solved, achieving realistic virtual environment interaction and supporting a variety of application scenarios.
Patent Information
- Application Number
- CN202110987554.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-01
- Filing Date
- 2021-08-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-08-26
AI Technical Summary
Existing remote interaction technologies lack user presence and realism, resulting in poor user interaction experiences in virtual environments and requiring expensive equipment and infrastructure support.
By inserting user graphical representations into cloud server computers and leveraging image processing technology to enable real-time multi-user collaboration and interaction in a virtual environment, and by using existing computing devices and cameras for background removal and graphic composition, a realistic and interactive experience is provided.
Without requiring expensive equipment, it enhances the user's sense of presence and interactive experience in the virtual environment, enabling realistic social interaction and collaboration, and supporting a variety of virtual environment applications.
Smart Images

Figure CN114205362B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Patent Application No. 17 / 006,327, filed August 28, 2020, and U.S. Patent Application No. 17 / 060,516, filed October 1, 2020, which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to virtual and / or mixed reality. More specifically, this disclosure relates to the presentation, transmission, and display of multimedia streams in virtual and / or mixed reality environments. Background Technology
[0004] As situations such as the 2020 COVID-19 pandemic forced global restrictions on the movement of people, changing the way we meet, learn, shop, and work, remote collaboration and interaction, especially social interaction, have become increasingly important. A wide variety of solutions are available on the market to enable real-time communication and collaboration, from chat applications to video calls such as Skype™ and Zoom™, or virtual offices for remote teams represented by 2D avatars, such as the virtual offices offered by Pragli™.
[0005] Given the current state of development of wearable immersive technologies such as extended reality (e.g., augmented and / or virtual reality) and their relatively low technology footprint, it's understandable that most solutions offer a flat, two-dimensional user interface on which most interactions occur. However, when comparing real-life experiences to these solutions, the low level of realism, lack of user presence, lack of shared space, and quality of interaction can lead many users to feel lonely or bored, which in turn can sometimes result in lower productivity than performing the same activities in person.
[0006] What is needed is a technological solution that, when interacting remotely, does not require purchasing expensive equipment (e.g., in head-mounted displays) or implementing new or costly infrastructure, but provides users with a sense of realism, their own presence as well as that of the participants, and an interactive experience as if in real life, all while utilizing existing computing devices and cameras. Summary of the Invention
[0007] This summary is provided to present, in a simplified form, the selection of concepts that will be further described in the following detailed description. This summary is not intended to identify key features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter.
[0008] This disclosure generally relates to computer systems, and more specifically to a system and method for realizing interaction, particularly social interaction, in a virtual environment; a system and method for virtual presence based on image processing; a user authentication system and method based on user graphical representations; a system and method for virtual broadcasting from within a virtual environment; a system and method for delivering applications within a virtual environment; a system and method for providing cloud computing-based virtual computing resources within a cloud server computer in a virtual environment; and a system and method for enabling self-organizing virtual communication between proximate user graphical representations.
[0009] The system disclosed herein enables interaction in a virtual environment, particularly including social interaction. The system comprises one or more cloud server computers, each containing at least one processor and memory storing data and instructions for implementing a virtual environment platform comprising at least one virtual environment. The one or more cloud server computers are configured to insert user graphical representations generated from live data feeds obtained from cameras at three-dimensional coordinate positions in at least one virtual environment, update user graphical representations in at least one virtual environment, and enable real-time multi-user collaboration and interaction within the virtual environment.
[0010] In an embodiment, the system further includes at least one camera that receives live data feeds from one or more users of a client device. Additionally, the system includes client devices communicatively connected to one or more cloud server computers and at least one camera. The system generates a user graphical representation from the live data feed, which is interpolated into three-dimensional coordinates of the virtual environment and updated therein using the live data feed. In the described embodiment, interpolating the user graphical representation into the virtual environment involves graphically combining the user graphical representation in the virtual environment such that the user graphical representation appears in the virtual environment (e.g., at a specified 3D coordinate location). The virtual environment platform serves the virtual environment to one or more client devices. The system enables real-time multi-user collaboration and (social) interaction in the virtual environment by accessing a graphical user interface via the client devices. Client or peer devices of this disclosure may include, for example, computers, headsets, mobile phones, glasses, transparent screens, tablet computers, and input devices that typically have a built-in camera, or input devices that can be connected to and receive data feeds from said camera.
[0011] In some embodiments, the virtual environment can be accessed by a client device via a downloadable client application or web browser application.
[0012] In some embodiments, the user graphical representation includes a user 3D virtual cutout with the background removed, or a user real-time 3D virtual cutout with the background removed, or a video with the background removed, or a video without the background removed. In some embodiments, the user graphical representation is a user 3D virtual cutout constructed from a user-uploaded or third-party source photo with the background removed, or a user real-time 3D virtual cutout with the background removed generated based on a feed of real-time 2D, stereo, depth data, or 3D live video stream data obtained from a camera, thus including the user's real-time video stream, or a video without the background removed, or a video with the background removed and displayed using a polygonal structure. Such a polygonal structure can be a quadrilateral structure or a more complex 3D structure, used as a virtual frame supporting the video. In other embodiments, one or more such user graphical representations are interpolated into three-dimensional coordinates within a virtual environment and graphically combined therein.
[0013] User 3D virtual clipping can include a virtual copy of the user constructed from a 2D photograph uploaded by the user or from a third-party source. In an embodiment, using a 2D photograph uploaded by the user or from a third-party source as input data, a user 3D virtual clipping is created via a 3D virtual reconstruction process using machine vision technology, resulting in a 3D mesh or 3D point cloud of the user with the background removed. User real-time 3D virtual clipping can include a virtual copy of the user based on a real-time 2D or 3D live video stream data feed obtained from a camera and after removing the user's background. In an embodiment, user real-time 3D virtual clipping is created by generating a 3D mesh or 3D point cloud of the user with the background removed, using a live user data feed as input data, and via a 3D virtual reconstruction process using machine vision technology. Video with background removed can include video streamed to a client device where the background removal process has been performed on the video so that only the user is visible, and then displayed using a polygonal structure on the receiving client device. Video without background removal can include video streamed to a client device where the video faithfully represents a camera capture, making the user and his or her background visible, and then displayed using a polygonal structure on the receiving client device.
[0014] In some embodiments, the data used as input data included in live data feeds and / or user-uploaded or third-party source 2D photographs includes 2D or 3D image data, 3D geometry, video data, media data, audio data, text data, haptic data, time data, 3D entities, 3D dynamic objects, metadata, priority data, security data, location data, lighting data, depth data, and infrared data, etc.
[0015] In some embodiments, the user graphical representation is associated with a top-down view, a third-person view, a first-person view, or a self-view. In embodiments, the user's perspective when accessing a virtual environment through a user graphical representation is a top-down view, a third-person view, a first-person view, a self-view, or a broadcast camera view. The self-view may include a user graphical representation seen by another user graphical representation, and optionally, a virtual background of the user graphical representation.
[0016] In other embodiments, the viewpoint is updated when the user manually navigates the virtual environment via a graphical user interface.
[0017] In other embodiments, the viewpoint is automatically established and updated using a virtual camera, wherein the viewpoint fed by live data is associated with the viewpoint of the user graphical representation and the virtual camera, and wherein the virtual camera is automatically updated by tracking and analyzing user eye and head tilt data or head rotation data, or a combination thereof. In embodiments, the viewpoint is automatically established and updated using one or more virtual cameras, which are virtually placed and aligned in front of the user graphical representation, for example, in front of a video without background removal, or a video with background removal, or a user 3D virtual cutout, or a user real-time 3D virtual cutout. In one embodiment, the one or more virtual cameras may point outwards from eye level. In another embodiment, two virtual cameras, one for each eye, point outwards from the level of both eyes. In another embodiment, the one or more virtual cameras may point outwards from the center of the head position in the user graphical representation. In another embodiment, the one or more virtual cameras may point outwards from the center of the user graphical representation. In another embodiment, the one or more virtual cameras may be placed in front of the user graphical representation, for example, at the head level of the user graphical representation, pointing towards the user graphical representation when in a self-view. The user's viewpoint captured by the camera is correlated with the user's graphical representation of the viewpoint and the associated virtual camera using computer vision, thereby manipulating the virtual camera. Furthermore, the virtual camera is automatically built and updated by tracking and analyzing user eye and head tilt data, or head rotation data, or combinations thereof.
[0018] In other embodiments, the self-perspective includes a clipping of the user graphical representation seen by removing the background (e.g., in a phone camera's "selfie mode"). Alternatively, the self-perspective includes a virtual background of the virtual environment behind the user graphical representation for understanding how other participants perceive themselves. When including a virtual background of the user graphical representation, the self-perspective can be set as an area around the user graphical representation that can be captured by a virtual camera and can produce a circle, square, rectangle, or any other suitable shape for constructing the self-perspective.
[0019] In some embodiments, updating the user's graphical representation within the virtual environment includes updating the user's status. In embodiments, available user statuses include absent, busy, available, offline, on a conference call, or in a meeting. User status can be manually updated via a graphical user interface. In other embodiments, user status is automatically updated via a connection to user calendar information that includes and synchronizes user status data. In other embodiments, user status is automatically updated by detecting usage of a specific program, such as a programming development environment, a 3D editor, or other productivity software that specifies a busy status, which can be synchronized based on the user status. In other embodiments, user status can be automatically updated based on data obtained from camera feeds using machine vision algorithms.
[0020] In some embodiments, interactions between users via corresponding user graphical representations, particularly social interactions, include: chatting; screen sharing; host options; telemetry; recording; voting; document sharing; sending emojis; agenda sharing and editing; virtual hugs; raising hands; handshakes; walking; content addition, including interactive applications or static or interactive 3D assets, animations, or 2D textures; meeting summary preparation; object movement; content projection; laser pointing; playing games; purchasing; participating in self-organized virtual communications; and participating in private or group conversations.
[0021] In some embodiments, the virtual environment is a persistent virtual environment stored in the persistent storage of one or more cloud server computers, or a temporary virtual environment stored in the temporary storage of one or more cloud server computers. In one embodiment, the virtual environment is a persistent virtual environment that records changes performed thereon, including customizations stored in the persistent storage of at least one cloud server computer assigned to the persistent virtual environment. In other embodiments, the virtual environment is a temporary virtual environment stored in the temporary storage of the cloud server.
[0022] In some embodiments, the arrangement of virtual environments is associated with a contextual theme that relates to one or more virtual environment verticals selected from the virtual environment platform. In embodiments, possible arrangements include those for education, meetings, work, shopping, services, social interaction, or entertainment, or combinations thereof. A complex of virtual environments within one or more verticals can represent a virtual environment cluster.
[0023] In other embodiments, the virtual environment cluster is one or more of the following: a virtual school containing at least a number of classrooms; or a virtual company containing at least a number of work areas and meeting rooms, some of which are shared as a common work or network space for members of different organizations; or an event facility containing at least one indoor or outdoor activity area that hosts live events including captures of live entertainment performers; or a virtual shopping mall containing at least a number of shops; or a virtual casino containing at least a number of gaming areas; or a virtual bank containing at least a number of service areas; or a virtual nightclub containing at least a number of VIP areas and / or party areas, including captures of live radio program host (DJ) performers; or a virtual karaoke entertainment facility containing multiple private or public karaoke rooms; or a virtual cruise ship containing multiple virtual areas inside the ship and areas outside the ship, the external areas including landscapes, islands, towns, and cities, allowing users to disembark from the virtual cruise ship to access those virtual areas; or an esports stadium or arena.
[0024] In other embodiments, the virtual environment further includes a virtual computer comprising virtual resources. In one embodiment, the virtual resources are from one or more cloud computing resources, which are accessed via client devices and assigned to the virtual computer resources via management tools.
[0025] In some embodiments, the virtual environment platform is configured to multicast or broadcast remote events to multiple instances of the virtual environment. For example, this can be done to accommodate a large number of users from around the world to experience the same multicast live event.
[0026] In some embodiments, a clickable link redirecting to a virtual environment is embedded in one or more third-party sources that include a third-party website, application, or video game.
[0027] In another aspect of this disclosure, a method for implementing interactions including social interactions in a virtual environment includes: providing a virtual environment platform containing at least one virtual environment in the memory of one or more cloud server computers including at least one processor; receiving live data feeds (e.g., live data feeds of users captured by at least one camera) from at least one corresponding client device; generating a user graphical representation from the live data feeds; inserting the user graphical representation into the three-dimensional coordinate position of the virtual environment; updating the user graphical representation within the virtual environment according to the live data feeds; and processing data generated by interactions in the virtual environment. Such interactions may particularly include social interactions within the virtual environment. For such interactions, the method may include enabling real-time multi-user collaboration and interaction in the virtual environment by directly communicating via peer-to-peer (P2P) or indirectly by using one or more cloud servers to serve the updated virtual environment to the client devices.
[0028] In some embodiments, the system (e.g., via a virtual environment platform) can also enable the creation of self-organizing virtual communication, which may include creating self-organizing voice communication channels between user graphical representations without changing the current viewpoint or position within the virtual environment. For example, a user graphical representation may approach another user graphical representation and participate in a self-organizing voice conversation at the location of the two user graphical representations' areas within the virtual environment. Such communication will be achieved, for example, by considering the distance, position, and orientation between the user graphical representations, and / or their current availability status (e.g., available or unavailable) or the state configuration of such self-organizing communication, or a combination thereof. The approaching user graphical representation will see visual feedback on the other user graphical representation, signaling that self-organizing communication is possible, and thus setting the start of a conversation between the two user graphical representations, where the approaching user can speak and the other user can hear and respond. In another embodiment, the virtual environment platform enables participation in self-organizing virtual communication within a virtual environment by processing data generated in response to steps performed by a client device, which may include the steps of: approaching a user graphical representation; selecting and clicking on a user graphical representation; sending or receiving an invitation to participate in self-organizing virtual communication from another user graphical representation; and accepting the received invitation. In such scenarios, the platform can open communication channels between user client devices, where user graphical representations maintain dialogue in a virtual space within a virtual environment.
[0029] In some embodiments, the method further includes engaging one or more users in a dialogue, transforming a user graphical representation from a user 3D virtual clip to a user real-time 3D virtual clip, or a video with or without background removal; and opening a peer-to-peer (P2P) communication channel between user client devices. In embodiments, the steps leading to two or more users engaging in a dialogue include: approaching a user graphical representation; selecting and clicking on the user graphical representation; sending or receiving a dialogue engagement invitation to or from another user graphical representation; and accepting the received invitation. The step of opening a communication channel between user client devices may be performed if processing and rendering are performed by the client devices, or the step of opening an indirect communication channel via one or more cloud server computers may be performed if processing and rendering are performed on at least one cloud server computer or between at least one cloud server and a client device. In other embodiments, the dialogue includes sending and receiving real-time audio to and from a participant's user 3D virtual clip. In other embodiments, the dialogue includes sending and receiving real-time audio and video from a user real-time 3D virtual clip or video with background removed, or a participant's video without background removal.
[0030] In some embodiments, the method of enabling interaction in a virtual environment further includes embedding a clickable link that redirects to the virtual environment into one or more third-party sources, including third-party websites, applications, or video games.
[0031] In another aspect of this disclosure, a data processing system includes one or more computing devices comprising at least one cloud server computer, the one or more computing devices including at least one processor and memory storing data and instructions for implementing image processing functions, wherein the one or more computing devices of the data processing system are configured to generate user graphic representations from live data feeds via one or more image processing combinations of at least one cloud server computer and two or more client devices in a hybrid system architecture. In an embodiment, the system includes: two or more client devices communicatively connected to each other via a network and connected to one or more cloud server computers, each including at least one processor and memory storing data and instructions for implementing image and media processing functions; and at least one camera that receives live data feeds from at least one user of at least one client device and is connected to at least one client device and one or more cloud server computers. User graphic representations are generated from live data feeds via one or more image and media processing combinations of one or more cloud server computers and one or more client devices. The one or more cloud server computers and one or more client devices interact through a hybrid system architecture.
[0032] In some embodiments, the data used as input data for the data processing system includes 2D or 3D image data, 3D geometry, video data, media data, audio data, text data, tactile data, time data, 3D entities, 3D dynamic objects, metadata, priority data, security data, location data, lighting data, depth data, and infrared data, etc.
[0033] In some embodiments, the hybrid system architecture includes a client-server side and a peer-to-peer (P2P) side. In one embodiment, the client-server side includes a network or application server. The client-server side may also be configured to include: a secure communication protocol; microservices; a database management system; a database; and / or a distributed messaging and resource distribution platform. Server-side components may be provided along with client devices that communicate with the server over a network. The client-server side defines the interaction between one or more clients and the server over a network, including any processing performed by the client, the server, or the receiving client. In one embodiment, one or more of the corresponding clients and servers perform necessary image and media processing according to various rule-based task assignment combinations. In one embodiment, the network or application server is configured to receive client requests using a secure communication protocol and process client requests by requesting microservices or data corresponding to the requests from the database using a database management system. The microservices are distributed using a publish-subscribe model leveraging a distributed messaging and resource distribution platform.
[0034] The P2P client includes: a P2P communication protocol enabling real-time communication between client devices in a virtual environment; and a rendering engine configured to enable client devices to perform real-time 3D rendering of live session elements (e.g., user graphical representations) included in the virtual environment. In an embodiment, the P2P client also includes a computer vision library configured to enable client devices to perform real-time computer vision tasks in the virtual environment. Using this hybrid communication model enables fast P2P communication between users, reduces latency issues, and provides network services and resources to each session, facilitating various interactions between users and with content in the virtual environment.
[0035] The P2P client defines the interaction between client devices and any processing that one or more client devices from the P2P client can perform. In some embodiments, the P2P client is used for video and data processing tasks, as well as synchronization, streaming, and rendering between client devices. In other embodiments, the P2P client is used for video streaming, rendering, and synchronization between client devices, while the client-server side is used for data processing tasks. In other embodiments, the client-server side is used for video streaming and data processing tasks, while the P2P client is used for video rendering and synchronization between client devices. In other embodiments, the client-server side is used for video streaming, rendering, and data processing tasks, as well as synchronization.
[0036] In one embodiment, the data processing task includes generating a user graphical representation and inserting that representation into a virtual environment. Generating the user graphical representation may include performing background removal or other processing or enhancements.
[0037] In some embodiments, data on the P2P side is sent directly from one client device to the peer client device and vice versa, or relayed through the server via the client-server connection.
[0038] In some embodiments, at least one cloud server may be an intermediary server, meaning that the server is used to facilitate and / or optimize data exchange between client devices. In such embodiments, at least one cloud server may manage, analyze, process, and optimize incoming image and multimedia streams, and manage, evaluate, and optimize the forwarding of outbound streams as a router topology (e.g., but not limited to SFU (Selective Forwarding Unit), SAMS (Spatial Analysis Media Server), multimedia router, etc.), or image and media processing (e.g., but not limited to decoding, combining, improving, mixing, enhancing, expanding, computing, manipulating, encoding) and forwarding server topology (e.g., but not limited to, multipoint control unit - MCU, cloud media mixer, cloud 3D renderer, etc.) or other server topologies.
[0039] In such embodiments, the intermediate server is a SAMS (Single Media Management System), which manages, analyzes, and processes incoming data sent to each client device (e.g., including but not limited to metadata, priority data, data categories, spatial structure data, 3D location, orientation or motion information, images, media, and video based on scalable video codecs). In this analysis, based on the specific receiving client device user's space, 3D orientation, distance, and priority relationship with such incoming data, the media is modified, scaled up, or scaled down for time (variable frame rates), space (e.g., different image sizes), quality (e.g., quality based on different compression or encoding), and color (e.g., color resolution and range). This optimizes the forwarding of outbound data streams to each receiving client device, achieving optimal bandwidth and computational resource utilization for receiving data from one or more user client devices.
[0040] In some embodiments, multiple image processing tasks are categorized based on whether they are performed by a client device, a cloud server, and / or a receiving client device, and are thus classified as client device image processing, server image processing, and receiving client device image processing. Multiple image processing tasks can be performed on a hybrid client-server architecture, a P2P architecture, or a combination thereof. Image processing tasks include background removal, further processing or improvement, and insertion and combination into a virtual environment. A combination of the three image processing tasks can be used to generate, improve, and insert / combine user graphical representations into a virtual environment. The image processing combinations of client device processing, server image processing, and receiving client device processing, and their corresponding usage levels, depend on the amount of data to be processed, the latency allowed to maintain a smooth user experience, the desired quality of service (QoS), the required services, etc. The following are eight such image processing combinations performed on the client-server side.
[0041] In some embodiments, at least one of the client devices is configured to, in a client-server image processing combination, generate a user graphical representation, perform background removal, and send the background-removed user graphical representation to at least one cloud server for further processing. In a first illustrative image processing combination, the client device generates a user graphical representation including background removal and sends the background-removed user graphical representation to at least one cloud server for further processing or improvement to generate an enhanced user graphical representation with the background removed. At least one cloud server sends the background-removed enhanced user graphical representation to a receiving client device, which inserts and combines the background-removed enhanced user graphical representation into a virtual environment.
[0042] In the second illustrative image processing combination, the client device generates a user graphical representation including background removal, performs further processing on it to generate an enhanced user graphical representation with the background removed, and then sends it to at least one cloud server. The at least one cloud server sends the enhanced user graphical representation with the background removed to a receiving client device, which then inserts and combines the enhanced user graphical representation with the background removed into the virtual environment.
[0043] In the third illustrative image processing combination, the client device generates a user graphic representation including background removal, performs further processing on it to generate an enhanced user graphic representation with the background removed, and inserts and combines the enhanced user graphic representation with the background removed into the virtual environment. The client device then sends the enhanced user graphic representation with the background removed and combined into the virtual environment to the cloud server for relay to the receiving client device.
[0044] In the fourth illustrative image processing combination, the client device generates a user graphic representation including background removal and sends the background-removed user graphic representation to at least one cloud server for further processing to generate an enhanced user graphic representation with the background removed. The at least one cloud server then inserts and combines the background-removed enhanced user graphic representation into a virtual environment and then sends it to the receiving client device.
[0045] In the fifth illustrative image processing combination, the client device generates a user graphic representation including background removal and sends the background-removed user graphic representation to at least one cloud server for relay to the receiving client device. The receiving client device performs further processing on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, which is then inserted and combined into the virtual environment.
[0046] In the sixth illustrative image processing assembly, the client device sends a camera live data feed received from at least one camera and unprocessed data to at least one cloud server. The at least one cloud server performs the generation of a user graphic representation including background removal, further processes the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, and sends it to the receiving client device. The receiving client device inserts and combines the enhanced background-removed user graphic representation into the virtual environment.
[0047] In the seventh illustrative image processing assembly, the client device sends a camera live data feed received from at least one camera and transmits unprocessed data to at least one cloud server. The at least one cloud server generates a user graphic representation containing background removal, performs further processing on the background-removed user graphic representation to generate an enhanced user graphic representation with the background removed, and then inserts and combines the enhanced user graphic representation with the background removed into the virtual environment sent to the receiving client device.
[0048] In the eighth illustrative image processing assembly, the client device sends a camera live data feed received from at least one camera to at least one cloud server and sends unprocessed data for relay to the receiving client device. The receiving client device uses this data to generate a user graphic representation including background removal, performs further processing on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, and then inserts and combines the enhanced background-removed user graphic representation into the virtual environment.
[0049] In some embodiments, when client-server data is relayed through at least one cloud server, at least one cloud server is configured to use a relay NAT traversal (TURN) server. TURN can be used in the case of symmetric NAT (Network Address Translation) and can remain in the media path after the connection is established, while processed and / or unprocessed data is relayed between client devices.
[0050] The following is a description of three illustrative image processing combinations performed on the P2P end by one or both of the first and second peer devices.
[0051] In the first image processing combination, the first peer device generates a user graphical representation including background removal, performs further processing on it to generate an enhanced user graphical representation with the background removed, and inserts and combines the enhanced user graphical representation with the background removed into the virtual environment. The first peer device then sends the enhanced user graphical representation with the background removed, inserted and combined into the virtual environment, to the second peer device.
[0052] In the second image processing combination, the first peer device generates a user graphic representation including background removal and sends the background-removed user graphic representation to the second peer device. The second peer device performs further processing on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, which is then inserted and combined into the virtual environment.
[0053] In the third image processing assembly, the first peer device sends a camera live data feed received from at least one camera and sends unprocessed data to the second peer device. The second peer device uses this data to generate a user graphic representation that includes background removal, performs further processing on the background-removed user graphic representation to generate an enhanced user graphic representation with the background removed, and then inserts and combines the enhanced user graphic representation with the background removed into the virtual environment.
[0054] In some embodiments, the three image processing combinations on the P2P side may further include data relay via at least one cloud server. In these embodiments, the at least one cloud server may be configured as a STUN server, which allows peer devices to discover their public IP addresses and the NAT types behind them; this information can be used to establish data connections and exchange data between peer devices. In another embodiment, the at least one cloud server computer may be configured for signaling, which can be used for peer device location and connection to each other, as well as for exchanging data through communication coordination performed by the at least one cloud server.
[0055] In some embodiments, media, video, and / or data processing tasks include one or more of encoding, code conversion, decoding, spatial or 3D analysis and processing, including one or more of image filtering, computer vision processing, image sharpening, background enhancement, background removal, foreground blurring, eye occlusion, face pixelation, speech distortion, image magnification, image cleansing, skeletal structure analysis, face or head counting, object recognition, marker or QR code tracking, eye tracking, feature analysis, 3D mesh or volume generation, feature tracking, face recognition, SLAM tracking, and facial expression recognition, or other modular plug-ins in the form of microservices running on such media routers or servers.
[0056] In some embodiments, background removal includes employing image segmentation through one or more of instance segmentation or semantic segmentation, as well as using a deep neural network.
[0057] In some embodiments, one or more computing devices of the data processing system are further configured to insert a user graphical representation into a virtual environment by generating virtual cameras, wherein generating virtual cameras includes associating captured viewpoint data with the viewpoint of the user graphical representation within the virtual environment. In embodiments, inserting and combining the user graphical representation into the virtual environment includes generating one or more virtual cameras that are virtually placed and aligned in front of the user graphical representation, for example, in front of a video with or without background removed, or a user 3D virtual cutout, or a user real-time 3D virtual cutout. In one embodiment, the one or more virtual cameras may point outwards from eye level. In another embodiment, two virtual cameras, one for each eye, may point outwards from both eye level. In another embodiment, the one or more virtual cameras may point outwards from the center of the user graphical representation's head position. In another embodiment, the one or more virtual cameras may point outwards from the center of the user graphical representation. In another embodiment, the one or more virtual cameras may be placed in front of the user graphical representation, for example, at head level, pointing towards the user graphical representation from a self-viewpoint.
[0058] In one embodiment, one or more virtual cameras are created by at least using computer vision to correlate captured user viewpoint data with the viewpoint of a user graphical representation within a virtual environment.
[0059] In another aspect of this disclosure, an image processing method includes: providing data and instructions for implementing image processing functions in the memory of at least one cloud server computer; and generating a user graphical representation in a virtual environment based on live data feeds from at least one client device through one or more image processing combinations of at least one cloud server computer and at least one client device, wherein the at least one cloud server computer interacts with at least one client device through a hybrid system architecture. In an embodiment, the method includes: acquiring live data feeds from at least one user of at least one corresponding client device from at least one camera; and generating a user graphical representation through one or more image processing combinations of one or more cloud server computers and at least one client device. The one or more cloud server computers and at least one client device can interact through a hybrid system architecture of this disclosure that includes P2P and client-server components.
[0060] In some embodiments, the method includes performing video and data processing, as well as synchronization, streaming, and rendering between client devices, by a P2P client. In other embodiments, the method includes performing video streaming, rendering, and synchronization between client devices by a P2P client, while a client-server side is used for data processing. In other embodiments, the method includes performing video streaming and data processing by a client-server side, while a P2P client is used for video rendering and synchronization between client devices. In other embodiments, the method includes performing video streaming, rendering, data processing, and synchronization by a client-server side.
[0061] In some embodiments, the data processing task includes generating a user graphical representation and inserting that representation into a virtual environment. In one embodiment, the data processing task includes first generating a user graphical representation that includes background removal, then performing further processing, and subsequently inserting and combining it into the virtual environment. In other embodiments, the image processing task is performed through a combination of multiple image processing operations on a client device and a cloud server computer in a client-server or P2P context.
[0062] In some embodiments, inserting a user graphical representation into a virtual environment includes generating virtual cameras, wherein generating virtual cameras includes associating captured viewpoint data with the viewpoint of the user graphical representation within the virtual environment. In embodiments, inserting and combining a user graphical representation into a virtual environment includes generating one or more virtual cameras that are virtually placed and aligned in front of the user graphical representation, for example, in front of a video with or without background removed, a user 3D virtual cutout, or a user real-time 3D virtual cutout. In one embodiment, the one or more virtual cameras may point outwards from eye level. In another embodiment, two virtual cameras, one for each eye, may point outwards from both eye levels. In another embodiment, the one or more virtual cameras may point outwards from the center of the user graphical representation's head position. In another embodiment, the one or more virtual cameras may point outwards from the center of the user graphical representation. In another embodiment, the one or more virtual cameras may be placed in front of the user graphical representation, for example, at the user's head level, pointing at the user graphical representation from a self-perspective. Virtual cameras are created at least by associating captured user viewpoint data with the viewpoint of the user graphical representation within the virtual environment using computer vision.
[0063] In some embodiments, the method further includes embedding a clickable link on the user's graphical representation, the link responding to a click and pointing to a third-party source containing profile information about the corresponding user.
[0064] In another aspect of this disclosure, a user authentication system based on user graphical representation includes: one or more cloud server computers including at least one processor and a memory storing data and instructions; a user database including user data associated with user accounts and one or more corresponding user graphical representations; and a face scanning and authentication module connected to the database; wherein the one or more cloud server computers are configured to perform steps including: authenticating a user by performing a face scan of the user through the face scanning and authentication module, wherein the face scan includes extracting facial feature data from camera data received from a client device, and checking the match between the extracted facial feature data and the user graphical representation associated with the user account in the user database; if a matching user graphical representation is found in the checking step, providing the user with access to the corresponding user account; and if no matching user graphical representation is found in the checking step, generating a new user graphical representation from the camera data and a new user account stored in the user database, and accessing the user account.
[0065] In one embodiment, the system includes at least one camera configured to obtain data from a user on at least one client device requesting access to a user account, wherein the at least one camera is connected to at least one client device and one or more cloud server computers. The one or more cloud server computers authenticate the user by: performing a facial scan of the user via the facial scanning and authentication module; checking the match between the user database and the user's graphical representation; and, if the user account is confirmed and available, providing the user with the corresponding user graphical representation and access to the user account; and, if the user account is unavailable, generating a new user graphical representation from the data and storing the new user account in the user database, along with access to the user account.
[0066] User accounts can be used, for example, to access a virtual environment platform or any other application (e.g., an application that can be linked to the environment platform), such as any interactive application, game, email account, university profile account, work account, etc. Given the additional authentication steps of generating a user graphical representation or retrieving an existing user graphical representation from a user database, the graphical representation-based user authentication system of this disclosure provides a higher level of security than standard camera-based face detection authentication systems.
[0067] In some embodiments, the user graphical representation is a user 3D virtual clip, or a real-time user 3D virtual clip with the background removed, or a video with the background removed, or a video without the background removed. In embodiments, the user graphical representation is a user 3D virtual clip constructed from a user-uploaded or third-party source photo, or a real-time user 3D virtual clip with the background removed, or a video with the background removed, or a video without the background removed, generated based on a real-time 2D or 3D live video stream data feed obtained from a camera. In some embodiments, one or more cloud server computers are also configured to animate the matched user graphical representation or a new user graphical representation. Animating the matched user graphical representation involves applying a machine vision algorithm to the corresponding user graphical representation by a client device or at least one cloud server computer to identify the user's facial expressions and graphically simulate facial expressions on the user graphical representation. In other embodiments, updating the user 3D virtual clip constructed from a user-uploaded or third-party source photo involves applying a machine vision algorithm to the generated user 3D virtual clip by a client device or at least one cloud server computer to identify the user's facial expressions and graphically simulate facial expressions on the user 3D virtual clip.
[0068] In some embodiments, one or more cloud server computers are also configured to check the date of a matching user graphic representation and determine whether the matching user graphic representation needs to be updated. In an embodiment, if a user account is available, and in response to one or more cloud server computers checking the date of available user graphic representations, the one or more cloud server computers determine whether an existing user graphic representation needs to be updated by comparing it with a corresponding threshold or security requirement. For example, if a system security update is required, it may be necessary to update all user graphic representations, or at least to update graphic representations created before a specified date. If a user graphic representation is required, the one or more cloud server computers generate a user graphic representation update request to the corresponding client device. If the user approves the request, the one or more cloud server computers or the client device continue to generate user graphic representations based on feeds from a live camera. If no update is required, the one or more cloud server computers continue to retrieve existing user graphic representations from the user database.
[0069] In some embodiments, the user graphical representation is inserted into a two-dimensional or three-dimensional virtual environment, or onto a third-party source linked to the virtual environment (e.g., by overlaying it on a screen of a third-party application or website integrated or coupled with the systems disclosed herein), and combined with the two-dimensional or three-dimensional virtual environment graphics.
[0070] In some embodiments, the generation of a user graphical representation occurs asynchronously with the user's access to their user account. For example, if the system determines that a user has been authenticated after performing a facial scan and detection based on the user graphical representation, the system can enable the user to access their user account while a new user graphical representation is being generated, so that it can be provided to the user once ready and then inserted and combined into the virtual environment.
[0071] In some embodiments, one or more cloud server computers are also configured to authenticate users via login authentication credentials, which include a personal identification number (PIN), or a username and password, or a combination thereof.
[0072] In some embodiments, authentication is triggered in response to the activation of an invitation link or deep link sent from one client device to another. In an embodiment, clicking the invitation link or deep link triggers at least one cloud server computer to request user authentication. For example, the invitation link or deep link can be used for telephone calls, conference calls, or video game session invitations, where the invited user can be authenticated through the user graph representation-based authentication system of this disclosure.
[0073] In another embodiment, facial scanning uses 3D authentication, which involves guiding the user to perform head movement patterns and extracting 3D facial data based on these patterns. This can be accomplished using application instructions stored in at least one server computer, which guide the user to perform head movement patterns to achieve 3D authentication, such as performing one or more head poses, tilting or rotating the head horizontally or vertically in a circular motion, performing user-generated pose patterns, or specific head movement patterns, or combinations thereof. 3D authentication identifies additional features based on data received from a live video feed from a camera, rather than simply comparing and analyzing a view or image. In this 3D authentication embodiment, the facial scanning process can identify additional features from data that may include facial data, including head movement patterns, facial volume, height, depth of facial features, facial scars, tattoos, eye color, facial skin parameters (e.g., skin color, wrinkles, pore structure, etc.), reflection parameters, and, for example, simply the location of such features on the facial topology, which may be the case in other types of facial detection systems. Capturing such facial data thus increases the capture of a realistic face, which can be used to generate a realistic graphical representation of the user. Facial scans for 3D authentication can be performed using high-resolution 3D cameras, depth cameras (e.g., LiDAR), light field cameras, etc. The facial scanning process and 3D authentication can utilize deep neural networks, convolutional neural networks, and other deep learning techniques to retrieve, process, and evaluate the user's authentication using facial data.
[0074] In another aspect of this disclosure, a user authentication method based on user graphical representation includes: providing a user database in the memory of one or more cloud server computers that stores user data associated with user accounts and one or more corresponding user graphical representations, and a face scanning and authentication module connected to the user database; receiving a request to access a user account from a client device; performing a face scan of the user on the client device by extracting facial feature data from camera data captured by at least one camera communicating with the client device through the face scanning and authentication module; checking the match between the extracted facial feature data and the user graphical representation associated with the user account in the user database; if a matching user graphical representation is found in the checking step, providing the user with access to the user account; and if no matching user graphical representation is found in the checking step, generating a new user graphical representation from the camera data and storing a new user account in the user database, and providing access to the user account.
[0075] In an embodiment, the method includes: performing a user's facial scan via a facial scanning and authentication module using image and / or media data received from at least one camera connected to at least one client device and one or more cloud server computers; checking for a match between a user database and user facial data associated with a user account; if the user account is available, providing the user with a corresponding user graphic representation and access to the user account; and if the user account is unavailable, generating a new user graphic representation from the facial data and storing a new user account in the user database and access to the user account.
[0076] In some embodiments, the user graphical representation is a 3D virtual clipping of the user constructed from user-uploaded or third-party source photos, or a real-time 3D virtual clipping of the user's real-time video stream with the background removed, generated based on a feed of real-time 2D or 3D live video stream data obtained from a camera, or a video with the background removed, or a video without the background removed. In other embodiments, the method includes animating the matched user graphical representation or a new user graphical representation, which may include applying machine vision algorithms to the corresponding user graphical representation by a client device or at least one cloud server computer to identify the user's facial expressions and graphically simulate facial expressions on the user graphical representation. In embodiments, updating the user 3D virtual clipping includes applying machine vision algorithms to the generated user 3D virtual clipping by a client device or at least one cloud server computer to identify the user's facial expressions and graphically simulate facial expressions on the user 3D virtual clipping.
[0077] In some embodiments, the method further includes: if a matching user graphical representation is found in the checking step, checking the date of the matching user graphical representation; determining, at least in part based on the date, whether the matching user graphical representation needs to be updated; and, if it is certain that the matching user graphical representation needs to be updated, generating a user graphical representation update request. In embodiments, the method includes: if a user account is available, checking the date of available user graphical representations; determining, by comparison with a corresponding threshold or security requirement, whether an existing user graphical representation needs to be updated; and, if it is certain that a user graphical representation needs to be updated, generating and transmitting a user graphical representation update request to the corresponding client device. If the user approves the request, one or more cloud server computers or client devices continue to generate user graphical representations based on live camera feeds. If no update is required, one or more cloud server computers continue to retrieve existing user graphical representations from the user database.
[0078] In some embodiments, the method further includes inserting a user graphical representation into a two-dimensional or three-dimensional virtual environment, or onto a third-party source linked to the virtual environment (e.g., by overlaying it onto a screen of a third-party application or website integrated or coupled with the system disclosed herein), and combining the user graphical representation with the two-dimensional or three-dimensional virtual environment graphics.
[0079] In some embodiments, the generation of a new user graphical representation occurs asynchronously with the user accessing their user account.
[0080] In some embodiments, the method further includes authenticating the user by logging in with authentication credentials containing at least a username and a password.
[0081] In some embodiments, the authentication is triggered in response to the activation of an invitation link. In another embodiment, the method further includes one client device providing an invitation link or a deep link to another client device, wherein clicking the invitation link triggers at least one cloud server computer to request authentication from the user.
[0082] In another aspect of this disclosure, a system for virtual broadcasting from within a virtual environment is provided. The system comprises: a server computer system including one or more server computers, each server computer including at least one processor and memory, the server computer system including data and instructions for implementing a data exchange management module configured to manage data exchange between client devices; and at least one virtual environment including a virtual broadcast camera positioned within the at least one virtual environment and configured to capture multimedia streams from within the at least one virtual environment. The server computer system is configured to receive live feed data captured by the at least one camera from at least one client device, and to broadcast the multimedia stream to the at least one client device based on the data exchange management, wherein the broadcast multimedia stream is configured to be displayed to a corresponding user graphical representation generated from the user live data feed from the at least one client device. The data exchange management between client devices performed by the data exchange management module includes analyzing the incoming multimedia stream, and evaluating and forwarding the outgoing multimedia stream based on the analysis of the incoming media stream.
[0083] In one embodiment, a multimedia stream is sent to at least one media server computer for broadcast to at least one client device. In another embodiment, the system includes at least one camera that acquires live feed data from a user on at least one client device and transmits the live feed data from the user to at least one media computer via the at least one client device; wherein, based on data exchange management from the at least one media server computer, the multimedia stream is broadcast to at least one client device and displayed to a corresponding user graphical representation generated by the live data feed from the user via the at least one client device; and wherein the data exchange management between client devices performed by the data exchange management module includes analyzing and optimizing the incoming multimedia stream, and evaluating and optimizing the forwarding of the outgoing multimedia stream.
[0084] In some embodiments, when forwarding outgoing multimedia streams, the server computer system utilizes a routing topology, including Selective Forwarding Unit (SFU), NAT traversal using relay (TURN), SAMS, or other suitable multimedia server routing topology, or media processing and forwarding server topology, or other suitable server topology.
[0085] In some embodiments, the server computer system uses a media processing topology to process outbound multimedia streams for viewing by a user graphical representation within at least one virtual environment via a client device. In embodiments, when utilizing the media processing topology, at least one media server computer is configured to decode, combine, improve, mix, enhance, expand, compute, manipulate, and encode multimedia streams to an associated client device for viewing by a user graphical representation within at least one virtual environment via the client device.
[0086] In some embodiments, when utilizing a forwarding server topology, the server computer system utilizes one or more of an MCU, a cloud media mixer, and a cloud 3D renderer.
[0087] In some embodiments, the incoming multimedia stream includes user priority data and distance relationship data, and the user priority data includes a higher priority score for user graphical representations closer to the incoming multimedia stream source and a lower priority score for user graphical representations farther from the incoming multimedia stream source. In embodiments, the multimedia stream includes data related to user priority and the distance relationship between the corresponding user graphical representation and the multimedia stream, including metadata, or priority data, or data category, or spatial structure data, or three-dimensional location, or orientation or motion information, or image data, or media data, and video data based on a scalable video codec, or a combination thereof. In other embodiments, the priority data includes a higher priority score for users closer to the multimedia stream source and a lower priority score for users farther from the virtual multimedia stream source. In other embodiments, the forwarding of the outbound multimedia stream is based on user priority data and distance relationship data. In embodiments, the forwarding of the outbound multimedia stream by the media server based on user priority and distance relationship data includes optimizing the bandwidth and computing resource utilization of one or more receiving client devices. In other embodiments, the forwarding of the outbound multimedia stream also includes modifying, enlarging, or reducing the temporal, spatial, quality, and / or color characteristics of the multimedia stream.
[0088] In some embodiments, the virtual broadcast camera is managed by a client device that accesses the virtual environment. In one embodiment, the virtual broadcast camera is configured to manipulate the viewpoint of the camera updated in the virtual environment and broadcast the updated viewpoint to at least one client device.
[0089] In some embodiments, at least one virtual environment includes a plurality of virtual broadcast cameras, each virtual broadcast camera providing a multimedia stream from a corresponding viewpoint within at least one virtual environment. In an embodiment, each virtual broadcast camera provides a multimedia stream from a corresponding viewpoint within the virtual environment, the multimedia stream being selectable by a user of at least one client device and alternating among themselves to provide a corresponding viewpoint to at least one corresponding user graphical representation.
[0090] In some embodiments, at least one virtual environment is hosted by at least one dedicated server computer connected to at least one media server computer via a network, or hosted in a peer-to-peer infrastructure and relayed through at least one media server computer.
[0091] In another aspect of this disclosure, a method for virtual broadcasting from within a virtual environment includes: providing data and instructions in the memory of at least one media server for implementing a client device data exchange management module that manages data exchange between client devices; capturing a multimedia stream by a virtual broadcast camera located within at least one virtual environment connected to at least one media server, and sending the multimedia stream to the at least one media server for broadcasting to at least one client device; obtaining live feed data from at least one client device (e.g., via at least one client device from at least one camera); performing data exchange management, including analyzing the incoming multimedia stream from the at least one virtual environment and the live feed data, and evaluating the forwarding of the outgoing multimedia stream; and broadcasting the corresponding multimedia stream to the client device based on the data exchange management, wherein the multimedia stream is displayed to a user graphical representation of a user of the at least one client device. In this context, this refers to what the user graphical representation can “see” based on their position in the virtual environment, which corresponds to what will be displayed to the user (via the client device) when viewing the virtual environment from the perspective of his or her user graphical representation.
[0092] In some embodiments, when forwarding outgoing multimedia streams, the method utilizes routing topologies, including SFU, TURN, SAMS, or other suitable multimedia server routing topologies, or media processing and forwarding server topologies, or other suitable server topologies.
[0093] In some embodiments, the method further includes decoding, combining, improving, mixing, enhancing, expanding, computing, manipulating, and encoding multimedia streams while utilizing a media processing topology.
[0094] In some embodiments, when utilizing a forwarding server topology, the method further includes utilizing one or more of a multipoint control unit (MCU), a cloud media mixer, and a cloud 3D renderer.
[0095] In some embodiments, the incoming multimedia stream includes user priority data and distance relationship data, and the user priority data includes a higher priority score for user graphic representations closer to the incoming multimedia stream source and a lower priority score for user graphic representations farther from the incoming multimedia stream source. In embodiments, the method further includes optimizing the forwarding of outbound multimedia streams implemented by a media server based on user priority and distance relationship data, which may include optimizing the bandwidth and computing resource utilization of one or more receiving client devices. In other embodiments, optimizing the forwarding of outbound multimedia streams implemented by a media server further includes modifying, amplifying, or reducing the temporal, spatial, quality, and / or color characteristics of the multimedia stream.
[0096] In some embodiments, at least one virtual environment includes a plurality of virtual broadcast cameras, each virtual broadcast camera providing a multimedia stream from a corresponding viewpoint within the at least one virtual environment. In another embodiment, the method further includes providing a plurality of virtual broadcast cameras, each virtual broadcast camera providing a multimedia stream from a corresponding viewpoint within the virtual environment, the multimedia streams being selectable by a user of at least one client device and alternating among themselves to provide a corresponding viewpoint to at least one corresponding user graphical representation.
[0097] In another aspect of this disclosure, a system for delivering applications within a virtual environment is provided, comprising: at least one cloud server computer including at least one processor and memory, the memory including data and instructions for implementing at least one virtual environment linked to an application module, the application module including one or more installed applications and corresponding application rules for multi-user interaction; wherein, in response to selection by a virtual environment host via a client device, one or more installed applications are displayed and activated during a virtual environment session, such that a user graphical representation of the virtual environment host and a user graphical representation of any participant within the virtual environment can interact with one or more installed applications via corresponding client devices, and wherein the at least one cloud server computer manages and processes received user interactions with one or more installed applications according to the application rules for multi-user interaction in the application module, and forwards the processed interactions accordingly (e.g., to each client device) to establish a multi-user session realizing a shared experience according to the multi-user interaction application rules.
[0098] In some embodiments, application rules for multi-user interaction are stored and managed on one or more independent application servers.
[0099] In some embodiments, one or more applications are installed from application installation packages available in the application library and provide application services through corresponding application programming interfaces.
[0100] In some embodiments, the application library is context-filtered. In these embodiments, context filtering is designed to provide relevant applications for specific contexts.
[0101] In some embodiments, one or more installed applications are shared with and viewed through a virtual display application installed on a corresponding client device. In one embodiment, after installation and activation, one or more installed applications are shared with and viewed through the virtual display application installed on a corresponding client device, wherein the virtual display application is configured to receive one or more installed applications from an application library and publish one or more selected applications to be displayed to the conference host user graphical representation and other participant user graphical representations in the virtual environment via their corresponding client devices. In other embodiments, application modules are represented as 2D screens or 3D volume application module graphical representations that display content from installed applications to user graphical representations in the virtual environment, and wherein the virtual display application is represented as a 2D screen or 3D volume that displays content from installed applications to user graphical representations in the virtual environment.
[0102] In some embodiments, one or more applications are installed directly within the virtual environment before or simultaneously with a multi-user session.
[0103] In some embodiments, one or more applications are installed using a virtual environment setup tool before starting a multi-user session.
[0104] In some embodiments, one or more of the application rules for multi-user interaction define synchronous interaction, asynchronous interaction, or a combination thereof. In embodiments, such rules are used accordingly to update user interactions and corresponding updated views of one or more applications.
[0105] In some embodiments, asynchronous interaction is achieved through at least one server computer or through a separate server computer dedicated to handling interaction with a single user who has at least one installed application.
[0106] In some embodiments, the virtual environment is a classroom, or office space, or meeting room, or conference room, or auditorium, or theater.
[0107] In another aspect of this disclosure, a method for delivering an application within a virtual environment is provided, comprising: providing at least one virtual environment in the memory of at least one cloud server computer, and an application module including one or more installed applications and corresponding application rules for multi-user interaction, wherein the application module is linked to and visible within the virtual environment; receiving a selection instruction from a virtual environment host; displaying and activating one or more installed applications during a session in the virtual environment, such that a user graphical representation of the virtual environment host and one or more participant user graphical representations within the virtual environment can interact with the one or more installed applications via corresponding client devices; managing and processing user interactions with the one or more installed applications according to the application rules for multi-user interaction in the application module; and forwarding the processed interactions to the client devices to establish a multi-user session for a shared experience according to the application rules.
[0108] In some embodiments, the method further includes storing and managing application rules for multi-user interaction in one or more independent application servers.
[0109] In some embodiments, the method further includes installing one or more applications from an application installer available in an application library; and providing application services through a corresponding application programming interface. In other embodiments, the application library is context-filtered to provide relevant applications. In other embodiments, one or more installed applications are shared with and viewed through a virtual display application installed on a corresponding client device. In an embodiment, the method includes, upon activation, sharing with and viewing one or more installed applications through a virtual display application installed on a corresponding client device, the virtual display application being configured to receive one or more installed applications from an application library and publish one or more selected applications to be displayed to the conference host user graphical representation and other participant user graphical representations in the virtual environment through their corresponding client devices.
[0110] In some embodiments, the method further includes installing one or more applications directly within the virtual environment before or simultaneously with the multi-user session. In other embodiments, the method further includes installing one or more applications using a virtual environment setup tool before initiating the multi-user session.
[0111] In some embodiments, the method further includes defining one or more application rules for multi-user interaction to include synchronous interaction, asynchronous interaction, or a combination thereof. In embodiments, the method also includes updating the user interaction and corresponding updated views of one or more applications accordingly.
[0112] In another aspect of this disclosure, a system for providing virtual computing resources within a virtual environment includes a server computer system comprising one or more server computers including at least one cloud server computer, the at least one cloud server computer including at least one processor and memory including data and instructions for implementing at least one virtual environment, and at least one virtual computer associated with the at least one virtual environment, wherein the at least one virtual computer receives virtual computing resources from the server computer system. This association may include connecting the virtual computer to the virtual environment. In an embodiment, the at least one virtual computer has a corresponding graphical representation in the virtual environment. The graphical representation can provide other benefits, such as facilitating user interaction with the virtual computer and increasing the realism of the user experience (e.g., for a home office experience). Therefore, in an embodiment, the at least one virtual computer includes at least one corresponding associated graphical representation located within the virtual environment, wherein the virtual computer receives virtual computing resources from at least one cloud server computer; and at least one client device connected to the at least one server computer via a network; wherein, in response to the at least one client device accessing one or more virtual computers (e.g., by interacting with the corresponding graphical representation), the at least one cloud server computer provides at least a portion of the available virtual computing resources to the at least one client device.
[0113] In some embodiments, a server computer system is configured to provide at least a portion of virtual computing resources to at least one client device in response to a user graphical representation interacting with at least one corresponding graphical representation of at least one virtual computer within at least one virtual environment. In other embodiments, one or more virtual computer graphical representations are spatially located within a virtual environment for access by the user graphical representation. In embodiments, the arrangement of the virtual environment is associated with a contextual theme, such as the arrangement of virtual items, furniture, floor plans, etc., for educational, meeting, work, shopping, service, social, or entertainment purposes. In other embodiments, one or more virtual computer graphical representations are located within the arrangement of the virtual environment for access by one or more user graphical representations. For example, a virtual computer may be located in a virtual room that a user graphical representation would access when participating in activities that might require or benefit from the ability to use resources associated with the virtual computer (such as conducting projects in a virtual classroom, laboratory, or office).
[0114] In some embodiments, the server computer system is configured to provide at least a portion of virtual computing resources to at least one client device in response to a user accessing at least one cloud server computer by logging into at least one client device without accessing a virtual environment. In an illustrative scenario, a user accessing at least one cloud server computer accesses virtual computing resources by physically logging into a client device connected to at least one cloud server computer via a network, triggering the provision of virtual computing resources to the client device without accessing a virtual environment.
[0115] In some embodiments, at least a portion of the virtual computing resources is assigned to client devices via a management tool. In other embodiments, the provision of at least a portion of the virtual computing resources is performed based on a stored user profile. In one embodiment, resource assignment is performed based on a stored user profile containing one or more parameters associated with and assigned to the user profile, including priority data, security data, QoS, bandwidth, storage space, or computing power, or a combination thereof.
[0116] In some embodiments, at least one virtual computer contains a downloadable application available from an application library. In illustrative cases involving multiple virtual computers, each virtual computer is a downloadable application available from an application library.
[0117] In another aspect of this disclosure, a method for providing virtual computing resources within a virtual environment includes: providing at least one virtual computer and a virtual environment associated with the at least one virtual computer in the memory of at least one cloud server computer; associating virtual computing resources with the at least one virtual computer; receiving an access request from at least one client device to access one or more virtual computers; and, in response to the access request received from the at least one client device, providing the at least one client device with a portion of the available virtual computing resources associated with the at least one virtual computer. In embodiments, the association of the virtual computing resources with the at least one virtual computer may include the virtual computer receiving the virtual computing resources from the at least one cloud server computer.
[0118] In some embodiments, the access request includes a request allowing a user graphical representation to interact with one or more graphical representations representing at least one virtual computer. In another embodiment, the method further includes: receiving from a user graphical representation an access request to one or more graphical representations of a virtual computer within the at least one virtual environment; and providing at least a portion of available virtual computing resources to a corresponding client device. In other embodiments, the arrangement of the virtual environment is associated with a contextual theme of the virtual environment, including arrangements for education, meetings, work, shopping, services, socializing, or entertainment, and wherein one or more virtual computers are located within the arrangement of the virtual environment for access by one or more user graphical representations.
[0119] In some embodiments, the access request is triggered by a user logged into at least one client device. In another embodiment, the method further includes: receiving an access request from a user physically logged into a client device connected to at least one cloud server computer via a network; and providing virtual computing resources to the client device without accessing the virtual environment.
[0120] In some embodiments, the method further includes assigning at least a portion of virtual computing resources to client devices via a management tool. In other embodiments, the assignment is performed based on a stored user profile containing one or more parameters associated with and assigned to the user profile, including priority data, security data, QoS, bandwidth, storage space, computing power, or a combination thereof.
[0121] In another aspect of this disclosure, a system for enabling self-organizing virtual communication between user graphical representations includes one or more cloud server computers, each including at least one processor and a memory storing data and instructions for implementing a virtual environment. The virtual environment is configured such that at least one proximate user graphical representation and at least one target user graphical representation within the virtual environment can open a self-organizing communication channel, enabling self-organizing dialogue between user graphical representations within the virtual environment via this self-organizing communication channel. In an embodiment, the system further includes two or more client devices accessing at least one virtual environment through corresponding user graphical representations and connected to one or more cloud server computers via a network; wherein the virtual environment enables at least one proximate user graphical representation and at least one target user graphical representation to open a self-organizing communication channel, thereby enabling self-organizing dialogue between user graphical representations within the virtual environment.
[0122] In some embodiments, the opening of a self-organizing communication channel is performed based on the distance, location, and orientation between user graphical representations, or the current availability status, privacy settings, or the state configuration of self-organizing communication, or a combination thereof.
[0123] In some embodiments, the self-organizing dialogue is performed at the location of two user graphical representation areas within the virtual environment. In other embodiments, the self-organizing dialogue is performed using the current viewpoint within the virtual environment.
[0124] In some embodiments, self-organizing dialogue enables the optional change of perspective, position, or combination thereof within the same or another connected virtual environment where the self-organizing dialogue takes place.
[0125] In some embodiments, one or more cloud server computers are also configured to generate visual feedback in the virtual environment, signaling the possibility of self-organizing communication. In an embodiment, a user graphical representation receives visual feedback that signals the possibility of self-organizing communication, thereby triggering the opening of a self-organizing communication channel and signaling the start of a self-organizing dialogue between user graphical representations.
[0126] In some embodiments, the self-organizing dialogue includes sending and receiving real-time audio and video. In illustrative cases, such video may be displayed from a user graphical representation.
[0127] In some embodiments, a user corresponding to a nearby user graphical representation selects and clicks on a target user graphical representation before opening the self-organizing communication channel. In other embodiments, one or more cloud server computers are also configured to open the self-organizing communication channel in response to an accepted invitation. For example, a user corresponding to a nearby user graphical representation sends a self-organizing communication participation invitation to the target user graphical representation and opens the self-organizing communication channel after receiving invitation approval from the target user graphical representation.
[0128] In some embodiments, the self-organizing communication channel is implemented via at least one cloud server computer or as a P2P communication channel.
[0129] In another aspect of this disclosure, a method for enabling self-organizing virtual communication between user graphical representations includes: providing a virtual environment in the memory of one or more cloud server computers including at least one processor; detecting two or more client devices accessing at least one virtual environment through corresponding graphical representations, wherein the client devices are connected to one or more cloud server computers via a network; and opening a self-organizing communication channel in response to at least one user graphical representation approaching another user graphical representation, thereby enabling self-organizing dialogue between user graphical representations in the virtual environment.
[0130] In some embodiments, the method further includes detecting and evaluating one or more of the following: distance, location and orientation, or current availability status, privacy settings, or state configuration of the self-organizing communication, or combinations thereof, between user graphical representations before opening the self-organizing communication channel.
[0131] In some embodiments, the method enables self-organizing dialogue to be executed at the locations of two user graphical representation areas within a virtual environment. In other embodiments, the self-organizing dialogue is executed using the current viewpoint within the virtual environment.
[0132] In some embodiments, the method includes optionally changing the viewpoint, position, or a combination thereof within the same or another connected virtual environment where a self-organizing dialogue is taking place.
[0133] In some embodiments, the method further includes generating visual feedback in a virtual environment to signal that self-organizing communication is possible. The method may also include sending visual feedback to a target user graphical representation to signal that self-organizing communication is possible, thereby triggering the opening of a self-organizing communication channel and signaling the start of a dialogue between user graphical representations.
[0134] In some embodiments, the conversation includes sending and receiving real-time audio and video from a user graphical representation display.
[0135] In some embodiments, the method further includes selecting and clicking a target user graphical representation from a nearby user graphical representation. In other embodiments, an ad hoc communication channel is opened in response to an accepted invitation. In embodiments, the method further includes sending or receiving an ad hoc virtual communication participation invitation to another user graphical representation before opening the ad hoc communication channel.
[0136] It also describes a computer-readable medium on which instructions are stored, which are configured to cause one or more computers to perform any of the methods described herein.
[0137] The foregoing summary does not constitute an exhaustive list of all aspects of this disclosure. It is conceivable that this disclosure encompasses all systems and methods that can be practiced from all suitable combinations of the aspects outlined above, as well as those disclosed in the following detailed description and specifically pointed to in the claims filed with this application. Such combinations have advantages not specifically described in the foregoing summary. Other features and advantages will become apparent from the accompanying drawings and the detailed description below. Attached Figure Description
[0138] The specific features, aspects, and advantages of this disclosure will be better understood with reference to the following description and accompanying drawings, in which:
[0139] Figure 1 A schematic diagram of a system for implementing interactions, including social interactions, in a virtual environment, according to an embodiment, is depicted.
[0140] Figure 2A-2BA schematic diagram depicts the deployment of a system that enables interaction, including social interaction within virtual environments across multiple vertical sectors that incorporate virtual environment platforms.
[0141] Figure 3 A schematic diagram of a hybrid system architecture employed in a system for implementing interaction in a virtual environment, according to an embodiment, is depicted.
[0142] Figure 4 A schematic diagram of a graphical user interface according to an embodiment is depicted, in which a user can interact in a virtual environment.
[0143] Figure 5 A block diagram depicts a method according to an embodiment for transforming a user's 3D virtual cutout into a user's real-time 3D virtual cutout, or a video with or without background removal.
[0144] Figures 6A-6C A schematic diagram depicts a combination of multiple image processing operations performed on the client-server side by the corresponding client device and cloud server.
[0145] Figures 7A-7C A schematic diagram depicts a combination of multiple image processing operations performed by corresponding peer-to-peer clients on the P2P side.
[0146] Figure 8 A schematic diagram of a user authentication system based on a user graphical representation according to an embodiment is depicted.
[0147] Figure 9 A schematic diagram depicting a third-person perspective of a virtual office environment according to an embodiment is provided.
[0148] Figures 10A-10B A schematic diagram depicts a virtual classroom environment according to an embodiment.
[0149] Figure 11 A schematic diagram depicting the positions of multiple virtual cameras according to an embodiment is provided.
[0150] Figure 12 A schematic diagram of a system for virtual broadcasting from within a virtual environment is depicted.
[0151] Figure 13 A schematic diagram of a system for delivering applications within a virtual environment is depicted.
[0152] Figure 14 The description illustrates the use of a method based on embodiments for... Figure 13 The diagram illustrates a virtual environment in which a system delivers applications within a virtual environment.
[0153] Figure 15 A schematic diagram of a system for providing virtual computing resources within a virtual environment, according to an embodiment, is depicted.
[0154] Figure 16 A schematic diagram of a system for implementing self-organizing virtual communication between user graphical representations according to an embodiment is depicted.
[0155] Figure 17 An embodiment of a method for implementing interaction in a virtual environment is described according to an embodiment.
[0156] Figure 18 An embodiment of the image processing method according to the embodiments is described.
[0157] Figure 19 A user authentication method 1900 based on user graphical representation according to an embodiment is described.
[0158] Figure 20 A block diagram illustrating a method for virtual broadcasting from within a virtual environment according to an embodiment is shown.
[0159] Figure 21 A block diagram illustrating a method for delivering an application from within a virtual environment, according to an embodiment, is shown.
[0160] Figure 22 A block diagram illustrating a method for providing virtual computing resources within a virtual environment according to an embodiment is shown.
[0161] Figure 23 A block diagram illustrating a method for implementing self-organizing virtual communication between user graphical representations is shown. Detailed Implementation
[0162] In the following description, reference is made to the accompanying drawings, which illustrate various embodiments by way of illustration. Furthermore, various embodiments will be described below with reference to several examples. It should be understood that embodiments may include changes in design and structure without departing from the scope of the claimed subject matter.
[0163] The systems and methods disclosed herein address at least some of the aforementioned drawbacks by providing a virtual environment platform comprising one or more virtual environments capable of enabling real-time multi-user collaboration and interaction similar to that available in real life, suitable for meetings, work, education, shopping, and services. The virtual environments can be selected from multiple virtual environments across different vertical domains available on the virtual environment platform. Combinations of virtual environments from the same and / or different vertical domains can form a virtual environment cluster, which can contain hundreds or even thousands of virtual environments. Virtual environments can be 2D or 3D, including an arrangement and visual appearance associated with the vertical domain of the virtual environment, which can be customized by users according to their preferences or needs. Users can access virtual environments through graphical representations that can be interpolated into and graphically combined with two-dimensional or three-dimensional virtual environments.
[0164] User graphical representations can be user 3D virtual cutouts constructed from user-uploaded or third-party source photos with backgrounds removed, or real-time user 3D virtual cutouts, or videos with backgrounds removed, or videos without backgrounds removed; any of these can be switched between at any time as needed by the user. User graphical representations can include user state providing further details about current availability or other data related to other users. Interactions such as dialogue and collaboration between users in the virtual environment and interactions with objects within the virtual environment are implemented. This disclosure also provides a data processing system and method comprising multiple image processing combinations that can be used to generate user graphical representations. This disclosure also provides: a user authentication system and method based on user graphical representations, which can be used to access a virtual environment platform or other applications linked to the virtual environment from the virtual environment platform; a system and method for virtual broadcasting from within a virtual environment; a system and method for delivering applications within a virtual environment; a system and method for providing cloud computing-based virtual computing resources within a virtual environment cloud server computer; and a system and method capable of enabling self-organized virtual communication between proximate user graphical representations.
[0165] Realizing virtual presence within the virtual environment and enabling realistic interaction and collaboration between users can increase the realism of remote activities, such as those required during pandemics or other situations restricting movement. The systems and methods of this disclosure also enable access to various virtual environments on client devices such as mobile devices or computers, without requiring more expensive immersive devices, such as extended reality head-mounted displays or costly new system infrastructures. Client or peer devices of this disclosure may include, for example, computers, headsets, mobile phones, glasses, transparent screens, tablet computers, and input devices that typically have a built-in camera, or input devices that can connect to and receive data feeds from said camera.
[0166] Figure 1 A schematic diagram of a system 100 for implementing social interaction in a virtual environment according to an embodiment is depicted.
[0167] The system 100 disclosed herein for implementing interaction in a virtual environment includes one or more cloud server computers 102, each cloud server computer including at least one processor 104 and a memory 106 storing data and instructions for implementing a virtual environment platform 108, the virtual environment platform including at least one virtual environment 110, such as a virtual environment AC. The one or more cloud server computers are configured to insert user graphical representations generated from live data feeds obtained from cameras at three-dimensional coordinate positions in at least one virtual environment, update the user graphical representations in at least one virtual environment, and enable real-time multi-user collaboration and interaction in the virtual environment. In the described embodiments, inserting user graphical representations into the virtual environment involves graphically combining the user graphical representations in the virtual environment such that the user graphical representations appear in the virtual environment (e.g., at a specified 3D coordinate position). Figure 1 In the example shown, system 100 also includes at least one camera 112 that receives a live data feed 114 from a user 116 on client device 118. One or more client devices 118 are communicatively connected to one or more cloud server computers 102 and at least one camera 112 via a network. A user graphical representation 120 generated from the live data feed 114 is interpolated to the three-dimensional coordinate position of the virtual environment 110 (e.g., virtual environment A) and combined with the virtual environment graphics, as well as updated using the live data feed 114. Updated virtual environments are provided to client devices either through direct P2P communication or indirectly through the use of one or more cloud servers 102. System 100 enables real-time multi-user collaboration and interaction in the virtual environment 110 by accessing a graphical user interface via client device 118.
[0168] exist Figure 1 In this scenario, two users 116 (e.g., users A and B) are accessing a virtual environment A and are interacting with its elements and each other via their respective user graphical representations 120 (e.g., user graphical representations A and B) accessed through their respective client devices 118 (e.g., client devices A and B). Although in Figure 1 Only two users 116, client device 118, and user graphical representation 120 are depicted in the figure, but those skilled in the art will understand that the system enables multiple users 116 to interact with each other via their respective client devices 118 and their corresponding graphical representations 120.
[0169] In some embodiments, the virtual environment platform 108 and the corresponding virtual environment 110 can enable the real-time sharing of multiple experiences, such as live performances, concerts, webinars, keynote speeches, etc., to multiple (e.g., thousands or even millions) user graphical representations 120. These virtual performances can be presented by multiple instances of the virtual environment 110 and / or multicast to multiple instances of the virtual environment 110 to accommodate a large number of users 116 from around the world.
[0170] In some embodiments, client device 118 may be one or more of a mobile device, personal computer, game console, media center, and head-mounted display. Camera 112 may be one or more of a 2D or 3D camera, 360-degree camera, webcam, RGBD camera, CCTV camera, professional camera, mobile phone camera, depth camera (e.g., LiDAR), or light field camera.
[0171] In some embodiments, virtual environment 110 refers to a virtual building (e.g., a virtual model) designed using computer-aided drafting (CAD) methods and any suitable 3D modeling technique. In other embodiments, virtual environment 110 refers to a virtual building scanned from a real building (e.g., a physical room) using any suitable scanning tool, comprising an image scanning pipeline of various photographic, video, depth measurement, and / or simultaneous localization and mapping (SLAM) scan inputs to generate virtual environment 110. For example, radar imaging such as synthetic aperture radar, real aperture radar, light detection and ranging (LIDAR), inverse aperture radar, monopulse radar, and other types of imaging techniques can be used to map and model real-world structures and transform them into virtual environment 110. In other embodiments, virtual environment 110 is a virtual building modeled from a real building (e.g., a room, building, or facility in the real world).
[0172] In some embodiments, the client device 118 is connected to at least one cloud server computer 102 via a wired or wireless network. In some embodiments, the network may include millimeter-wave (mmW) or a combination of mmW and sub-6 GHz communication systems, such as fifth-generation wireless communication (5G). In other embodiments, the system may be connected via a wireless local area network (Wi-Fi). In other embodiments, the system may be connected via fourth-generation wireless communication (4G), may be supported by a 4G communication system, or may include other wired or wireless communication systems.
[0173] In some embodiments, when a live data feed 114 is received from user 116, at least one processor of client device 118 performs processing and rendering including the generation, updating, and insertion of user graphical representation 120 into a selected virtual environment 110 and combinations thereof. One or more cloud server computers 102 may receive the client-rendered user graphical representation 120, insert the client-rendered user graphical representation 120 into the three-dimensional coordinates of the virtual environment 110, combine the inserted user graphical representation 120 with the virtual environment 110, and then continue to transmit the client-rendered user graphical representation 120 to the receiving client device. For example, as... Figure 1 As shown, client device A can receive live data feed 114 from a corresponding camera 112, process and render the data from the live data feed 114 to generate a user graphic representation A, and then transmit the client-rendered user graphic representation A to at least one cloud server computer 102, which can position the user graphic representation A in three-dimensional coordinates of the virtual environment 110, and then transmit the user graphic representation A to client device B. A similar process applies to client device B and user graphic representation B from user B. User graphic representations A and B can therefore view and interact with each other in the virtual environment A. However, various other image processing combinations can be implemented using the systems and methods of this disclosure, as referenced in [reference]. Figures 6A-7C As shown and described.
[0174] In some embodiments, when client device 118 sends an unprocessed live data feed 114 to user 116, at least one processor 104 of one or more cloud server computers 102 performs processing and rendering including the generation, updating, and insertion of the user graphical representation 120 and its combination with the virtual environment. One or more cloud server computers 102 thus receive the unprocessed live data feed 114 from client device 118, and then generate, process, and render the user graphical representation 120 located in three-dimensional coordinates within the virtual environment 110 from the unprocessed live data feed, and then transmit the cloud-rendered user graphical representation within the virtual environment to other client devices 118. For example, as... Figure 1 As shown, client device A can receive live data feed 114 from a corresponding camera 112, and can then transmit unprocessed user live data feed 114 to at least one cloud server computer 102, which can generate, process, and render user graphical representation A and position user graphical representation A in three-dimensional coordinates of virtual environment 110, and then transmit user graphical representation A to client device B. A similar process applies to client device B and user graphical representation B from user B. User graphical representations A and B can therefore view and interact with each other in virtual environment A.
[0175] In some embodiments, the virtual environment platform 108 is configured to embed clickable links that redirect to a virtual environment into one or more third-party sources, including third-party websites, applications, or video games. The link may be, for example, an HTML link. The linked virtual environment 110 may be associated with the content of the website into which the link is embedded. For example, the link may be embedded on a car dealership or manufacturer's website, where the clickable link redirects to a virtual environment 110 representing a car dealership showroom that a user can access through a user graphical representation 120.
[0176] In some embodiments, the user graphical representation 120 includes clickable links embedded thereon, such as links to third-party sources containing profile information about the corresponding user. For example, the clickable link could be an HTML link embedded in the source code of the user graphical representation 120 that can authorize access to social media (e.g., professional social media websites such as LinkedIn). TM Access to the corresponding user's graphical representation can provide additional information about the user. In some embodiments, if the user permits, at least some basic information about the user can be displayed when another user clicks or hovers the cursor over the corresponding user's graphical representation. This can be done by accessing and retrieving user data from a database or a third-party source.
[0177] In some embodiments, the user graphical representation is a 3D virtual clip of the user constructed from a photo uploaded by the user or from a third-party source (e.g., from a social media website), or a real-time 3D virtual clip of the user 116 including a real-time video stream with the background removed, or a video with the background removed, or a video without the background removed. In other embodiments, the client device 118 generates the user graphical representation 120 by processing and analyzing the live data feed 114 of the user 116, generating animated data that is sent to other peer client devices 118 via a peer-to-peer (P2P) system architecture or a hybrid system architecture, as referenced. Figure 3 Further described. The receiving peer client device 118 uses animation data to locally build and update the user's graphical representation.
[0178] User 3D virtual clipping may include a virtual copy of the user constructed from a 2D photograph uploaded by the user or from a third-party source. In one embodiment, using a 2D photograph uploaded by the user or from a third-party source as input data, a user 3D virtual clipping is created via a 3D virtual reconstruction process using machine vision technology, resulting in a 3D mesh or 3D point cloud of the user with the background removed. In one embodiment, the user 3D virtual clipping may have a static facial expression. In another embodiment, the user 3D virtual clipping may include facial expressions updated via camera feeds. In yet another embodiment, the user 3D virtual clipping may include expressions that can be changed via buttons on the user's graphical interface, such as buttons allowing the user 3D virtual clipping to smile, frown, be serious, etc. In yet another embodiment, the user 3D virtual clipping uses a combination of the above techniques to display facial expressions. After the user 3D virtual clipping is generated, the state and / or facial expression of the user 3D virtual clipping can be continuously updated, for example, by processing camera feeds from the user. However, if the camera is not turned on, the user 3D virtual clipping may still be visible to other users with an unavailable state and a static facial expression. For example, a user might be focused on a task and might not want to be disturbed (e.g., having a "Do Not Disturb" or "Busy" status), so they turn off their camera. At this point, the user's 3D virtual cutout might simply be sitting at their desk, possibly stationary, or might be performing pre-configured movements, such as typing. However, when the user's camera is turned back on, the user's 3D virtual cutout can be updated again in real-time regarding the user's facial expressions and / or movements. Standard 3D facial model reconstruction techniques (e.g., 3D facial fitting and texture fusion) used to create the user's 3D virtual cutout can be used, making the resulting user graphical representation clearly recognizable as the user.
[0179] Real-time 3D virtual clipping of a user can include a virtual copy of the user based on a real-time 2D or 3D live video stream data feed obtained from a camera and after removing the user's background. In embodiments, real-time 3D virtual clipping of a user is created by generating a 3D mesh or 3D point cloud of the user with the background removed, using the user live data feed as input data, and via a 3D virtual reconstruction process using machine vision techniques. For example, real-time 3D virtual clipping of a user can be generated from 2D video from a camera (e.g., a webcam), which can be processed to create a holographic 3D mesh or 3D point cloud. In another example, real-time 3D virtual clipping of a user can be generated from 3D video from a depth camera (e.g., a LiDAR or any depth camera), which can be processed to create a holographic 3D mesh or 3D point cloud. Thus, real-time 3D virtual clipping of a user represents the user graphically in three dimensions and in real time.
[0180] Video with background removed can include video streamed to a client device where the background removal process has been performed, making it visible only to the user, and then displayed using a polygonal structure on the receiving client device. Video without background removal can include video streamed to a client device where the video faithfully represents a camera capture, making the user and their background visible, and then displayed using a polygonal structure on the receiving client device. The polygonal structure can be a quadrilateral structure or a more complex 3D structure, used as virtual frames to support the video.
[0181] Videos without background removal can include video streamed to a client device, where the video faithfully represents a camera capture, making the user and his or her background visible, and then displayed using polygonal structures on the receiving client device. The polygonal structures can be quadrilateral structures or more complex 3D structures, used as virtual frames to support the video.
[0182] In some embodiments, the data used as input data included in live data feeds and / or user-uploaded or third-party source 2D photographs includes 2D or 3D image data, 3D geometry, video data, media data, audio data, text data, haptic data, time data, 3D entities, 3D dynamic objects, metadata, priority data, security data, location data, lighting data, depth data, and infrared data, etc.
[0183] In some embodiments, the background removal process required to achieve real-time 3D virtual cropping for the user is performed using image segmentation and deep neural networks, which can be implemented by one or more processors of the client device 118 or at least one cloud server computer 102. Image segmentation is the process of dividing a digital image into multiple objects, which can help locate objects and boundaries that can separate the foreground (e.g., real-time 3D virtual cropping for the user) obtained from the live data feed 114 of the user 116 from the background. Sample image segmentation that can be used in embodiments of this disclosure may include, for example, watershed transform algorithms available from OpenCV.
[0184] Suitable image segmentation processes that can be used for background removal in this disclosure employ artificial intelligence (AI) techniques, such as computer vision, to achieve such background removal and may include instance segmentation and / or semantic segmentation. Instance segmentation assigns a distinct label to each individual instance of one or more multi-object classes. In some examples, instance segmentation is performed using a masked R-CNN, such as one fed from user live data to detect objects in an image while generating a high-quality segmentation mask for each instance, with an additional branch added for predicting object masks, which runs in parallel with an existing branch for bounding box recognition. The segmentation masks created for the user and the background are then extracted, and the background can be removed. Semantic segmentation uses deep learning or deep neural network (DNN) techniques to implement an automatic background removal process. Semantic segmentation divides an image into semantically meaningful parts by assigning each pixel a category label from one or more categories, such as color, texture, and smoothness, according to predefined rules. In some examples, semantic segmentation can leverage end-to-end, pixel-to-pixel semantic segmentation trained with fully convolutional networks (FCNs), as disclosed in the paper "Fully Convolutional Networks for Semantic Segmentation" by Evan Shelhamer, Jonathan Long, and Trevor Darrell in IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 39, No. 4 (April 2017), which is incorporated herein by reference. Following the aforementioned background removal process, a point cloud within the user's face and body boundaries can be preserved. This point cloud can be processed by one or more processors of the client device 118 or at least one cloud server computer 102 to generate a 3D mesh or 3D point cloud for the user, which can be used in the construction of the user's real-time 3D virtual cutout. The user's real-time 3D virtual cutout is then updated from the live data feed 114 from the camera 112.
[0185] In some embodiments, updating the user graphical representation 120 involves applying machine vision algorithms to the generated user graphical representation 120 by the client device 118 or at least one cloud server computer 102 to recognize facial expressions of the user 116 and graphically simulate facial expressions on the user graphical representation 120 within the virtual environment 110. Generally, such facial expression recognition can be performed using the principles of affective computing, which handles the recognition, interpretation, processing, and simulation of human emotions. Daniel Canedo and António J.R. Neves provide a review of conventional facial expression recognition (FER) techniques in “Facial Expression Recognition Using Computer Vision: A Systematic Review” in *Applied Sciences*, Vol. 9, No. 21 (2019), which is incorporated herein by reference.
[0186] Conventional FER techniques include steps of image acquisition, preprocessing, feature extraction, and classification or regression. In some embodiments of this disclosure, image acquisition is performed by feeding image data from a camera feed 114 to one or more processors. Preprocessing steps are often necessary to provide the most relevant data to the feature classifier and typically include face detection techniques capable of creating bounding boxes that define the target user's face, which are desired regions of interest (ROIs). ROIs are preprocessed through intensity normalization for illumination variations, noise filtering for image smoothing, data augmentation to increase training data, rotation correction for rotated faces, image resizing for different ROI sizes, and image cropping with better background filtering. After preprocessing, the algorithm retrieves relevant features from the preprocessed ROI, including action units (AUs), motion of certain facial landmarks, distances between facial landmarks, facial texture, gradient features, etc. These features can then be fed into a classifier, which may be, for example, a support vector machine (SVM) or a convolutional neural network (CNN). After training the classifier, emotions can be detected in real-time in the user and constructed in a user graphical representation 120 by, for example, connecting all facial feature relationships.
[0187] In some embodiments, the user graphical representation is associated with a top-down view, a third-person view, a first-person view, or a self-view. In an embodiment, the user 116's viewpoint when accessing a virtual environment through the user graphical representation is a top-down view, a third-person view, a first-person view, a self-view, or a broadcast camera view. The self-view may include the user graphical representation as seen by another user graphical representation, and optionally, the virtual background of the user graphical representation.
[0188] In some embodiments, the viewpoint is updated when the user 116 manually navigates the virtual environment 110 via a graphical user interface.
[0189] In other embodiments, the viewpoint is automatically established and updated using a virtual camera, wherein the viewpoint fed by live data is associated with the viewpoint of the user graphical representation and the virtual camera, and wherein the virtual camera is automatically updated by tracking and analyzing user eye and head tilt data or head rotation data, or a combination thereof. In embodiments, the viewpoint is automatically established and updated using one or more virtual cameras, which are virtually placed and aligned in front of the user graphical representation 120, for example, in front of a video with or without background removed, or a user 3D virtual cut, or a user real-time 3D virtual cut. In one embodiment, the one or more virtual cameras can point outward from eye level. In another embodiment, two virtual cameras, one for each eye, can point outward from both eye level. In another embodiment, the one or more virtual cameras can point outward from the center of the head position in the user graphical representation. The viewpoint of the user 116 captured by camera 112 is associated with the viewpoint of the user graphical representation 120 and the associated virtual camera using computer vision, thereby manipulating the virtual camera.
[0190] The virtual camera provides a virtual representation of the user's graphical representation 120's viewpoint associated with the user's 116's viewpoint, allowing the user 116 to view an area of the virtual environment 110 that the user graphical representation 120 may be looking at from one of many viewpoints. The virtual camera updates automatically by tracking and analyzing the user's eye and head tilt data, or head rotation data, or a combination thereof. The virtual camera position can also be manually changed by the user 116 based on the viewpoint selected by the user 116.
[0191] A self-perspective is the viewpoint of another user graphical representation 120 (e.g., in a phone camera's "selfie mode") as seen by that user graphical representation 120 with its background removed. Alternatively, the self-perspective may include a virtual background for the user graphical representation 120 to understand the perception of the user 116 as seen by other participants. When a virtual background for the user graphical representation is included, the self-perspective can be set to an area surrounding the user graphical representation, which can be captured by a virtual camera and can produce a circle, square, rectangle, or any other suitable shape for constructing the self-perspective. For example, in a scenario where the user graphical representation 120 is virtually located in a house, behind the user, there may be windows from which trees can be seen; the self-perspective displays the user graphical representation and, alternatively, displays a background that includes the windows and trees.
[0192] In other embodiments, tracking and analysis of user eye and head tilt data, or head rotation data, or a combination thereof, includes using computer vision to capture and analyze viewing positions and orientations captured by at least one camera 112, thereby manipulating a virtual camera within a virtual environment 110. For example, such manipulation may include receiving and processing eye and head tilt data captured by at least one camera using computer vision methods; extracting viewing positions and orientations from the eye and head tilt data; identifying one or more coordinates of the virtual environment contained within the position and orientation from the eye tilt data; and manipulating the virtual camera based on the identified coordinates.
[0193] In some embodiments, instructions in the memory 106 of at least one cloud server computer 102 are also capable of performing data analysis of user activities within at least one virtual environment 110. Data analysis can be used for interactions, including participating in conversations with other users, interacting with objects within the virtual environment 110, making purchases, downloading, engaging with content, etc. Data analysis can utilize a variety of known machine learning techniques to collect and analyze data from interactions in order to perform recommendations, optimizations, predictions, and automation. For example, data analysis can be used for marketing purposes.
[0194] In some embodiments, at least one processor 104 of one or more cloud server computers 102 is further configured to implement the trading and monetization of content added in at least one virtual environment 110. At least one cloud server computer 102 may communicatively connect to an application and object library, where users can find, select, and insert content in at least one virtual environment via a suitable application programming interface (API). One or more cloud server computers 102 may further connect to one or more payment gateways capable of executing corresponding transactions. Content may include, for example, interactive applications or static or interactive 3D assets, animations, or 2D textures.
[0195] Figure 2A-2B Schematic diagrams depict the deployment of systems 200a and 200b that enable interaction in a virtual environment, which encompasses multiple vertical domains of a virtual environment platform.
[0196] Figure 2A A schematic diagram depicts the deployment 200a of a system for implementing interaction in a virtual environment according to an embodiment, the virtual environment including multiple vertical domains 202 of a virtual environment platform 108. Figure 2A Some components can refer to Figure 1 The same or similar elements may be used, and therefore the same reference numerals may be used.
[0197] Vertical domain 202 is associated with the contextual theme of the virtual environment and includes uses related to the contextual theme, such as the virtual environment vertical domain of meeting 204 as, for example, a virtual meeting room; the virtual environment vertical domain of work 206 as, for example, a virtual office space; the virtual environment vertical domain of learning 208 as, for example, a virtual classroom; and the virtual environment vertical domain of shopping 210 as, for example, a virtual store. Figure 2A Other vertical sectors not shown may include, for example, services such as banking, reservations (e.g., hotels, travel agencies, or restaurants) and government agency services (e.g., consulting on starting a new company for a fee); and entertainment (e.g., karaoke, event halls or arenas, theaters, nightclubs, sports stadiums, museums, cruise ships, etc.) and so on.
[0198] Each virtual environment vertical domain 202 may contain multiple available virtual environments 110 (e.g., virtual environment AL), each virtual environment having one or more available layouts and visual appearances associated with the context of the corresponding vertical domain 202. For example, virtual environment A of the virtual environment vertical domain 202 of meeting 204 may contain a conference table with seating, a whiteboard, and a projector. At least one cloud server computer may provide corresponding resources (e.g., memory, network, and computing power) to each virtual environment 110. Vertical domain 202 may be obtained from virtual environment platform 108, which may be accessed by one or more users 116 via client device 118 through graphical user interface 212. Graphical user interface 212 may be contained in a downloadable client application or web browser application, providing the application data and instructions required to execute the selected virtual environment 110, and enabling multiple interactions therein. Furthermore, each virtual environment 110 may include one or more human or artificial intelligence (AI) hosts or assistants that may provide the necessary data and / or services to assist users within the virtual environment through their corresponding user graphical representations. For example, human or AI banking service staff can assist virtual bank users by providing the necessary information in the form of presentations, tables, lists, etc., according to the user's request.
[0199] In some embodiments, each virtual environment 110 is a persistent virtual environment containing customized changes performed thereon, wherein the changes are stored in persistent storage on at least one cloud server computer 102. For example, returning to the example of virtual environment A, the seating arrangement around the table, the color of the walls, or even the size and capacity of the room can be modified to suit the user's needs or preferences. The changes performed can be saved in persistent storage and are available thereafter during subsequent sessions in the same virtual environment A. In some examples, the persistent storage of modifications implemented in virtual environment 110 may require a subscription fee to be paid to the room owner or host (e.g., via virtual environment platform 108 connected to a payment gateway).
[0200] In other embodiments, virtual environment 110 is a temporary virtual environment stored in temporary storage on at least one cloud server computer 102. In these embodiments, changes performed on virtual environment 110 may not be stored and therefore may not be available in future sessions. For example, the temporary virtual environment may be selected from predefined available virtual environments from different vertical domains 202 from virtual environment platform 108. Changes in the arrangement (such as decorations or modifications) may or may not be implemented, but if changes are implemented, these changes may be lost after the session ends.
[0201] In some embodiments, a composite of virtual environments 110 within one or more vertical domains 202 can represent a virtual environment cluster 214. For example, some virtual environment clusters 214 may contain hundreds or even thousands of virtual environments 110. To a user, the virtual environment cluster 214 can appear as part of the same system, where users can interact with each other or seamlessly access other virtual environments within the same virtual environment cluster 214. For example, virtual environments D and E from the virtual environment vertical domain of work 206, plus virtual environment B from the virtual environment vertical domain of meeting 204, can form a virtual environment cluster 214 representing a company. In this example, users can have two different workspaces, such as a game development room and a business development room, and a meeting room for video conferencing. Users from either the game development room or the business development room can meet in the meeting room and conduct private virtual meetings, while the rest of the staff can continue their current activities in their original workspaces.
[0202] In other examples, the virtual environment cluster 214 may represent a theater or event facility, where each virtual environment 110 represents an indoor or outdoor event area (e.g., an auditorium or event venue) where one or more performers are giving a live performance via their corresponding user graphical representations. For example, an orchestra and / or singers could hold a concert by having their performance recorded live via cameras and via their user graphical representations (e.g., via their user-generated 3D virtual cutouts). Each performer's user graphical representation can be interpolated into the corresponding three-dimensional coordinates of the stage where they can perform. Audiences can watch the performance from the auditorium via their corresponding user graphical representations and can engage in various interactions, such as virtual clapping, singing along to songs, virtual dancing, virtual jumping, or cheering.
[0203] In other examples, the virtual environment cluster 214 can represent a casino containing multiple game areas (e.g., blackjack, poker, roulette, and slot machine areas), token purchase areas, and activity rooms. Machines in each game area can be configured as casino applications to provide a user experience relevant to each game. The casino operator can include a corresponding user graphical representation 120 (such as s) or a real-time 3D virtual cutout of the user. The casino operator represented by s can be a real human operator or an AI program assisting the user in the virtual casino. Each casino game can be coupled to a payment gateway from the casino company operating the virtual casino, enabling payments to and from the user.
[0204] In other examples, virtual environment cluster 214 could represent a shopping mall with multiple floors, each floor containing multiple virtual environments such as shops, showrooms, public areas, food courts, etc. Each virtual room can be managed by a corresponding virtual room administrator. For example, each shop can be managed by a corresponding shop administrator. Salespeople can be available in each area as 3D live virtual avatars or real-time 3D virtual cutouts of users, and can be real people or AI assistants. In the current example, each virtual shop and restaurant in the sample food court can be configured to purchase goods online through a corresponding payment gateway and delivery system and have the goods delivered to the user's address.
[0205] In another example, virtual environment cluster 214 contains multiple virtual gathering areas of a virtual nightclub, where users can meet and socialize through their corresponding user graphical representations. For example, each virtual gathering area may contain different themes and associated music and / or decorations. In addition to talking and texting, some other interactions in the virtual nightclub may include, for example, virtual dancing or drinking, sitting in different lounge areas (e.g., a lounge or bar), etc. Furthermore, in this example, an indoor concert can be held in the virtual nightclub. For example, an electronic music concert could be played by a radio host (DJ) performing behind a virtual table on a stage, where the DJ can be represented by a 3D live virtual avatar or a user-represented real-time 3D virtual cutout. If the DJ is represented by a user-represented real-time 3D virtual cutout, the real-time movement of the DJ playing the audio mixing console can be projected onto the real-time 3D virtual cutout from a live data feed obtained by a camera capturing images of the DJ at the DJ's location (e.g., from the DJ's house or recording studio). In addition, each member of the audience can also be represented by their own user graphical representation, some of whom can be represented by 3D live virtual avatars, while others can be represented by real-time 3D virtual cutouts based on user preferences.
[0206] In other examples, virtual environment cluster 214 may represent a virtual karaoke entertainment facility containing multiple private or public karaoke rooms. Each private karaoke room may contain a virtual karaoke machine, a virtual screen, a stage, microphones, speakers, decorations, sofas, tables, and drinks and / or food. Users can select songs via the virtual karaoke machine, which can connect to a song database, trigger the system to play songs for the user, and project lyrics onto the virtual screen for the user to sing along with using their user graphical representation. Public karaoke rooms may also contain humans or AIDJs who select songs for users, call users to the stage, and mute or unmute users as needed to listen to the performance. Users can sing remotely from their client devices via microphones.
[0207] In other examples, the virtual environment cluster 214 could represent a virtual cruise ship comprising multiple zones, such as bedrooms, engine room, activity room, bow, stern, port side, starboard side, bridge, and multiple decks. Some zones may have humans or AI assistants attending to users through corresponding user graphical representations, such as providing additional information or services. If available, a virtual environment or simple graphical representation of the cruise ship's exterior could be provided, such as depicting the landscapes of islands, towns, or cities that might be visited upon arrival at a specific destination. Users can thus experience traveling on the high seas and discovering new places through their user graphical representations, while being able to interact virtually with each other.
[0208] In other examples, virtual environment cluster 214 may represent an esports stadium or arena containing multiple virtual environments, which represent sports fields, courts, or rooms where users can play games via their user graphical representation using appropriate input / output devices (e.g., computer keyboards, game controllers, etc.). The mechanics of each esports event may depend on the sport being played. The esports stadium or arena may include public areas where users can choose which sports areas to access. Available sports schedules may also be available, informing users which sports activities are available at what times.
[0209] Figure 2B This represents a deployment 200b of a virtual school 216, which combines multiple virtual environments from various vertical domains 202. The virtual school 216 includes four classrooms (e.g., classrooms A-D218-224), an auditorium 226, a sports area 228, a cafeteria 230, a teachers' lounge 232, a library 234, and a bookstore 236. Each virtual environment may contain virtual objects associated with that environment, represented by corresponding graphical representations.
[0210] For example, a virtual classroom (e.g., any of virtual classrooms AD 218-224) enables students to attend lectures and can be configured to allow students to participate in the class through various interactions (e.g., raising hands, content projection, presentations, expressing questions or contributions verbally or via text, etc.) and can provide teachers with special administrative permissions (e.g., allowing someone to speak, muting one or more students during a lecture, sharing content via a digital whiteboard, etc.). An auditorium allows speakers to deliver presentations or can host multiple events. Sports area 228 can be configured to allow students to play multiple esports using their corresponding user graphical representations. Cafeteria 230 allows students to order food online and socialize using user graphical representations. Teacher lounge 232 can be configured for teachers to meet and discuss agendas, student progress, etc., using their corresponding teacher user graphical representations. Library 234 allows students to borrow ebooks for their classes or leisure reading. Finally, bookstore 236 can be configured to allow students to purchase books (e.g., ebooks or physical books) and / or other school materials.
[0211] Figure 3 A schematic diagram of a sample hybrid system architecture 300, according to an embodiment, can be used in a system implementing interaction in a virtual environment. In some embodiments, the hybrid system architecture 300 is a hybrid communication model for interacting with other peer clients (e.g., other attendees in a virtual meeting, classroom, etc.), comprising a client-server side 304 and a P2P side 306, each in... Figure 3 The area is defined by a dashed line. Using this hybrid communication model enables fast P2P communication between users, reducing waiting time issues, while providing network services, data, and resources to each session, allowing for various interactions between users and with content in the virtual environment. Figure 3 Some components can refer to Figure 1-2A The same or similar elements may be used, and therefore the same reference numerals may be used.
[0212] In various embodiments, the level and ratio of client-server 304 to P2P client 306 usage depends on the amount of data to be processed, the latency allowed to maintain a smooth user experience, the desired quality of service (QoS), the required services, etc. In one embodiment, P2P client 306 is used for video and data processing, streaming, and rendering. This hybrid system architecture 300 may be suitable, for example, when low latency and low data volumes are required, and when there are “heavy” clients, meaning that the client devices contain sufficient computing power to perform such operations. In another embodiment, a combination of client-server 304 and P2P client 306 may be used, such as P2P client 306 for video streaming and rendering, while client-server 304 is used for data processing. This hybrid system architecture 300 may be suitable, for example, when there is a large amount of data to process, or when additional microservices may be required. In other embodiments, client-server 304 may be used for video streaming and data processing, while P2P client 306 is used for video rendering. For example, this mode of the hybrid system architecture 300 may be suitable when there is a larger volume of data to process and / or when only thin clients are available. In other embodiments, the client-server side 304 can be used for video streaming, rendering, and data processing. This mode of the hybrid system architecture 300 may be suitable when very thin clients are available. The hybrid system architecture 300 can be configured to alternate between different usage forms of the client-server side 304 and the P2P side 306 within the same session as needed.
[0213] In some embodiments, at least one cloud server from the client-server side 304 can be an intermediary server, meaning that the server is used to facilitate and / or optimize data exchange between client devices. In such embodiments, at least one cloud server can manage, analyze, process, and optimize incoming image and multimedia streams, and manage, evaluate, and / or optimize the forwarding of outbound streams as a router topology (e.g., but not limited to SFU (Selective Forwarding Unit), SAMS (Spatial Analysis Media Server), multimedia router, etc.), or can use an image and media processing server topology (e.g., for tasks including but not limited to decoding, combining, improving, mixing, enhancing, expanding, computing, manipulating, encoding) or a forwarding server topology (including but not limited to MCU, cloud media mixer, cloud 3D renderer, media server) or other server topologies.
[0214] In such embodiments, where the intermediate server is a SAMS, this media server manages, analyzes, and processes incoming data (e.g., metadata, priority data, data categories, spatial structure data, 3D location, orientation or motion information, images, media, video based on scalable video codecs, or combinations thereof) sent to each client device, and manages and / or optimizes the forwarding of outbound data streams to each receiving client device in such analysis. This may include modifying, scaling up, or scaling down the media for time (e.g., varying frame rates), space (e.g., different image sizes), quality (e.g., quality based on different compression or encoding), and color (e.g., color resolution and range), and may achieve optimal bandwidth and computational resource utilization for receiving one or more user client devices based on factors such as the space, 3D orientation, distance, and priority relationship of a particular receiving client device user with respect to such incoming data.
[0215] In some embodiments, media, video, and / or data processing tasks include one or more of encoding, code conversion, decoding, spatial or 3D analysis and processing, including one or more of image filtering, computer vision processing, image sharpening, background enhancement, background removal, foreground blurring, eye occlusion, face pixelation, speech distortion, image magnification, image cleansing, skeletal structure analysis, face or head counting, object recognition, marker or QR code tracking, eye tracking, feature analysis, 3D mesh or volume generation, feature tracking, face recognition, SLAM tracking, and facial expression recognition, or other modular plug-ins in the form of microservices running on such media routers or servers.
[0216] The client-server side 304 employs a secure communication protocol 308 to achieve secure end-to-end communication between the client device 118 and the network / application server 310 over the network. A suitable secure communication protocol 308 may include, for example, Datagram Transport Layer Security (DTLS), which is itself compatible with Secure User Datagram Protocol (UDP), Secure Real-Time Transport Protocol (SRTP), Hypertext Transfer Protocol Security (HTTPS), and Network Sockets Security (WSS: / / ), providing full-duplex authentication for application access, privacy protection, and the integrity of data exchanged in transit. A suitable network / application server 310 may include, for example, a Jetty web application server, which is a Java HTTP web server and Java Servlet container, enabling proper deployment of machine-to-machine communication and web application services.
[0217] Although the network / application server 310 is Figure 3While depicted as a single element, those skilled in the art will understand that the web server and application server can be separate elements. For example, the web server can be configured to receive client requests via a secure communication protocol 308 and route the requests to the application server. The web / application server 310 can thus receive and process client requests using the secure communication protocol 308, which may include requests for one or more microservices 312 (e.g., Java-based microservices) and / or retrieving data from a database 314 using a corresponding database management system 316. The application / web server 310 can provide session management and many other services, such as 3D content and application logic, as well as session state persistence (e.g., for persistent storage of shared documents, synchronizing interactions and changes in a virtual environment, or maintaining the visual state and modifications of a virtual environment). A suitable database management system 316 could be, for example, an object-relational mapping (ORM) database management system, which, given the ability of ORMs to transform data between incompatible type systems using object-oriented programming languages, may be suitable for database management using both open-source and commercial (e.g., proprietary) services. In other embodiments, by using a publish-subscribe model, the distributed spatial data bus 318 can be further used as a distributed messaging and resource distribution platform between microservices and client devices.
[0218] The P2P client 306 can use a suitable P2P communication protocol 320, which enables real-time communication between peer client devices 118 in a virtual environment through a suitable application programming interface (API), allowing for real-time interaction and synchronization, thus enabling a multi-user collaborative environment. For example, through the P2P client 306, contributions from one or more users can be directly transmitted to other users, who can observe the changes made in real time. An example of a suitable P2P communication protocol 320 could be the Web Real-Time Communication (WebRTC) protocol, a collection of standards, protocols, and JavaScript APIs that combine to enable the sharing of P2P audio, video, and data between peer client devices 118. The client devices 118 in the P2P client 306 can employ one or more rendering engines 322 to perform real-time 3D rendering for the live session. An example of a suitable rendering engine 322 could be a WebGL-based 3D engine. WebGL is a JavaScript API for rendering 2D and 3D graphics within any compatible web browser without the use of plugins, allowing one or more processors (e.g., one or more graphics processing units (GPUs)) of the client device 118 to accelerate physics and image processing and the use of effects. Furthermore, the client device 118 in the P2P end 306 can perform image and video processing and machine learning computer vision techniques via one or more suitable computer vision libraries 324. In one embodiment, the image and video processing performed by the client device in the P2P end 306 includes a background removal process used in the creation of the user's graphical representation before inserting it into the virtual environment. This process can be performed in real-time or near real-time on the received media stream, or non-real-time on, for example, a photograph. An example of a suitable computer vision library 324 could be OpenCV, a programming function library configured primarily for real-time computer vision tasks.
[0219] Figure 4 A schematic diagram of a graphical user interface 400 of a virtual environment live session module 402 according to an embodiment is depicted, through which a user can interact in a virtual environment.
[0220] Before a user can access the graphical user interface 400 of the virtual environment live session module 402, the user can first receive an invitation from a peer client device to participate in a dialogue with a peer user. This can open a P2P communication channel between user client devices when processing and rendering are performed by the client device, or alternatively, can open an indirect communication channel through the cloud server computer when processing and rendering are performed by at least one cloud server computer. Furthermore, as will be discussed later... Figure 5As shown in the description, transitions can occur from user 3D virtual cut to user real-time 3D virtual cut, or a video with the background removed, or a video without the background removed.
[0221] The virtual environment live session module 402 may include a virtual environment screen 404, which includes a graphical user interface (GUI) displaying the selected virtual environment. This GUI may include the arrangement of the virtual environment associated with a selected vertical domain of the virtual environment, as well as corresponding virtual objects, applications, other user graphical representations, etc. The GUI 400 of the virtual environment live session module 402 can implement and display multiple interactions 406, configured for users to interact with each other, for example, through their real-time 3D virtual cutaways. The virtual environment live session module 402 may include one or more data models associated with a corresponding task for implementing each interaction 406, plus the computer instructions required to implement said task. Each interaction 406 may be represented in different ways; in Figure 4 In the example shown, each interaction 406 is represented as a button on the graphical user interface 400 of the virtual environment live session module 402, where clicking each interaction button requests the corresponding service to perform the task associated with interaction 406. The virtual environment live session module 402 can be referenced, for example, through... Figure 3 The publicly available hybrid system architecture 300 is used to implement this.
[0222] Interaction 406 may include, for example, chat 408, screen sharing 410, host options 412, remote sensing 414, recording 416, voting 418, document sharing 420, sending emoticons 422, agenda sharing and editing 424, or other interactions 426. Other interactions 426 may include, for example, virtual hugs, raising hands, handshakes, walking, adding content, preparing meeting summaries, moving objects, projection, laser pointing, playing games, purchasing, and other social interactions that facilitate communication, competition, cooperation, and conflict resolution among users. The various interactions 406 are described in more detail below.
[0223] Chat 408 opens a chat window, allowing you to send and receive text comments and instant resources.
[0224] Screen sharing 410 enables users to share their screens in real time with any other participant.
[0225] Host option 412 is configured to provide the conversation host with additional options, such as muting one or more users, inviting or removing one or more users, ending the conversation, etc.
[0226] Remote Sensing 414 can view a user's current status, such as whether they are not present, busy, available, offline, in a conference call, or in a meeting. User status can be updated manually via a graphical user interface or automatically using machine vision algorithms based on data received from camera feeds.
[0227] Recorder 416 can record audio and / or video from conversations.
[0228] Voting 418 allows for voting on one or more suggestions published by any other participant. Through Voting 418, the host or other participants with such permissions can initiate a voting session at any time. Topics and options can be displayed for each participant. Depending on the configuration of the Voting 418 interaction, the results can be shown to all attendees at the end of the timeout period or at the end of each person's response.
[0229] Document sharing 420 enables the sharing of documents in any suitable format with other participants. These documents can also be permanently stored in persistent storage on one or more cloud server computers and can be associated with the virtual environment in which the virtual communication takes place.
[0230] Emoji Sending 422 enables sending emojis to other participants.
[0231] Agenda sharing and editing 424 enables the sharing and editing of agendas that may have been prepared by any participant. In some embodiments, the host can configure a checklist of agenda items before the meeting. Hosts or other participants with such permissions can bring the agenda to the forefront at any time. Through the agenda editing options, items can be deprecated or postponed when consensus is reached.
[0232] Other interactions 426 provides a non-exhaustive list of possible interactions that can be provided within the virtual environment based on its vertical orientation. Raising a hand enables users to raise their hands during virtual communications or meetings, allowing hosts or other participants with such rights to speak. Walking enables users to move within the virtual environment via real-time 3D virtual clipping. Content addition allows users to add interactive applications or static or interactive 3D assets, animations, or 2D textures to the virtual environment. Meeting summary preparation enables the automatic preparation of virtual meeting results and their distribution to participants at the end of the session. Object movement enables the movement of objects within the virtual environment. Projection enables the projection of content from attendees' screens onto screens or walls available in the virtual environment. Laser pointing enables the use of laser pointing to highlight desired content during presentations. Playing games enables the playing of one or more games or other types of applications that can be shared during a live session. Purchasing enables the purchase of content during the session. Other interactions not mentioned herein may also be configured according to the specific purpose of the virtual environment platform.
[0233] In some embodiments, the system can also enable the creation of self-organizing virtual communication, which may include creating self-organizing voice communication channels between user graphical representations without changing the current viewpoint or location within the virtual environment. For example, a user graphical representation can approach another user graphical representation and engage in a self-organizing voice conversation at the location where the two user graphical representations' areas are within the virtual environment. Such communication will be achieved, for example, by considering the distance, location, and orientation between the user graphical representations, and / or their current availability status (e.g., available or unavailable) or the state configuration of such self-organizing communication, or a combination thereof. In this example, the approaching user graphical representation will see visual feedback on the other user graphical representation, signaling that self-organizing communication is possible, and thus setting the start of a conversation between the two user graphical representations, where the approaching user can speak and the other user can hear and respond. In another example, a user graphical representation can approach another user graphical representation, click on the user graphical representation, send a conversation invitation, and, after the invitee's approval, engage in a self-organizing voice conversation at the location within the virtual environment where the two user graphical representations' areas are located. Depending on the privacy settings between the two user graphical representations, other users can view the interactions, facial expressions, hand gestures, etc., between the user graphical representations, regardless of whether they can hear the conversation. Any of the aforementioned interactions or other interactions 426 can also be performed directly within the virtual environment screen 404.
[0234] Figure 5 A method 500 according to an embodiment is described, which enables a transformation from one type of user graphical representation to another type of user graphical representation, such as from user 3D virtual clipping to user real-time 3D virtual clipping, or a transformation to a video with or without background removal.
[0235] This transition can be achieved when a user engages in a dialogue with another user's graphical representation. For example, a user might currently be sitting in an office chair and working on a computer in a virtual office. The user's current graphical representation could be a graphical representation of the user in a 3D virtual cutout. At this point, the camera might not be on because a live data feed from the user might not be needed. However, if the user decides to turn on the camera, the user's 3D virtual cutout could include facial expressions provided through facial analysis captured from the user's live data feed, as explained in more detail herein.
[0236] When a user engages in a conversation with another user's graphical representation and initiates a live session, if the user's camera is not active, it can be activated and a live data feed capture can be initiated. This capture can provide the user's live stream, thereby transforming the user's 3D virtual cut into a real-time 3D virtual cut or a video with or without background removal. Further, as... Figure 1 As described, the live stream of a user's real-time 3D virtual clipping 504 can be processed and rendered on the client or server, or it can be sent to other peer client devices in a P2P or hybrid system architecture for their own real-time processing and rendering (e.g., via reference). Figure 3 The hybrid system architecture described (300).
[0237] Figure 5 Method 500 can begin at step 502, approaching a user graphical representation. Then, in step 504, method 500 can continue by selecting and clicking on the user graphical representation. In step 506, method 500 can continue by the client device sending a dialogue participation invitation to another user graphical representation or receiving a dialogue participation invitation from another user graphical representation. In step 508, method 500 continues by the corresponding client device accepting the received invitation. Then, method 500 continues in step 510, transitioning from user 3D virtual clipping to user real-time 3D virtual clipping or video with or without background removal. Finally, in step 512, method 500 ends by opening a P2P communication channel between user client devices when processing and rendering are performed by the client device, or by opening an indirect communication channel through the cloud server computer when processing and rendering are performed by at least one cloud server computer. In some embodiments, the dialogue includes sending and receiving real-time audio and video displayed from the participants' user real-time 3D virtual clipping.
[0238] Figures 6A-6C A schematic diagram depicts a combination of multiple image processing operations performed by the corresponding client device 118 and cloud server 102 on the client-server side 304. The client-server side can be, for example, part of a hybrid system architecture, such as... Figure 3 The hybrid system architecture 300 is described in the text.
[0239] exist Figures 6A-6C In one embodiment, at least one cloud server 102 may be configured to use a relay network address translation (NAT) traversal (sometimes called TURN) server, which may be suitable for situations where the server cannot establish a connection between client devices 118. TURN is an extension of the NAT session traversal tool (STUN).
[0240] NAT is a method of remapping the Internet Protocol (IP) address space to another address space by modifying the network address information in the IP header of a data packet as it is transported through a service routing device. Therefore, NAT can provide access to a network (such as the Internet) for a private IP address and allow a single device (such as a routing device) to act as a proxy between the Internet and the private network. NAT can be symmetric or asymmetric. A framework called Interactive Connection Establishment (ICE) is configured to find the best path for connecting client devices, and this framework determines whether symmetric or asymmetric NAT is needed. Symmetric NAT is responsible not only for translating IP addresses from private to public addresses and vice versa, but also for port translation. Asymmetric NAT, on the other hand, uses a STUN server, allowing clients to discover their public IP address and the type of NAT behind them, which can be used to establish a connection. In many cases, STUN may only be used during connection setup, and once the session is established, data can begin flowing between client devices.
[0241] TURN can be used in the case of symmetric NAT and remains in the media path after the connection is established, while processed and / or unprocessed data is relayed between client devices.
[0242] Figure 6A A client-server error 304 is described, involving client device A, cloud server 102, and client device B. Figure 6A In this context, client device A is the sender of the data to be processed, and client device B is the receiver of the data. Several image processing tasks are described and categorized based on whether they are performed by client device A, cloud server 102, or client device B, and are thus classified as client device A processing 602, server image processing 604, and client device B processing 606.
[0243] The image processing task includes background removal 608, further processing or improvement 610, and insertion and combination into the virtual environment 612. (As from...) Figure 6B and 6C And from Figure 7B It will become apparent that the combination of the three image processing tasks illustrated in this paper can be used to generate, improve, and insert / combine user graphical representations into virtual environments. Furthermore, for simplicity, in Figures 6B-6C and Figures 7B-7C In the text, background removal 08 is described as "BG" 608, further processing or improvement 610 is described as "++" 610, and insertion and combination into the virtual environment 612 is described as "3D" 612.
[0244] In some embodiments, inserting and combining a user graphical representation into a virtual environment includes generating one or more virtual cameras that are virtually placed and aligned in front of the user graphical representation, for example, in front of a video with or without background removed, a user 3D virtual cutout, or a user real-time 3D virtual cutout. In one embodiment, the one or more virtual cameras may point outward from eye level. In another embodiment, two virtual cameras, one for each eye, may point outward from both eye levels. In another embodiment, the one or more virtual cameras may point outward from the center of the user graphical representation's head position. In another embodiment, the one or more virtual cameras may point outward from the center of the user graphical representation. In another embodiment, the one or more virtual cameras may be placed in front of the user graphical representation, for example, at head level, pointing towards the user graphical representation from a self-perspective. The one or more virtual cameras are created at least by using computer vision to correlate captured user perspective data with the perspective of the user graphical representation within the virtual environment. The one or more virtual cameras are automatically updated by tracking and analyzing user eye and head tilt data, or head rotation data, or a combination thereof, and may also be manually changed by the user based on a user-selected perspective.
[0245] The combination of image processing by client device A (processing 602), server image processing (processing 604), and client device B (processing 606), and the corresponding usage level, depend on the amount of data to be processed, the allowed latency to maintain a smooth user experience, the expected quality of service (QoS), the required services, etc.
[0246] Figure 6B Describe the image processing combination 1-4.
[0247] In image processing assembly 1, client device A generates a user graphic representation including background removal 608, and sends the background-removed user graphic representation to at least one cloud server 102 for further processing or improvement 610, generating an enhanced user graphic representation with background removal. The at least one cloud server sends the background-removed enhanced user graphic representation to client device B, which then inserts and combines the background-removed enhanced user graphic representation into the virtual environment.
[0248] In image processing assembly 2, client device A generates a user graphic representation including background removal 608, performs further processing or improvement 610 on it to generate an enhanced user graphic representation with the background removed, and then sends it to at least one cloud server 102. At least one cloud server 102 sends the enhanced user graphic representation with the background removed to client device B, and client device B inserts and combines the enhanced user graphic representation with the background removed into the virtual environment.
[0249] In image processing assembly 3, client device A generates a user graphic representation including background removal 608, performs further processing or improvement 610 on it to generate an enhanced user graphic representation with the background removed, and inserts and combines the enhanced user graphic representation with the background removed into the virtual environment. Client device A then sends the enhanced user graphic representation with the background removed and combined into the virtual environment to the cloud server for relay to client device B.
[0250] In image processing assembly 4, client device A generates a user graphic representation including background removal 608, and sends the background-removed user graphic representation to at least one cloud server 102 for further processing or improvement 610 to generate an enhanced user graphic representation with background removal. The at least one cloud server then inserts and combines the enhanced user graphic representation with background removal into a virtual environment, and then sends it to client device B.
[0251] Figure 6C Describe the image processing combination 5-8.
[0252] In image processing assembly 5, client device A generates a user graphic representation including background removal 608 and sends the background-removed user graphic representation to at least one cloud server 102 for relay to client device B. Client device B performs further processing or improvement 610 on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, which is then inserted and combined into the virtual environment.
[0253] In image processing assembly 6, client device A sends a camera live data feed received from at least one camera and unprocessed data to at least one cloud server 102. The at least one cloud server performs the generation of a user graphic representation including background removal 608, and performs further processing or improvement 610 on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, which is then sent to client device B. Client device B inserts and combines the enhanced background-removed user graphic representation into the virtual environment.
[0254] In image processing assembly 7, the client device sends a camera live data feed received from at least one camera and sends unprocessed data to at least one cloud server 102. At least one cloud server 102 generates a user graphic representation including background removal 608, performs further processing or improvement 610 on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, and then inserts and combines the enhanced background-removed user graphic representation into the virtual environment sent to the client device B.
[0255] In image processing assembly 8, client device A sends a camera live data feed received from at least one camera to at least cloud server 102 and sends unprocessed data for relay to client device B. Client device B uses this data to generate a user graphic representation including background removal 608, and performs further processing or improvement 610 on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, which is then inserted and combined into the virtual environment. In some embodiments, it may be understood that at least cloud server 102 may be an intermediary server, meaning that the server uses an intermediary server topology to facilitate and / or optimize data exchange between client devices.
[0256] In such embodiments, at least one cloud server may be an intermediary server, meaning that the server is used to facilitate and / or optimize data exchange between client devices. In such embodiments, at least one cloud server may manage, analyze, and optimize incoming multimedia streams, and manage, evaluate, and optimize the forwarding of outbound streams as a router topology (e.g., SFU, SAMS, multimedia router, etc.), or media processing (e.g., performing tasks including decoding, combining, improving, mixing, enhancing, expanding, computing, manipulating, or encoding) or forwarding server topology (e.g., but not limited to multipoint control units, cloud media mixers, cloud 3D renderers) or other server topologies.
[0257] In such embodiments, where the intermediate server is a SAMS, this media server manages, analyzes, and processes incoming data (e.g., metadata, priority data, data categories, spatial structure data, 3D location, orientation or motion information, images, media, or video based on a scalable video codec) from sending client devices, and manages or optimizes the forwarding of outbound data streams to receiving client devices during such analysis. This may include modifying, scaling up, or down the media based on factors such as the space, 3D orientation, distance, and priority relationship of the particular receiving client device user with respect to such incoming data, for time (e.g., varying frame rates), space (e.g., different image sizes), quality (e.g., quality based on different compression or encoding), and color (e.g., color resolution and range), thereby achieving optimal bandwidth and computational resource utilization for receiving one or more user client devices.
[0258] Intermediate server topology might be suitable for, for example, image processing combinations 1-8, where at least one cloud server 102 processes data between client devices A and B, such as... Figures 6A-6C As shown.
[0259] Figures 7A-7C A schematic diagram depicts a combination of multiple image processing operations performed by the corresponding client device in the P2P terminal 306. Figures 7A-7B In this context, it is described as peer-to-peer devices (AB) to distinguish it from situations where communication and processing occur via a client-server architecture. A P2P endpoint (306) can be, for example, part of a hybrid system architecture, such as... Figure 3 The hybrid system architecture 300 is described in the text.
[0260] Figure 7A A P2P endpoint 306 is depicted, comprising peer device A and peer device B, where peer device A is the sender of data to be processed, and peer device B is the receiver of that data. Multiple image and media processing tasks are depicted and categorized based on whether they are performed by peer device A or peer device B, and are thus classified as peer device A processing 702 and peer device B processing 704. Image and media processing tasks may include (but are not limited to) background removal 608, further processing or improvement 610, and insertion and combination into a virtual environment 612.
[0261] Figure 7B Describe the image processing combination 1-3.
[0262] In image processing assembly 1, peer device A generates a user graphic representation including background removal 608, performs further processing or improvement 610 on it to generate an enhanced user graphic representation with background removed, and inserts and combines the enhanced user graphic representation with background removed into a virtual environment with three-dimensional coordinates. Peer device A then sends the enhanced user graphic representation with background removed and combined into the virtual environment to peer device B.
[0263] In image processing assembly 2, peer device A generates a user graphic representation including background removal 608 and sends the background-removed user graphic representation to peer device B. Peer device B performs further processing or improvement 610 on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, which is then inserted and combined into the virtual environment.
[0264] In image processing assembly 3, peer device A sends a camera live data feed received from at least one camera to peer device B and sends encoded data. Peer device B decodes and uses the data to generate a user graphic representation including background removal 608, and performs further processing or improvement 610 on the background-removed user graphic representation to generate an enhanced user graphic representation with the background removed, and then inserts and combines the enhanced user graphic representation with the background removed into the virtual environment.
[0265] Figure 7C Describe the image processing combination 4-6.
[0266] exist Figure 7CIn one embodiment, at least one cloud server 102 can be configured as a STUN server, which allows peer devices to discover their public IP addresses and the NAT types behind them, information that can be used to establish data connections and exchange data between peer devices. Figure 7C In another embodiment, at least one cloud server 102 may be configured for signaling, which may be used for peer device location and connection to each other, and for exchanging data through communication coordination performed by at least one cloud server.
[0267] In all image and processing combinations 4-6, at least one cloud server 102 can use SAMS, SFU, MCU or other functional server topologies because at least one cloud server 102 serves between peer devices A and B.
[0268] In image processing assembly 4, peer device A generates a user graphic representation including background removal 608, performs further processing or improvement 610 on it to generate an enhanced user graphic representation with the background removed, and inserts and combines the enhanced user graphic representation with the background removed into the virtual environment. Peer device A then sends the enhanced user graphic representation with the background removed, which has been inserted and combined into the virtual environment, to peer device B via at least one cloud server acting as a STUN or signaling server.
[0269] In image processing assembly 5, peer device A generates a user graphic representation including background removal 608 and sends the background-removed user graphic representation to peer device B via at least one cloud server acting as a media router server. Peer device B performs further processing or improvement 610 on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, which client device B then inserts and combines into the virtual environment.
[0270] In the image processing assembly 6, peer device A sends a camera live data feed received from at least one camera and transmits unprocessed data to peer device B via at least one cloud server acting as a STUN or signaling server. Peer device B uses this data to generate a user graphic representation including background removal 608, and performs further processing or improvement 610 on the background-removed user graphic representation to generate an enhanced background-removed user graphic representation, which is then inserted and combined into the virtual environment.
[0271] Figure 8 This disclosure demonstrates a user authentication system 800 based on user graphical representation that can be used in embodiments of this disclosure. For example, the user authentication system 800 based on user graphical representation can be used to access user accounts authorized to access a virtual environment platform, such as... Figure 1 and Figure 2A The virtual environment platform 108.
[0272] A user authentication system 800 based on user graphical representation includes one or more cloud server computers 802, each containing at least one processor 804 and a memory 806 storing data and instructions. The memory includes a user database 808 storing user data associated with user accounts 810 and one or more corresponding user graphical representations 812. The user authentication system 800 also includes a face scanning and authentication module 814 connected to the database 808 storing data associated with user accounts 810. The one or more cloud server computers 802 are configured to authenticate users by performing a face scan via the face scanning and authentication module 814. The face scan includes extracting facial feature data from camera data received from a client device 822 and checking for a match between the extracted facial feature data and the user graphical representations in the user database 808.
[0273] exist Figure 8 In the example shown, system 800 also includes at least one camera 816 configured to acquire image data 818 from user 820 of at least one client device 822 requesting access to user account 810. The at least one camera 816 is connected to at least one client device 822, which is configured to transmit the data captured by camera 816 to one or more cloud server computers 802 for further processing. Alternatively, camera 816 may be directly connected to one or more cloud server computers 802. One or more cloud server computers 802 are configured to perform a facial scan of the user via a facial scanning and authentication module 814, check for a match between the user database 808 and existing user graphic representations, and authenticate the user by providing the corresponding user graphic representation 812 and access to user account 810 if user account 810 is confirmed and available. Alternatively, if user account 810 is unavailable, one or more cloud server computers 802 are configured to authenticate the user by generating a new user graphic representation 812 and a new user account 810 stored in user database 808 using data 818 obtained from a live data feed.
[0274] User account 810 can be used, for example, to access a virtual environment platform or any other application (e.g., an application that can be linked to the environment platform), such as any interactive application, game, email account, university profile account, work account, etc. Given steps such as generating a user graphical representation 812 or retrieving an existing user graphical representation 812 from a user database 808, the graphical representation-based user authentication system 800 of this disclosure provides a higher level of convenience and security than standard camera-based face detection authentication systems.
[0275] In some embodiments, one or more cloud server computers are also configured to check the date of a matching user graphic representation and determine whether the matching user graphic representation needs to be updated. In an embodiment, if user account 810 is available, and in response to one or more cloud server computers 802 checking the date of available user graphic representations 812, the one or more cloud server computers 802 determine whether an existing user graphic representation 814 needs to be updated by comparing it with a corresponding threshold or security requirement. For example, if a system security update is required, it may be necessary to update all user graphic representations, or at least to update graphic representations created before a specified date. If user graphic representation 814 is required, the one or more cloud server computers 802 generate a user graphic representation update request to the corresponding client device 822. If user 820 approves the request, the one or more cloud server computers 802 or client device 822 continue to generate user graphic representation 814 based on data 818 fed from the live camera. If no update is required, after authentication, the one or more cloud server computers 802 continue to retrieve existing user graphic representations 812 from the user database 808.
[0276] In some embodiments, the user graphical representation 812 is inserted into a two-dimensional or three-dimensional virtual environment, or onto a third-party source linked to the virtual environment, and combined with the two-dimensional or three-dimensional virtual environment. For example, the user graphical representation 812 can be inserted onto a third-party source linked to the virtual environment by overlaying it on the screen of a third-party application or website that is integrated or coupled to the system of this disclosure.
[0277] In one example, overlaying a user graphical representation 812 on the screen of a third-party source is done on top of a 2D website or application linked to a virtual environment. For example, two or more friends visiting a shopping website together can overlay their user graphical representations on the shopping website to explore the website's content and / or interact with it. In another example, overlaying a user graphical representation 812 on the screen of a third-party source is done on top of a 3D gaming session linked to a virtual environment. For example, a user can access an esports gaming session linked to a virtual environment through his or her user graphical representation 812, which can be overlaid on top of the esports gaming session along with the user graphical representations 812 of other team members. In these examples, such overlays of user graphical representations 812 can enable a coherent and multi-point delivery view of the expressions and communications of all users during a visit to a 2D website or an experience of a 3D gaming session.
[0278] In some embodiments, the generation of the user graphical representation 812 occurs asynchronously with the user 820 accessing the user account 810. For example, if the user graphical representation-based authentication system 800 determines that the user 820 has already been authenticated after performing a facial scan, the user graphical representation-based authentication system 800 can enable the user 820 to access the user account 810 while a new user graphical representation 812 is being generated, so that it can be provided to the user 812 once it is ready and then inserted and combined into the virtual environment.
[0279] In some embodiments, one or more cloud server computers 802 further authenticate user 802 through login authentication credentials, which include a personal identification number (PIN), or a username and password, or a combination of camera authentication and PIN, or a username and password.
[0280] In some embodiments, the user-graphical representation-based authentication system 800 triggers authentication in response to activation of an invitation link or deep link sent from one client device 822 to another client device. Clicking the invitation link or deep link triggers at least one cloud server computer 802 to request user authentication. For example, the invitation link or deep link can be used for telephone calls, conference calls, or video game session invitations, where the invited user can be authenticated through the user-graphical representation-based authentication system 800.
[0281] In another embodiment, facial scanning uses 3D authentication, which involves guiding the user to perform head movement patterns and extracting 3D facial data based on these patterns. This can be accomplished using application instructions stored in at least one server computer, which guide the user to perform head movement patterns to achieve 3D authentication, such as performing one or more head poses, tilting or rotating the head horizontally or vertically in a circular motion, performing user-generated pose patterns, or specific head movement patterns, or combinations thereof. 3D authentication identifies additional features based on data received from a live video feed from a camera, rather than simply comparing and analyzing a view or image. In this 3D authentication embodiment, the facial scanning process can identify additional features from data that may include facial data, including head movement patterns, facial volume, height, depth of facial features, facial scars, tattoos, eye color, facial skin parameters (e.g., skin color, wrinkles, pore structure, etc.), reflection parameters, and, for example, simply the location of such features on the facial topology, which may be the case in other types of facial detection systems. Capturing such facial data thus increases the capture of a realistic face, which can be used to generate a realistic graphical representation of the user. Facial scans for 3D authentication can be performed using high-resolution 3D cameras, depth cameras (e.g., LiDAR), light field cameras, etc. The facial scanning process and 3D authentication can utilize deep neural networks, convolutional neural networks, and other deep learning techniques to retrieve, process, and evaluate the user's authentication using facial data.
[0282] Figure 9 A schematic diagram of a third-person perspective 900 of a virtual environment 110, depicted by user graphical representation 120, is shown, where the virtual environment 110 is a virtual office.
[0283] The virtual office includes one or more desks 902, office chairs 904, office computers 906, projection surfaces 908 for projecting content 910, and multiple user graphical representations 120 representing corresponding users accessing the virtual environment 110 through their client devices.
[0284] The user graphical representation 120 can initially be a user 3D virtual cutout, and can be transformed into a user real-time 3D virtual cutout after an invitation approval process, comprising a user real-time video stream with background removed, or a video with background removed, or a video without background removed, based on a feed of real-time 2D or 3D live video stream data obtained from a camera. (See reference...) Figure 5 The process may include opening a communication channel to enable multiple interactions within a live session, as described in the reference. Figure 4For example, a user might initially sit in office chair 904 and work on the corresponding office computer 906, which could represent the actual actions the user is performing in real life. Other users could be able to view this (e.g., through...). Figure 4 The remote sensing (414) provides information on the current user status, such as whether the user is absent, busy, available, offline, in a conference call, or in a meeting. If the user is available, another user's graphical representation can approach the user being discussed and send an invitation to participate in the conversation. For example, both users could decide to move to a private meeting room in a virtual office and begin a live session with multiple interactions. Users can also be able to project desired content onto the projection surface 908 (e.g., via screen sharing).
[0285] In some embodiments, the virtual office also includes virtual computers comprising virtual resources from one or more cloud computing resources, which are accessed via client devices and assigned to the virtual computer resources via management tools. The virtual computers may be associated with office computer 906. However, virtual computers may also be associated with personal home computers or computers from any other location that have access to cloud-based virtual computing resources. Resources may include storage, networking, and processing power required to perform various tasks. Furthermore, in the example of the office space, the virtual computer associated with virtual office computer 906 may then be coupled to a user's physical office computer, such that, for example, when a user logs into such a virtual computer, data stored in virtual office computer 906 can be obtained from the physical office computer in the physical office or any other space with a physical computer. The virtual infrastructure including all virtual computers associated with virtual office computer 906 can be managed through a virtual environment platform using administrator options based on exclusive administrator privileges (e.g., provided to the organization's IT team using virtual environment 110). Therefore, the virtual environment platform disclosed herein enables virtual office management and provides multiple options that expand the possibilities of typical virtual meetings and conferencing applications, increase the realism of collaboration and interaction, and simplify the way collaboration occurs.
[0286] Figures 10A-10B A schematic diagram depicts a virtual environment viewed through a corresponding user graphical representation according to an embodiment, wherein the virtual environment is a virtual classroom 1000. Figures 10A-10B The user graphical representation of students and teachers can be any of the following: a user 3D virtual cutout constructed from user-uploaded or third-party source photos; a user real-time 3D virtual cutout with background removed, generated based on a real-time 2D or 3D live video stream feed obtained from a camera; a video with background removed; or a video without background removed.
[0287] exist Figure 10AIn this virtual classroom, multiple user graphical representations of student 1002 are remotely participating in a classroom lecture provided by a user graphical representation of teacher 1004. Teacher 1004 can project classroom content 1006 onto one or more projection surfaces 1008, such as a virtual classroom whiteboard. The virtual classroom 1000 may also include multiple virtual classroom desks 1010 that can support user learning. See reference... Figure 4 The disclosed options allow for multiple interactive choices for student 1002, such as raising hands, screen sharing (e.g., on projection surface 1008), and pointing a laser at specific content, depending on the situation. Figure 10A In the image, the user graphic representation of teacher 1004 is projected onto the projection surface.
[0288] Figure 10B Depicting and Figure 10A A similar embodiment, the difference being that the user graphical representation of teacher 1004 is seated behind virtual desk 1012, while only content 1006 is shared or projected onto the virtual classroom whiteboard projection surface 1008. Because teacher 1004 shares the same virtual space with student 1002 and can move back and forth within classroom 1000, a more realistic and interactive experience is created for both student 1002 and teacher 1004.
[0289] Figure 11 A schematic diagram depicts multiple virtual camera positions 1100 according to an embodiment.
[0290] exist Figure 11 In the two user graphical representations 1102, user 3D virtual clipping 1104, and user real-time 3D virtual clipping 1106, there are one or more virtual camera positions 1100 for one or more virtual cameras, each position containing a viewing direction, angle, and field of view that generate a viewpoint for the user graphical representation.
[0291] In one embodiment, one or more virtual cameras may be located at eye level 1108, pointing outwards from the eye level of the user graphic representation 1102. In another embodiment, two virtual cameras, one for each eye, may point outwards from the two eye levels 1110 of the user graphic representation 1102. In another embodiment, one or more virtual cameras may point outwards from the center of the head position 1112 of the user graphic representation. In another embodiment, one or more virtual cameras may point outwards from the center 1114 of the user graphic representation 1102. In another embodiment, one or more virtual cameras may be positioned in front of the user graphic representation 1102, for example at head level, pointing towards the user graphic representation 1102 when in self-view 1116. (See reference...) Figures 6A-7C As explained, one or more virtual cameras can be created during the process of inserting and combining user graphical representations into a virtual environment.
[0292] In one embodiment, the user's viewpoint captured by the camera is associated with the user's graphical representation of the viewpoint and a virtual camera linked using computer vision, thus manipulating the virtual camera. Furthermore, the virtual camera can be automatically updated, for example, by tracking and analyzing user eye and head tilt data, or head rotation data, or a combination thereof.
[0293] Figure 12 A schematic diagram of a system 1200 for virtual broadcasting from within a virtual environment is depicted.
[0294] System 1200 may include one or more server computers. Figure 12 The illustrated illustrative system 1200 includes at least one media server computer 1202, which includes at least one processor 1204 and memory 1206. Memory 1206 includes data and instructions for a data exchange management module 1208 that implements data exchange between client devices 1210. System 1200 also includes at least one virtual environment 1212 connected to the at least one media server computer 1202. The at least one virtual environment 1212 includes a virtual broadcast camera 1214 located within the at least one virtual environment 1212 and configured to capture multimedia streams from within the at least one virtual environment 1212. The at least one virtual environment 1212 may be hosted by at least one dedicated server computer connected to the at least one media server computer 1202 via a network, or it may be hosted in a peer-to-peer infrastructure and relayed through the at least one media server computer 1202. Multimedia streams are sent to the at least one media server computer 1202 for broadcast to at least one client device 1210. System 1200 also includes at least one camera 1216 that receives live feed data from a user 1218 of at least one client device 1210 and transmits the live feed data from the user to at least one media computer 1202 via at least one client device 1210. (See reference...) Figures 6A-7C The disclosed live feed data received by at least one media computer 1202 can be generated by a combination of multiple image processing methods.
[0295] At least one virtual broadcast camera 1214 sends a multimedia stream to at least one media server computer 1202 for broadcasting the corresponding multimedia stream to a receiving client device 1210 based on data exchange management from the at least one media server computer 1202. The multimedia stream is displayed on a corresponding display to the corresponding user graphical representation 1220 of the user 1218 of the at least one client device 1210. The data exchange management between the client devices 1210, performed by the data exchange management module 1208, includes analyzing the incoming multimedia stream and evaluating and forwarding the outgoing multimedia stream.
[0296] In some embodiments, when forwarding outgoing multimedia streams, at least one media server computer 1202 utilizes a routing topology including Selective Forwarding Unit (SFU), NAT traversal using relay (TURN), Spatial Analysis Media Server (SAMS), or other suitable multimedia server routing topology, or media processing and forwarding server topology, or other suitable server topology. In other embodiments, when utilizing a media processing topology, at least one media server computer 1202 is configured to decode, combine, improve, mix, enhance, expand, compute, manipulate, and encode multimedia streams. In other embodiments, when utilizing a forwarding server topology, at least one media server computer 1202 utilizes one or more of a multipoint control unit (MCU), a cloud media mixer, and a cloud 3D renderer.
[0297] In some embodiments, the incoming multimedia stream includes user priority data and distance relationship data, and the user priority data includes a higher priority score for user graphical representations closer to the incoming multimedia stream source and a lower priority score for user graphical representations farther from the incoming multimedia stream source. In embodiments, a multimedia stream sent by at least one client device 1210 and / or broadcast camera 1214 to at least one media server includes data related to user priority and the distance relationship between the corresponding user graphical representation 1202 and the multimedia stream, including metadata, or priority data, or data category, or spatial structure data, or three-dimensional position, or orientation or motion information, or image data, or media data, and video data based on a scalable video codec, or a combination thereof. In other embodiments, the priority data includes a higher priority score for users closer to the virtual multimedia stream source 1224 and a lower priority score for users farther from the virtual multimedia stream source 1224. In other embodiments, the forwarding of outbound multimedia streams is based on user priority data and distance relationship data. In embodiments, the forwarding of outbound multimedia streams by the media server based on user priority and distance relationship data includes optimizing the bandwidth and computing resource utilization of one or more receiving client devices.
[0298] In some embodiments, at least one virtual broadcast camera 1214 is virtually regarded as a virtual broadcast camera 1214 configured to broadcast multimedia streams within at least one virtual environment 1212. The virtual broadcast camera 1214 may be located near a virtual multimedia stream source 1224 and may also move back and forth within the virtual environment 1212. In other embodiments, the virtual broadcast camera 1214 may be managed by a client device 1210 accessing the virtual environment and may be configured to manipulate the viewpoint of the camera updated in the virtual environment, broadcasting the updated viewpoint to at least one client device associated with the virtual broadcast camera 1214.
[0299] In some embodiments, the virtual multimedia streaming source 1224 includes live virtual events, including one or more of group discussions, speeches, meetings, presentations, webinars, entertainment programs, sporting events, and performances, wherein multiple user graphical representations of real speakers speaking remotely (e.g., being recorded onto their corresponding cameras 1216 while speaking from their homes) are placed within the virtual environment 1212.
[0300] In some embodiments, the multimedia stream may be viewed as a real-time 3D view in a web browser rendered on a client or cloud computer, or it may be streamed for live viewing on a suitable video platform (e.g., YouTube). TM Live streaming, Twitter TM Facebook TM Live streaming, Zoom TM wait).
[0301] exist Figure 12 In the example shown, user ACs access a virtual environment 1212 through their corresponding client devices. Each user AC has a camera 1216 that transmits a multimedia stream corresponding to each user AC. This camera can be used to generate a user graphical representation AC and insert and combine it into the virtual environment 1212, as described in embodiments of this disclosure. Therefore, in the virtual environment 1212, each user AC has a corresponding user graphical representation AC. The multimedia streams transmitted by at least one camera 1216 through at least one client device 1210, and the multimedia streams transmitted by at least one broadcast camera 1214 to at least one media server 1202, contain data related to user priority and the distance relationship between the corresponding user graphical representation and the multimedia stream. This data includes, for example, metadata, priority data, data categories, spatial structure data, three-dimensional position, orientation or motion information, image data, media data, video data based on a scalable video codec, etc. This data can be used by a data exchange management module 1208 to manage data exchange between client devices 1210, including analyzing and optimizing incoming multimedia streams and evaluating and optimizing the forwarding of outgoing multimedia streams.
[0302] Therefore, for example, when user graphical representation A is closer to virtual multimedia streaming source 1224 in virtual environment 1212, the forwarding of outbound media streams can be optimized to include, for example, images with a higher resolution for user graphical representation A than those provided to user graphical representations B and C. The multimedia stream can be viewed in first person by users via their client device 1210 through their user graphical representation 1222, for example, within virtual environment 1212. In some examples, the multimedia stream is viewed as a real-time 3D view in a web browser rendered on a client or cloud computer. Users can watch live multimedia streams of events (e.g., webinars, conferences, panel discussions, presentations, etc.) as a real-time 3D view in a web browser rendered on a client or cloud computer, or they can be streamed for live viewing on suitable video platforms and / or social media.
[0303] Figure 13 A schematic diagram of a system 1300 for delivering applications within a virtual environment is depicted.
[0304] System 1300 includes at least one cloud server computer 1302, which includes at least one processor 1304 and memory 1306. Memory 1306 contains data and instructions implementing at least one virtual environment 1308 linked to application module 1310. Application module 1310 includes one or more installed applications 1312 and corresponding multi-user interaction application rules 1314. In response to selection by virtual environment host 1316 via client device 1318, one or more installed applications 1312 are displayed and activated during a session of virtual environment 1302, enabling virtual environment host user graphical representation 1320 and any participant user graphical representation 1322 within virtual environment 1308 to interact with one or more installed applications 1312 via their respective client devices 1318. At least one cloud server computer 1302 manages and processes received user interactions with one or more installed applications 1312 according to the multi-user interaction application rules 1314 in application module 1310. At least one cloud server computer 1302 further forwards the processed interaction to each client device 1318 accordingly to establish a multi-user session in the virtual environment 1308, thereby realizing a shared experience according to the multi-user interaction application rules 1314.
[0305] In some embodiments, the multi-user interaction application rules 1314 are stored and managed in one or more separate application servers that can be connected to at least one cloud server computer 1302 via a network.
[0306] In some embodiments, one or more applications are installed from application installation packages available in an application library, providing application services through corresponding application programming interfaces. In other embodiments, the application library is context-filtered. In these embodiments, context filtering is designed to provide only applications relevant to specific contexts. For example, host 1316 can context-filter the application library (e.g., an app store) to find applications related to specific contexts (e.g., learning, entertainment, sports, reading, purchasing, weather, work, etc.), and can select an application of interest to install within application module 1310. In other embodiments, the application library is hosted on one or more third-party server computers, or on at least one cloud server computer 1302.
[0307] In some embodiments, one or more installed applications are shared with and viewed through a virtual display application installed on a corresponding client device. In another embodiment, after installation and activation, one or more installed applications 1312 are shared with and viewed through a virtual display application 1324 installed on a corresponding client device 1318. The virtual display application 1324 can be configured to receive one or more installed applications 1312 from an application library and publish one or more selected installed applications 1312 to be displayed on their respective client devices 1318 to the conference host user graphical representation 1320 and other participant user graphical representations 1322 in the virtual environment 1308. The virtual display application 1324 can be an online or installed file viewer application that can be configured to receive and display the installed applications 1312.
[0308] In some embodiments, application module 1310 is represented as a 2D screen or 3D volume application module graphical representation 1326 within a virtual environment, displaying content from the installed application 1312 to a user graphical representation 1322 within the virtual environment. In other embodiments, virtual display application 1324 is represented as a 2D screen or 3D volume, displaying content from the installed application to a user graphical representation within a virtual environment 1308.
[0309] In some embodiments, one or more applications 1312 are installed directly within the virtual environment 1308 before or simultaneously with the multi-user session. In other embodiments, one or more applications 1312 are installed using a virtual environment setup tool before the multi-user session begins.
[0310] In some embodiments, application rules for multi-user interaction can define synchronous or asynchronous interaction, or a combination thereof, to update user interactions and corresponding updated views of one or more applications. Both synchronous and asynchronous interactions can be configured through multi-user interaction application rule 1314 and can be implemented via parallel processing by at least one server computer 1302, or via a dedicated server computer for processing interactions with a single user of at least one installed application 1312.
[0311] For example, if host 1316 is a teacher, the teacher can choose to display a workbook application containing book content to users. The teacher can edit the workbook, and students can view the same workbook with the teacher's edits via their virtual display application 1324, either with synchronous interaction and a corresponding updated view, or without teacher edits when synchronous interaction is selected. In another example, in a presentation application containing a presentation file with multiple slides, asynchronous interaction allows each user to view individual slides asynchronously. In another example, in the case of an educational application, presenting cardiac anatomy while testing students, where student interactions are synchronous so that other students witness and observe the interactions performed by the students. In another example, a teacher can write on a whiteboard, allowing students to synchronously view the text written on the whiteboard via their virtual display application. In yet another example, a video player application can synchronously display video to all students.
[0312] In some exemplary embodiments, the virtual environment 1308 is a classroom, or office space, or meeting room, or conference room, or auditorium, or theater.
[0313] Figure 14 The description illustrates the use of a method based on embodiments for... Figure 13 A schematic diagram of the virtual environment 1308 of the system 1300 for delivering applications within a virtual environment is depicted.
[0314] Virtual environment 1308 includes an application module graphical representation 1326, which contains at least one installed application 1312 selected by host 1316 of virtual environment 1308, and two users A and B who view and interact with the installed application 1312 through their corresponding virtual display application 1324. It can be understood that user A can view a page (e.g., page 1) of a book application through virtual display application 1324, which may be the same page selected by host 1316 through application module graphical representation 1326, representing synchronous interaction and management of the installed application 1312. On the other hand, through asynchronous interaction and management of the installed application 1312 by virtual display application 1324, user B can view a different page than host 1316 and user A.
[0315] Figure 15 A schematic diagram of a system 1500 that provides virtual computing resources within a virtual environment according to an embodiment is depicted.
[0316] System 1500 includes a server computer system comprising one or more server computers, including at least one cloud server computer 1502. The at least one cloud server computer 1502 includes at least one processor 1504 and memory 1506, the memory 1506 including data and instructions for implementing at least one virtual environment 1508, and at least one virtual computer 1510 associated with the at least one virtual environment 1508. The at least one virtual computer 1510 receives virtual computing resources from the server computer system. In an embodiment, the at least one virtual computer has a corresponding graphical representation 1512 within the virtual environment 1508. The graphical representation 1512 can provide other benefits, such as facilitating user interaction with the virtual computer and increasing the realism of the user experience (e.g., for a home office experience). Therefore, in an embodiment, the at least one virtual computer includes at least one corresponding associated graphical representation 1512 located within the virtual environment 1508, wherein the at least one virtual computer 1510 receives virtual computing resources from the at least one cloud server computer 1502. System 1500 also includes at least one client device 1514 connected via a network to the at least one server computer 1510. In response to at least one client device 1514 accessing one or more virtual computers 1510 (e.g., by interacting with a corresponding graphical representation), at least one cloud server computer 1502 provides at least a portion of available virtual computing resources to at least one client device 1514.
[0317] In some embodiments, virtual computing resources are accessed by a user graphical representation 1516 of user 1518, who accesses (e.g., interacts with) one or more graphical representations of a virtual computer 1512 within a virtual environment 1508 via a corresponding client device 1514, and thereby provides them to the corresponding client device 1514.
[0318] In some embodiments, a virtual computer graphical representation 1512 is spatially positioned within a virtual environment for access by a user graphical representation. In one embodiment, the arrangement of the virtual environment 1508 is associated with a contextual theme of the virtual environment 1508 and may include arrangements of virtual objects, furniture, floor plans, etc., for educational, meeting, work, shopping, service, social, and entertainment purposes, respectively. In other embodiments, one or more virtual computer graphical representations are positioned within the arrangement of the virtual environment 1508 for access by one or more user graphical representations 1516. For example, the virtual computer may be positioned in a virtual room that a user graphical representation 1516 would access when participating in activities that may require or benefit from the ability to use resources associated with the virtual computer, such as conducting projects in a virtual classroom, laboratory, or office.
[0319] In some embodiments, the server computer system is configured to provide at least a portion of virtual computing resources to at least one client device in response to a user accessing at least one cloud server computer by logging in to at least one client device without accessing a virtual environment. In an illustrative scenario, a user 1518 accessing at least one cloud server computer 1502 accesses virtual computing resources by physically logging in to a client device 1514 connected to at least one cloud server computer 1502 via a network, triggering the provision of virtual computing resources to client device 1514 without accessing a virtual environment. For example, user 1518 can log in to cloud server computer 1502 from his or her home computer and access virtual computer 1510 to receive virtual computing resources accordingly. In another example, user 1518 can log in to cloud server computer 1502 from his or her work computer to access virtual computer 1510 and receive virtual computing resources accordingly.
[0320] In some embodiments, at least a portion of the virtual computing resources are assigned to client devices via management tools. Therefore, the virtual infrastructure, including all associated virtual computers, can be managed using administrator options based on exclusive administrator privileges (e.g., providing a virtual environment to an organization's IT team).
[0321] In some embodiments, the provision of virtual computing resources is performed based on a stored user profile. In this embodiment, the assignment of virtual computing resources is performed based on a stored user profile containing one or more parameters associated with and assigned to the user profile, including priority data, security data, QoS, bandwidth, storage space, or computing power, or a combination thereof. For example, a user accessing a work virtual computer from home can have a personal profile configured to provide the user with specific virtual computing resources associated with that profile.
[0322] In some embodiments, each virtual computer is a downloadable application available from an application library.
[0323] Figure 16 A schematic diagram of a system 1600 for implementing self-organizing virtual communication between user graphical representations according to an embodiment is depicted.
[0324] System 1600 includes one or more cloud server computers 1602, each including at least one processor 1604 and a memory 1606. The memory 1606 stores data and instructions for implementing a virtual environment 1608. The virtual environment 1608 is configured such that at least one proximate user graphical representation and at least one target user graphical representation within the virtual environment 1608 can open an ad hoc communication channel, enabling ad hoc dialogue via the ad hoc communication channel between user graphical representations within the virtual environment 1608. Figure 16 In the example shown, the system also includes two or more client devices 1610 that access at least one virtual environment through corresponding user graphical representations and are connected to one or more cloud server computers 1602 via a network 1612. The virtual environment 1608 enables at least one proximate user graphical representation 1614 and at least one target user graphical representation 1616 to open a self-organizing communication channel 1618 from their respective users 1620, allowing for self-organizing dialogue between user graphical representations within the virtual environment 1608.
[0325] In some embodiments, the opening of the self-organizing communication channel 1618 is performed based on the distance, position and orientation between user graphical representations, or the current availability status, privacy settings, or the state configuration of the self-organizing communication, or a combination thereof.
[0326] In some embodiments, the self-organizing dialogue is performed within the virtual environment 1608 where the two user graphical representations are located. For example, if a nearby user graphical representation 1614 encounters a target user graphical representation 1614 in a specific area of a lounge or office space, self-organizing communication can be enabled to allow the two users to maintain a dialogue within that specific area of the lounge or office space without changing their locations. In other embodiments, the self-organizing dialogue is performed using the current viewpoint within the virtual environment. In the example above, self-organizing communication can be enabled to allow the two users to maintain a dialogue without changing their viewpoint. In other embodiments, the self-organizing dialogue enables optional changes in viewpoint, location, or a combination thereof within the same or another connected virtual environment where the self-organizing dialogue occurs.
[0327] In some embodiments, one or more cloud server computers are also configured to generate visual feedback in the virtual environment, signaling the possibility of self-organizing communication. In an embodiment, a user graphical representation receives visual feedback that signals the possibility of self-organizing communication, thereby triggering the opening of a self-organizing communication channel and signaling the start of a self-organizing dialogue between user graphical representations.
[0328] In some embodiments, the self-organizing dialogue includes sending and receiving real-time audio and video from a user graphical representation display.
[0329] In some embodiments, a user corresponding to a nearby user graphical representation 1614 selects and clicks on a target user graphical representation 1616 before opening the self-organizing communication channel 1618.
[0330] In some embodiments, one or more cloud server computers are also configured to open a self-organizing communication channel in response to an accepted invitation. For example, a user corresponding to a nearby user graphical representation 1614 further sends a self-organizing communication participation invitation to a target user graphical representation 1616, and opens a self-organizing communication channel 1618 after receiving invitation approval from the target user graphical representation 1614.
[0331] In some embodiments, the self-organizing communication channel 1618 is implemented via at least one cloud server computer or as a P2P communication channel.
[0332] Figure 17 An embodiment of a method 1700 for implementing interaction in a virtual environment according to an embodiment is described.
[0333] The method 1700 for implementing interaction in a virtual environment according to this disclosure begins in steps 1702 and 1704 by providing a virtual environment platform containing at least one virtual environment in the memory of one or more cloud server computers containing at least one processor.
[0334] As shown in steps 1706 and 1708, the method receives a live data feed from a user on a client device from at least one camera, and then generates a user graphical representation based on the live data feed. Method 1700 then inserts the user graphical representation into the three-dimensional coordinates of the virtual environment, as shown in step 1710.
[0335] Subsequently, in step 1712, the method updates the user's graphical representation within the virtual environment based on the live data feed. Finally, in step 1714, the method processes data generated from interactions within at least one virtual environment using the corresponding graphical representation located within the virtual environment, and concludes in step 1716.
[0336] Figure 18An embodiment of the image processing method 1800 according to an embodiment is described.
[0337] Method 1800 begins with steps 1802 and 1804, providing data and instructions for implementing image processing functions in the memory of at least one cloud server computer. In step 1806, method 1800 continues by acquiring live data feeds from at least one user of at least one corresponding client device, obtained from at least one camera. Then, in step 1808, method 1800 continues with one or more image processing combinations from one or more cloud server computers and at least one client device (e.g., ...). Figures 6A-7C The image processing combination generates a user graphical representation, and the process can then end in step 1810. One or more cloud server computers and at least one client device can utilize a hybrid system architecture from this disclosure that includes P2P and client-server components (e.g., ...). Figure 3 Interact with the hybrid system architecture 300.
[0338] Figure 19 A user authentication method 1900 based on user graphical representation according to an embodiment is described.
[0339] Method 1900 begins with steps 1902 and 1904, providing a user database in the memory of one or more cloud server computers that stores user data associated with user accounts and corresponding user graphical representations, and a face scanning and authentication module connected to the user database. Method 1900 continues in step 1906, receiving a request to access a user account from a client device, and then in step 1908, performing a face scan on a user of at least one client device via the face scanning and authentication module using images received from at least one camera, which may be connected to at least one client device and / or one or more cloud server computers. In check 1910, method 1900 continues by checking for a match between the user data associated with the user account in the user database. If the user account is available, method 1900 continues in step 1912, providing the user with the corresponding user graphical representation and access to the user account. In the negative case, if the user account is unavailable, method 1900 may continue in step 1914, generating a new user graphical representation and storing a new user account in the user database, and access to the user account. This process may end in step 1916.
[0340] Figure 20 A block diagram of a method 2000 for virtual broadcasting from within a virtual environment, according to an embodiment, is shown.
[0341] Method 2000 begins in step 2002 by providing data and instructions in the memory of at least one media server for a client device data exchange management module that implements the management of data exchange between client devices. Method 2000 continues in step 2004 by capturing a multimedia stream using a virtual broadcast camera located within at least one virtual environment connected to at least one media server.
[0342] In step 2006, method 2000 continues by sending a multimedia stream to at least one media server for broadcasting to at least one client device. In step 2008, method 2000 continues by acquiring live feed data from at least one camera via at least one client device.
[0343] In step 2010, the method continues to perform data exchange management, including analyzing and optimizing incoming multimedia streams from at least one virtual environment and live feed data from users, as well as evaluating and optimizing the forwarding of outgoing multimedia streams. Finally, in step 2012, method 2000 ends, and the corresponding multimedia stream is broadcast to the client device based on the data exchange management, wherein the multimedia stream is displayed to the user's graphical representation on at least one client device.
[0344] Figure 21 A block diagram of a method 2100 for delivering an application within a virtual environment, according to an embodiment, is shown.
[0345] Method 2100 begins at step 2102, providing at least one virtual environment in the memory of at least one cloud server computer, and an application module including one or more installed applications and corresponding application rules for multi-user interaction, wherein the application module is linked to and visible within the virtual environment. In step 2104, method 2100 continues by receiving a selection instruction from the virtual environment host. Then, in step 2106, method 2100 continues by displaying and activating one or more installed applications during a session in the virtual environment, enabling the user graphical representation of the virtual environment host and the user graphical representation of any participant within the virtual environment to interact via corresponding client devices.
[0346] In step 2108, method 2100 continues by receiving user interactions with one or more installed applications. Subsequently, method 2100 continues, as shown in step 2110, to manage and process user interactions with one or more installed applications according to application rules for multi-user interactions in the application module. Finally, method 2100 terminates in step 2112, forwarding the processed interactions to each client device accordingly to establish a multi-user session for a shared experience based on application rules.
[0347] Figure 22 A block diagram illustrating a method 2200 for providing virtual computing resources within a virtual environment according to an embodiment is shown.
[0348] Method 2200 begins at step 2202, providing at least one virtual computer in the memory of at least one cloud server computer, and a virtual environment containing one or more graphical representations of the virtual computer. The method continues in step 2204, where the virtual computer receives virtual computing resources from the at least one cloud server computer. Then, in step 2206, the method continues by receiving an access request for the one or more virtual computers from at least one client device. Finally, in step 2208, the method ends by providing a portion of the available virtual computing resources to the at least one client device based on the client device's request.
[0349] Figure 23 A block diagram of a method 2300 for implementing self-organizing virtual communication between user graphical representations is shown.
[0350] Method 2300 begins in step 2302 by providing a virtual environment in the memory of one or more cloud server computers containing at least one processor. Then, in step 2304, the method continues by detecting two or more client devices accessing at least one virtual environment via corresponding graphical representations, wherein the client devices are connected to the one or more cloud server computers via a network. Finally, in step 2306, method 2300 ends by opening a self-organizing communication channel in response to at least one user graphical representation approaching another user graphical representation, enabling self-organizing dialogue between user graphical representations within the virtual environment.
[0351] The document also describes computer-readable media on which instructions are stored, configured to cause one or more computers to perform any of the methods described herein. As used herein, the term "computer-readable medium" includes volatile and non-volatile, removable and non-removable media implemented in any method or technology capable of storing information such as computer-readable instructions, data structures, program modules, or other data. In general, the functionality of the computing devices described herein can be implemented in computing logic embodied in hardware or software instructions that can be written in programming languages such as C, C++, COBOL, and JAVA. TM , PHP, Perl, Python, Ruby, HTML, CSS, JavaScript, VBScript, ASPX, Microsoft.NET TMLanguages such as C# are used. The computational logic can be compiled into an executable program or written in an interpreted programming language. Typically, the functionality described herein can be implemented as a logic module, which can be replicated to provide greater processing power, merged with other modules, or divided into submodules. The computational logic can be stored in any type of computer-readable medium (e.g., non-transitory media such as memory or storage media) or computer storage device, and can be stored on and executed by one or more general-purpose or special-purpose processors, thereby creating a special-purpose computing device configured to provide the functionality described herein.
[0352] While specific embodiments have been described and illustrated in the accompanying drawings, it should be understood that such embodiments are merely illustrative and not intended to limit the scope of the invention, and that the invention is not limited to the specific constructions and arrangements shown and described, as various other modifications will be apparent to those skilled in the art. Therefore, this description is to be regarded as illustrative rather than restrictive.
Claims
1. A system for virtual broadcasting from within a virtual environment, characterized in that, Include: A server computer system comprising one or more server computers, each server computer including at least one processor and memory, the server computer system comprising: Data and instructions, the implementation of which is a data exchange management module configured to manage data exchange between client devices; and At least one virtual environment comprising a virtual broadcast camera located within the at least one virtual environment and configured to capture a multimedia stream from a virtual multimedia stream source within the at least one virtual environment, wherein a server computer system is configured to receive live feed data captured by at least one camera from at least one client device and broadcast the multimedia stream to the at least one client device based on data exchange management, wherein the broadcast multimedia stream is configured to be displayed to a corresponding user graphical representation generated from the user live data feed from the at least one client device, and wherein the data exchange management between client devices performed by the data exchange management module includes analyzing the incoming multimedia stream and evaluating the forwarded outgoing multimedia stream based on the analysis of the incoming multimedia stream; in, The incoming multimedia stream includes user priority data and distance relationship data, wherein the user priority data includes a higher priority score for user graphic representations that are closer to the incoming multimedia stream source and a lower priority score for user graphic representations that are farther away from the incoming multimedia stream source. The forwarding of the outbound multimedia stream is based on the user priority data and the distance relationship data.
2. The system according to claim 1, characterized in that, The server computer system utilizes a routing topology that includes Selective Forwarding Units (SFUs), Relay NAT Traversal (TURN), or Spatial Analysis Media Servers (SAMS).
3. The system according to claim 1, characterized in that, The server computer system uses a media processing topology to process the outbound multimedia stream for viewing by the user's graphical representation within the at least one virtual environment via the client device.
4. The system according to claim 1, characterized in that, The server computer system described therein uses a forwarding server topology that includes one or more multipoint control units (MCUs), a cloud media mixer, and a cloud 3D renderer.
5. The system according to claim 1, characterized in that, The forwarding of the outbound multimedia stream includes modifying, scaling up, or scaling down the multimedia stream for time, space, quality, or color characteristics.
6. The system according to claim 1, characterized in that, The virtual broadcast camera is managed by a client device that accesses the virtual environment.
7. The system according to claim 1, characterized in that, The at least one virtual environment includes a plurality of virtual broadcast cameras, each virtual broadcast camera providing a multimedia stream from a corresponding viewpoint within the at least one virtual environment.
8. The system according to claim 1, characterized in that, The at least one virtual environment is hosted by at least one dedicated server computer connected to at least one media server computer via a network, or hosted in a peer-to-peer infrastructure and relayed through at least one media server computer.
9. A method for virtual broadcasting from within a virtual environment, characterized in that, Include: Multimedia streams are captured by a virtual broadcast camera positioned within at least one virtual environment connected to at least one media server; The multimedia stream is sent to the at least one media server for broadcast to at least one client device; Obtain real-time feed data from the at least one client device; Performing data exchange management includes analyzing incoming multimedia streams and live feed data from the at least one virtual environment, and evaluating the forwarding of outgoing multimedia streams; as well as Based on the data exchange management, the outbound multimedia stream is broadcast to the client device, wherein the outbound multimedia stream is displayed to a user graphical representation in the at least one virtual environment; in, The incoming multimedia stream includes user priority data and distance relationship data, wherein the user priority data includes a higher priority score for user graphic representations that are closer to the incoming multimedia stream source and a lower priority score for user graphic representations that are farther away from the incoming multimedia stream source. The forwarding of the outbound multimedia stream is based on the user priority data and the distance relationship data.
10. The method according to claim 9, characterized in that, The forwarding of the outbound multimedia stream includes modifying, scaling up, or scaling down the multimedia stream for time, space, quality, or color characteristics.
11. The method according to claim 9, characterized in that, The at least one virtual environment includes a plurality of virtual broadcast cameras, each virtual broadcast camera providing a multimedia stream from a corresponding viewpoint within the at least one virtual environment.
12. A computer-readable medium having instructions stored thereon, characterized in that, The instructions are configured to cause at least one server computer, including a processor and memory, to perform the following steps: Provides data exchange, management data and instructions, and at least one virtual environment; Obtain a live data feed containing the incoming multimedia stream from a virtual multimedia stream source; Perform data exchange management, including analyzing the incoming multimedia streams and evaluating the forwarding of outgoing multimedia streams; as well as Based on the data exchange management, the outbound multimedia stream is broadcast to a receiving client device within the at least one virtual environment, wherein the outbound multimedia stream is displayed to a user graphical representation within the at least one virtual environment. in, The incoming multimedia stream includes user priority data and distance relationship data, wherein the user priority data includes a higher priority score for user graphic representations that are closer to the incoming multimedia stream source and a lower priority score for user graphic representations that are farther away from the incoming multimedia stream source. The forwarding of the outbound multimedia stream is based on the user priority data and the distance relationship data.
13. The computer-readable medium according to claim 12, characterized in that, The forwarding of the outbound multimedia stream includes modifying, scaling up, or scaling down the multimedia stream for time, space, quality, or color characteristics.
14. The computer-readable medium according to claim 12, characterized in that, The at least one virtual environment includes a plurality of virtual broadcast cameras, each virtual broadcast camera providing a multimedia stream from a corresponding viewpoint within the at least one virtual environment.
Citation Information
Patent Citations
Spatially faithful telepresence supporting varying geometries and moving users
US20200099891A1