User authentication system and method using a graphic representation base

The system enhances virtual reality interactions by inserting user graphic representations into 3D virtual environments, improving presence and collaboration, addressing the limitations of existing technologies in providing realistic and cost-effective interactions.

JP7698300B2Active Publication Date: 2025-06-25TMRW FOUNDATION IP SARL
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021138957
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-01
Filing Date
2021-08-27
Publication Date
2025-06-25
Estimated Expiration
2041-08-27

AI Technical Summary

Technical Problem

Existing virtual reality and augmented reality technologies provide low-quality interactions with a lack of user presence and shared space, leading to feelings of loneliness and reduced productivity, often requiring expensive equipment and infrastructure.

Method used

A system that enables real-time multi-user collaboration and interaction in virtual environments using cloud server computers to insert user graphic representations generated from live data feeds into three-dimensional coordinates, allowing interactions without the need for expensive equipment, and supports ad-hoc communication and application delivery within virtual environments.

Benefits of technology

Enhances user presence and interaction quality in virtual environments, improving productivity by providing realistic interactions and collaboration, accessible on existing devices without the need for costly hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698300000001
    Figure 0007698300000001
  • Figure 0007698300000002
    Figure 0007698300000002
  • Figure 0007698300000003
    Figure 0007698300000003
Patent Text Reader

Abstract

To provide a system, method, and computer-readable medium for providing a graphical representation-based user authentication system.SOLUTION: A system 100 comprises: at least one processor; one or more cloud server computers 102 including memory for storing data and instructions for implementing a virtual environment platform including at least one virtual environment; at least one camera 112 for acquiring a live data feed from a user of a client device; and the client device 118 communicatively connected to the one or more cloud server computers and at least one camera. The system generates user graphical representation from the live data feed, which is inserted into a selected virtual environment, which is updated in the virtual environment, and enables real-time multi-user collaboration and interactions in the virtual environment.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application is a part of U.S. Patent Application No. 17 / 006,327, filed on August 28, 2020, which is related to U.S. Patent Application No. 17 / 005,767, filed on August 28, 2020. Each of the aforementioned patent applications is hereby incorporated by reference in its entirety.

Background Art

[0002] Due to situations such as the pandemic of the novel coronavirus in 2020, movement has been restricted worldwide, and changes have occurred in the ways of meetings, learning, shopping, and work, making real - time communication and collaboration, especially interactions including social interactions, increasingly important. Various solutions that enable real - time communication and collaboration are already on the market, ranging from chat applications to videophones such as Skype (trademark) and Zoom (trademark), or virtual offices for remote teams represented by 2D avatars provided by Pragli (trademark).

Summary of the Invention

Problems to be Solved by the Invention

[0003] Considering the current state of development of wearable immersive technologies such as extended reality (e.g., augmented reality and / or virtual reality) and the relatively low rate of technological diversion, it is understandable that most solutions provide a flat 2 - dimensional user interface where most interactions take place. However, when these solutions are compared with real - world experiences, the low level of presence, lack of user presence, lack of shared space, and quality of interactions that can be performed can bring a sense of loneliness or boredom to many users, and as a result, productivity may be lower than directly performing the same actions.

[0004] What is needed is a technical solution that provides the user with a sense of presence, a sense of the physical presence of themselves and the participants, and a feeling of interacting as if in reality when interacting remotely, without the need to purchase expensive equipment (e.g., head-mounted displays, etc.) and implement new or costly infrastructure, while using all existing computing devices and cameras.

Means for Solving the Problem

[0005] This summary is provided to introduce a selection of concepts in a simplified form that will be further described in the following detailed description. This summary is not intended to identify key features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0006] The present disclosure generally refers to computer systems, and more particularly to systems and methods that enable interaction in virtual environments, particularly social interactions, image processing-based virtual presence systems and methods, user authentication systems and methods based on user graphic representations, systems and methods for virtual broadcasting from within a virtual environment, systems and methods for delivering applications within a virtual environment, systems and methods for provisioning cloud computing-based virtual computing resources within a virtual environment cloud server computer, and systems and methods that enable ad-hoc virtual communication between approaching user graphic representations.

[0007] The system of the present disclosure that enables interactions in a virtual environment, particularly including social interactions, comprises one or more cloud server computers comprising at least one processor and a memory storing data and instructions for implementing a virtual environment platform including at least one virtual environment. The one or more cloud server computers are configured to insert a user graphic representation generated from a live data feed obtained by a camera into a three-dimensional coordinate position of at least one virtual environment, update the user graphic representation in at least one virtual environment, and enable real-time multi-user collaboration and interaction in the virtual environment.

[0008] In one embodiment, the system further comprises at least one camera for obtaining a live data feed from one or more users of a client device. Additionally, the system comprises a client device communicatively connected to the one or more cloud server computers and the at least one camera. The system generates a user graphic representation from the live data feed, which is inserted into the three-dimensional coordinates of the virtual environment and updated there using the live data feed. In the described embodiment, inserting the user graphic representation into the virtual environment involves graphically combining the user graphic representation with the virtual environment such that the user graphic representation appears in the virtual environment (e.g., at a specified 3D coordinate position). The virtual environment platform provides the virtual environment to one or more client devices. The system enables real-time multi-user collaboration and (social) interaction in the virtual environment by accessing a graphical user interface through the client device. The client device or peer device of the present disclosure may include, for example, a computer, a headset, a mobile phone, glasses, a transparent screen, a tablet, and a general input device with a built-in camera or connected to a camera to receive a data feed from the camera.

[0009] In some embodiments, the virtual environment is accessible by a client device via a downloadable client application or a web browser application.

[0010] In some embodiments, the user graphic representation includes a user 3D virtual cutout with the background removed, or a user real-time 3D virtual cutout with the background removed, or a video with the background removed, or a video with the background not removed. In some embodiments, the user graphic representation is a user 3D virtual cutout with the background removed constructed from a photo uploaded by the user or provided by a third party, or a user real-time 3D virtual cutout with the background removed generated based on real-time 2D stereo depth data or 3D live video stream data feed obtained from a camera, and thus including a user real-time video stream, or a video with the background not removed, or a video with the background removed, and is displayed using a polygon structure. Such a polygon structure can be a cut structure or a more complex 3D structure used as a virtual frame for corresponding to the video. In still other embodiments, one or more of such user graphic representations are inserted into three-dimensional coordinates within the virtual environment and graphically combined there.

[0011] The user 3D virtual cutout may include a virtual replica of the user constructed from 2D photos uploaded by the user or provided by a third party. In one embodiment, the user 3D virtual cutout is created by a virtual reconstruction process through machine vision technology that uses 2D photos uploaded by the user or provided by a third party as input data to generate a 3D mesh or 3D point cloud of the user with the background removed. The user real-time 3D virtual cutout may include a virtual replica of the user after the background has been removed, based on real-time 2D or 3D live video stream data feeds obtained from a camera. In one embodiment, the user real-time 3D virtual cutout is created by a virtual reconstruction process through machine vision technology by using the user's live data feed as input data to generate a 3D mesh or 3D point cloud of the user with the background removed. The video with the background removed includes the video streamed to the client device, and the background removal process is performed on the video so that only the user can be seen and is displayed using a polygon structure on the receiving client device. The video without the background removed includes the video streamed to the client device, and the video faithfully represents the camera capture, and thus the user and the user's background can be seen and is displayed using a polygon structure on the receiving client device.

[0012] In some embodiments, the data used as input data included in the live data feed and / or 2D photos uploaded by the user or provided by a third party may include 2D or 3D image data, 3D geometry, video data, media data, audio data, text data, tactile data, time data, 3D entities, 3D dynamic objects, text data, time data, metadata, priority data, security data, location data, lighting data, depth data, and infrared data, etc.

[0013] In some embodiments, the user graphic representation is associated with a top - view viewing perspective, or a third - person viewing perspective, or a first - person viewing perspective, or a self - view viewing perspective. In one embodiment, the user's viewing perspective when accessing the virtual environment through the user graphic representation is a top - view viewing perspective, or a third - person viewing perspective, or a first - person viewing perspective, or a self - view viewing perspective, or a broadcast camera perspective. The self - view viewing perspective may include the user graphic representation as seen by another user graphic representation, and optionally the virtual background of the user graphic representation.

[0014] In yet another embodiment, the viewing perspective is updated when the user manually navigates the virtual environment via the graphical user interface.

[0015] In yet another embodiment, the viewing perspective is automatically established and updated using a virtual camera, and the viewing perspective of the live data feed is associated with the viewing perspective of the user graphic representation and the virtual camera, where the virtual camera is automatically updated by tracking and analyzing the tilt data of the user's eyes and head, or the rotation data of the head, or a combination thereof. In one embodiment, the viewing perspective is automatically established and updated using one or more virtual cameras that are virtually positioned and aligned in front of the user graphic representation, such as in front of a video with an underexposed background, or a video with a removed background, or a user 3D virtual cutout, or a user real-time 3D virtual cutout. In one embodiment, one or more virtual cameras can be directed outward from eye level. In another embodiment, two (one for each eye) virtual cameras are directed outward from the height of both eyes. In yet another embodiment, one or more virtual cameras can be directed outward from the center of the head position of the user graphic representation. In yet another embodiment, one or more virtual cameras can be directed outward from the center of the user graphic representation. In yet another embodiment, when in the self-viewing perspective, one or more virtual cameras are positioned in front of the user graphic representation, for example, at the height of the head of the user graphic representation and directed towards the user graphic representation. The viewing perspective of the user captured by the camera is associated with the viewing perspective of the user graphic representation and the associated virtual camera that uses computer vision to operate the virtual camera. Further, the virtual camera is automatically established and updated by tracking and analyzing the tilt data of the user's eyes and head, or the rotation data of the head, or a combination thereof.

[0016] In yet another embodiment, the self-viewing perspective includes a cutout of a graphical representation (such as in the "selfie mode" of a phone camera) as viewed by another user graphical representation with the background removed. The self-viewing perspective alternatively includes a virtual background of a virtual environment behind the user graphical representation to understand how the user is perceived by other participants. When the self-viewing perspective includes a virtual background of the user graphical representation, it can be set as the area around the user graphical representation that can be captured by a virtual camera, and as a result, it can be circular, square, rectangular, or any other shape suitable for framing the self-viewing perspective.

[0017] In some embodiments, updating the user graphical representation within the virtual environment includes updating the user status. In one embodiment, available user statuses include away, busy, available, offline, in a phone call, or in a meeting. The user status can be updated manually through a graphical user interface. In other embodiments, the user status is automatically updated by connecting to and synchronizing with user calendar information that includes user status data. In yet other embodiments, the user status is automatically updated through detection of the usage of a particular program, such as a programming development environment, 3D editor, or other productivity software, that specifies a busy status that can be synchronized with the user status. In yet another embodiment, the user status can be automatically updated through a machine vision algorithm based on data feeds obtained from a camera.

[0018] In some embodiments, interactions between users through corresponding user graphic representations, especially those involving social interactions, include chat, screen sharing, host options, remote sensing, recording, voting, document sharing, emoji sending, sharing and editing of topics, virtual hugs, raising hands, waving, walking, interactive applications or static or interactive 3D assets, animations, or addition of content including 2D textures, preparation of meeting summaries, movement of objects, projection of content, laser pointing, gameplay, purchasing, participating in ad-hoc virtual communication, and participating in private or group conversations.

[0019] In some embodiments, the virtual environment is a persistent virtual environment stored in the persistent memory storage of one or more cloud server computers, or a temporary virtual environment stored in the temporary memory storage of one or more cloud server computers. In one embodiment, the virtual environment is a persistent virtual environment that records changes made therein, and the changes include customizations stored in the persistent memory storage of at least one cloud server computer designated for the persistent virtual environment. In other embodiments, the virtual environment is a temporary virtual environment stored in the temporary memory storage of a cloud server.

[0020] In some embodiments, the configuration of the virtual environment is associated with a context theme of the virtual environment related to one or more virtual environment verticals selected from the virtual environment platform. In one embodiment, possible configurations include those for use in education, meetings, work, shopping, services, socializing, or entertainment, or combinations thereof. A composite of virtual environments within one or more verticals can represent a virtual environment cluster.

[0021] In further embodiments, the virtual environment cluster is a virtual school with at least a plurality of classrooms, or a virtual company with at least a plurality of work areas and reception rooms where a part of it is shared as a co-working or networking space for members of different organizations, or an event facility with at least one indoor or outdoor event area that hosts live events including live entertainment performances, or a virtual shopping mall with at least a plurality of stores, or a virtual casino with at least a plurality of gaming areas, or a virtual bank with at least a plurality of service areas, or a virtual nightclub with at least a plurality of VIP areas and / or party areas including live disc jockey (DJ) performances, or a virtual karaoke entertainment facility with a plurality of private or public karaoke rooms, or a virtual cruise ship with a plurality of virtual areas inside the cruise ship and areas outside the cruise ship including landscapes, islands, towns, and cities that users can visit after getting off the virtual cruise ship, or one or more of an e-sports stadium or a gymnasium.

[0022] In further embodiments, the virtual environment further comprises a virtual computer including virtual resources. In one embodiment, the virtual resources are from one or more cloud computer resources accessed through a client device and are allocated to the virtual computer resources by a management tool.

[0023] In some embodiments, the virtual environment platform is configured to enable multicasting or broadcasting of remote events to multiple instances of the virtual environment. This can be done to allow a large number of users from various regions of the world to experience the same live event being multicast.

[0024] In some embodiments, clickable links that redirect to the virtual environment are embedded in one or more third-party sources including third-party websites, applications, or video games.

[0025] In another aspect of the present disclosure, a method for enabling interactions including social interactions in a virtual environment includes providing a virtual environment platform including at least one virtual environment in the memory of one or more cloud server computers comprising at least one processor; receiving a live data feed from at least one corresponding client device (e.g., of a user captured by at least one camera); generating a user graphic representation from the live data feed; inserting the user graphic representation into a three-dimensional coordinate position in the virtual environment; updating the user graphic representation in the virtual environment from the live data feed; and processing data generated from interactions in the virtual environment. Such interactions may include, in particular, social interactions in the virtual environment. For such interactions, the method may include providing the updated virtual environment to the client device directly via P2P communication or indirectly through the use of one or more cloud servers, enabling real-time multi-user collaboration and interaction in the virtual environment.

[0026] In some embodiments, the system further enables the creation of ad-hoc virtual communications (e.g., via a virtual environment platform), which may include creating an ad-hoc voice communication channel between user graphic representations without the need to change the current viewing perspective or position in the virtual environment. For example, user graphic representations can approach each other and conduct an ad-hoc voice conversation at a location within the virtual environment where both user graphic representation areas exist. Such communications can be enabled, for example, by taking into account the distance, position, and orientation between user graphic representations, and / or their current availability status (e.g., available or unavailable), or the status configuration of such ad-hoc communications, or combinations thereof. Approaching user graphic representations will see visual feedback regarding other user graphic representations signaling that ad-hoc communication is possible, thus setting the start of a conversation between both user graphic representations, in which case the approaching user can speak and the other user can listen and respond. In another embodiment, the virtual environment platform enables ad-hoc virtual communications in the virtual environment through the processing of data generated in response to steps performed by client devices, which may include steps of approaching a user graphic representation, selecting and clicking on a user graphic representation, sending or receiving an invitation to participate in an ad-hoc virtual communication with another user graphic representation, and accepting the received invitation. In such scenarios, the platform can open a communication channel between user client devices, and user graphic representations can converse in the virtual space of the virtual environment.

[0027] In some embodiments, the method further includes engaging one or more users in a conversation, transitioning a user graphic representation from a user 3D virtual cutout to a user real-time 3D virtual cutout, or a video with the background removed, or a video with the background not removed, and opening a peer-to-peer (P2P) communication channel between user client devices. In one embodiment, the step of engaging two or more users in a conversation includes approaching the user graphic representation, selecting and clicking on the user graphic representation, sending or receiving an invitation to participate in a conversation with another user graphic representation, and accepting the received invitation. The step of opening a communication channel between user client devices can be performed when processing and rendering are done by the client devices, or the step of opening an indirect communication channel through one or more cloud server computers can be performed when processing and rendering are done on at least one cloud server computer or between at least one cloud server and a client device. In further embodiments, the conversation includes sending and receiving real-time audio between the participating users' 3D virtual cutouts. In further embodiments, the conversation includes sending and receiving real-time audio and video displayed from the participating users' real-time 3D virtual cutouts, or videos with the background removed, or videos with the background not removed.

[0028] In some embodiments, the method for enabling interaction in a virtual environment further includes embedding clickable links that redirect to the virtual environment in one or more third-party sources including a third-party website, application, or video game.

[0029] In another aspect of the present disclosure, a data processing system comprises one or more computing devices including at least one cloud server computer, the one or more computing devices comprising at least one processor and a memory storing data and instructions for implementing an image processing function, and the one or more computing devices of the data processing system are configured to generate a user graphic representation from a live data feed by a combination of image processing of at least one cloud server computer and one or more of two or more client devices in a hybrid system architecture. In one embodiment, the system comprises two or more client devices communicatively connected to each other and to one or more cloud server computers via a network, the two or more client devices comprising at least one processor, a memory storing data and instructions for implementing an image and media processing function, and at least one camera connected to at least one of the client devices and to one or more cloud server computers for obtaining a live data feed from at least one user of the at least one client device. The user graphic representation is generated from the live data feed by a combination of image and media processing of one or more cloud server computers and one or more of the two or more client devices. The one or more cloud server computers and the one or more client devices interact through a hybrid system architecture.

[0030] In some embodiments, the data used as input data for the data processing system may include 2D or 3D image data, 3D geometry, video data, media data, audio data, text data, tactile data, time data, 3D entities, 3D dynamic objects, text data, time data, metadata, priority data, security data, location data, lighting data, depth data, and infrared data, among others.

[0031] In some embodiments, the hybrid system architecture comprises a client-server side and a peer-to-peer (P2P) side. In one embodiment, the client-server side comprises a web or application server. The client-server side may further be configured to include a secure communication protocol, microservices, a database management system, a database, and / or a distributed message and resource delivery platform. The server-side components may be provided together with client devices that communicate with the server over the network. The client-server side defines the interaction between one or more clients and the server over the network, including any processing performed by the client side, the server side, or the receiving client side. In one embodiment, one or more of the corresponding clients and servers perform the necessary image and media processing according to various combinations of rule-based task assignments. In one embodiment, the web or application server is configured to receive client requests using a secure communication protocol and process the client requests by requesting microservices or data corresponding to requests from a database using a database management system. The microservices are delivered using a distributed message and resource delivery platform using a publish-subscribe model.

[0032] The P2P side includes a P2P communication protocol that enables real-time communication between client devices in a virtual environment, and a rendering engine configured to enable a client device to perform real-time 3D rendering of live session elements (e.g., user graphic representations) included in the virtual environment. In one embodiment, the P2P side further includes a computer vision library configured to enable a client device to perform real-time computer vision tasks in the virtual environment. By using such a hybrid model of communication, rapid P2P communication between users becomes possible, the problem of latency is reduced while providing web services and resources to each session, and multiple interactions between users and with content in the virtual environment can be enabled.

[0033] The P2P side defines interactions between client devices and any processes that can be executed by one or the other client device from the P2P side. In some embodiments, the P2P side is used for video and data processing tasks, and synchronization, streaming, and rendering between client devices. In other embodiments, the P2P side is used for video streaming, rendering, and synchronization between client devices, while the client-server side is used for data processing tasks. In further embodiments, the client-server side is used for video streaming together with data processing tasks, while the P2P side is used for video rendering and synchronization between client devices. In yet another embodiment, the client-server side is used for video streaming, rendering, data processing tasks, and synchronization.

[0034] In one embodiment, the data processing tasks include generating a user graphic representation and inserting the user graphic representation into the virtual environment. Generating the user graphic representation may include performing background removal or other processing or enhancements.

[0035] In some embodiments, the P2P-side data is sent directly from one client device to a peer client device, or vice versa, or relayed through a server via the client-server side.

[0036] In some embodiments, at least one cloud server is an intermediate server, which means that the server is used to facilitate and / or optimize the exchange of data between client devices. In such embodiments, at least one cloud server manages, analyzes, processes, and optimizes incoming images and multimedia streams, and can manage, evaluate, and optimize the transfer of outgoing streams as a router topology (e.g., but not limited to, SFU (Selective Forwarding Units), SAMS (Spatially Analyzed Media Server), multimedia server router, or processing of images and media (e.g., but not limited to, decode, combine, improve, mix, enhance, extend, compute, operate, encode)), and a transfer server topology (e.g., but not limited to, multipoint control unit (MCU), cloud media mixer, cloud 3D renderer, etc.), or other server topologies.

[0037] In such an embodiment where the intermediate server is a SAMS, such a media server manages, analyzes, and processes the incoming data (e.g., but not limited to, metadata, priority data, data class, spatial structure data, three-dimensional position, orientation, or movement information, images, media, scalable video codec-based video) of each transmitting client device. In such an analysis, based on the priority relationship of such incoming data to achieve the optimal bandwidth and computing resource utilization rate for receiving one or more user client devices in terms of the spatial three-dimensional orientation, distance, and a specific receiving client device user, the media is changed, upscaled, or downscaled in terms of time (various frame rates), space (e.g., different image sizes), quality (e.g., quality based on different compression or encoding), and color (e.g., color resolution and range), so as to optimize the transfer of the outgoing data stream to each receiving client device.

[0038] In some embodiments, a plurality of image processing tasks are classified based on which of the client device, cloud server, and / or receiving client device performs them, and thus are classified as client device image processing, server image processing, and receiving client device image processing. The plurality of image processing tasks can be executed on the client-server side, P2P side, or a combination thereof of a hybrid architecture. The image processing tasks include background removal, further processing or improvement, and insertion and combination into a virtual environment. A combination of three image processing tasks can be used in generating, improving, and inserting / combining into a virtual environment the user graphic representation. The combination of image processing and the corresponding usage levels of client device processing, server image processing, and receiving client device processing depend on the amount of data to be processed, the waiting time allowed to maintain a smooth user experience, the desired quality of service (QOS), the required services, and the like. The following are eight such combinations of image processing performed on the client-server side.

[0039] In some embodiments, at least one of the client devices is configured to generate a user graphic representation with a combination of image processing on the client - server side, perform background removal, and transmit the user graphic representation with the background removed to at least one cloud server for further processing. In a first exemplary combination of image processing, the client device generates a user graphic representation including background removal and transmits the user graphic representation with the background removed to at least one cloud server for further processing or improvement to generate an enhanced user graphic representation with the background removed. The at least one cloud server transmits the enhanced user graphic representation with the background removed to the receiving client device, and the receiving client device inserts and combines the enhanced user graphic representation with the background removed into a virtual environment.

[0040] In a second exemplary combination of image processing, the client device generates a user graphic representation including background removal, performs further processing to generate an enhanced user graphic representation with the background removed, and then transmits this to at least one cloud server. The at least one cloud server transmits the enhanced user graphic representation with the background removed to the receiving client device, and the receiving client device inserts and combines the enhanced user graphic representation with the background removed into a virtual environment.

[0041] In a third exemplary combination of image processing, the client device generates a user graphic representation including background removal, performs further processing to generate an enhanced user graphic representation with the background removed, inserts and combines the enhanced user graphic representation with the background removed into a virtual environment. Then, the client device transmits the enhanced user graphic representation with the background removed inserted and combined into the virtual environment to a cloud server for relaying to the receiving client device.

[0042] In a fourth exemplary combination of image processing, the client device generates a user graphic representation including background removal and transmits the user graphic representation with the background removed to at least one cloud server for further processing to generate an enhanced user graphic representation with the background removed. The at least one cloud server then inserts and combines the enhanced user graphic representation with the background removed into a virtual environment and transmits it to the receiving client device.

[0043] In a fifth exemplary combination of image processing, the client device generates a user graphic representation including background removal and transmits the user graphic representation with the background removed to at least one cloud server for relaying to the receiving client device. The receiving client device performs further processing on the user graphic representation with the background removed to generate an enhanced user graphic representation with the background removed, inserts it into a virtual environment, and combines them.

[0044] In a sixth exemplary combination of image processing, the client device transmits a camera live data feed received from at least one camera, transmits the unprocessed data to at least one cloud server, and the at least one cloud server generates a user graphic representation including background removal, performs further processing on the user graphic representation with the background removed to generate an enhanced user graphic representation with the background removed, and transmits it to the receiving client device. The receiving client device inserts and combines the enhanced user graphic representation with the background removed into a virtual environment.

[0045] In a seventh exemplary combination of image processing, a client device transmits a camera live data feed received from at least one camera and sends the unprocessed data to at least one cloud server. The at least one cloud server generates a user graphic representation including background removal, performs further processing on the user graphic representation with the background removed to generate an enhanced user graphic representation with the background removed, and then inserts and combines the enhanced user graphic representation with the background removed into a virtual environment and sends it to the receiving client device.

[0046] In an eighth exemplary combination of image processing, a client device transmits a camera live data feed received from at least one camera and sends the unprocessed data to at least one cloud server for relaying to the receiving client device. The receiving client device uses the data to generate a user graphic representation including background removal, performs further processing on the user graphic representation with the background removed to generate an enhanced user graphic representation with the background removed, and then inserts and combines the enhanced user graphic representation with the background removed into a virtual environment.

[0047] In some embodiments, when client-server side data is relayed through at least one cloud server, the at least one cloud server is configured as a Traversal Using Relay NAT (TURN) server. TURN can be used in the case of symmetric Network Address Translation (NAT) and can remain in the media path even after a connection is established while processed and / or unprocessed data is being relayed between client devices.

[0048] The following is an explanation of three exemplary combinations of image processing performed on the P2P side by either or both of the first and second peer devices.

[0049] In the first combination of image processing, the first peer device generates a user graphic representation including background deletion, performs further processing to generate an enhanced user graphic representation with the background deleted, and inserts and combines the enhanced user graphic representation with the background deleted into the virtual environment. Then, the first peer device transmits the enhanced user graphic representation with the background deleted that has been inserted and combined into the virtual environment to the second peer device.

[0050] In the second combination of image processing, the first peer device generates a user graphic representation including background deletion and transmits the user graphic representation with the background deleted to the second peer device. The second peer device performs further processing to generate an enhanced user graphic representation with the background deleted for the user graphic representation with the background deleted, inserts it into the virtual environment, and combines it.

[0051] In the third combination of image processing, the first peer device transmits the camera live data feed received from at least one camera and transmits the unprocessed data to the second peer device. The second peer device uses the data to generate a user graphic representation including background deletion, performs further processing to generate an enhanced user graphic representation with the background deleted for the user graphic representation with the background deleted, and then inserts and combines the enhanced user graphic representation with the background deleted into the virtual environment.

[0052] In some embodiments, the combination of the three image processes on the P2P side may further include relaying data through at least one cloud server. In these embodiments, at least one cloud server can be configured as a STUN server, whereby the peer devices can detect their public IP addresses and the types of NATs behind them that can be used to establish data connections and data exchanges between the peer devices. In another embodiment, at least one cloud server computer can be configured for signaling, which can be used for peer devices to identify and connect to each other and to exchange data through communication coordination performed by at least one cloud server.

[0053] In some embodiments, media, video, and / or data processing tasks include one or more of image filtering, computer vision processing, image sharpening, background improvement, background removal, foreground blurring, eye covering, face pixelation, voice distortion, image upscaling, image cleansing, skeleton analysis, face or head counting, object recognition, marker or QR code tracking, fiducial tracking, feature analysis, 3D mesh or volume generation, feature tracking, face recognition, SLAM tracking, and face expression recognition, or other modular plugins in the form of microservices executed on such media routers or servers, including one or more of encoding, transcoding, decoding, spatial or 3D analysis and processing.

[0054] In some embodiments, background removal includes the adoption of image segmentation through one or more of instance segmentation or semantic segmentation and the use of a deep neural network.

[0055] In some embodiments, one or more computing devices of the data processing system are further configured to insert a user graphic representation into a virtual environment by generating a virtual camera, and generating the virtual camera includes associating the captured viewing perspective data with the viewing perspective of the user graphic representation within the virtual environment. In one embodiment, inserting and combining the user graphic representation into the virtual environment includes virtually generating and aligning one or more virtual cameras that are placed in front of the user graphic representation, for example, in front of a video with a removed background, or a video with an unremoved background, or a user 3D virtual cutout, or a user real-time 3D virtual cutout. In one embodiment, one or more virtual cameras can be directed outward from eye level. In another embodiment, two (one for each eye) virtual cameras can be directed outward from the height of both eyes. In yet another embodiment, one or more virtual cameras can be directed outward from the center of the position of the head of the user graphic representation. In yet another embodiment, one or more virtual cameras can be directed outward from the center of the user graphic representation. In yet another embodiment, one or more virtual cameras, when in a self-viewing perspective, may be placed in front of the user graphic representation, for example, at the height of the head of the user graphic representation, and directed towards the user graphic representation.

[0056] In one embodiment, one or more virtual cameras are created by, at least, using computer vision to associate the captured viewing perspective data of the user with the viewing perspective of the user graphic representation within the virtual environment.

[0057] In another aspect of the present disclosure, an image processing method includes providing data and instructions implementing an image processing function to a memory of at least one cloud server computer, and generating a user graphic representation in a virtual environment based on a live data feed from at least one client device by one or more combinations of image processing of at least one cloud server computer and at least one client device, wherein the at least one cloud server computer interacts with the at least one client device through a hybrid system architecture. In one embodiment, the method includes obtaining a live data feed from at least one camera from at least one user of at least one corresponding client device, and generating a user graphic representation by one or more combinations of image processing of one or more cloud server computers and at least one client device. The one or more cloud server computers and the at least one client device can interact through the hybrid system architecture of the present disclosure including a P2P side and a client-server side.

[0058] In some embodiments, the method includes, on the P2P side, performing video and data processing and synchronization, streaming, and rendering between client devices. In further embodiments, the method includes, on the P2P side, performing video streaming, rendering, and synchronization between client devices while the client-server side is being used for data processing. In further embodiments, the method includes, on the client-server side, performing video streaming together with data processing while the P2P side is being used for video rendering and synchronization between client devices. In yet another embodiment, the method includes, on the client-server side, performing video streaming, rendering, and data processing and synchronization.

[0059] In some embodiments, the data processing task includes generating a user graphic representation and inserting the user graphic representation into a virtual environment. In one embodiment, the data processing task includes generating a user graphic representation that first performs background removal, then performing further processing, and then inserting and combining it into a virtual environment. In yet another embodiment, the image processing task is executed through a combination of multiple image processes of a client device and a cloud server computer on the client-server side or the P2P side.

[0060] In some embodiments, inserting a user graphic representation into a virtual environment includes generating a virtual camera, and generating the virtual camera includes associating the captured viewing perspective data with the viewing perspective of the user graphic representation within the virtual environment. In one embodiment, inserting and combining a user graphic representation into a virtual environment includes generating one or more virtual cameras that are virtually positioned and aligned in front of the user graphic representation, for example, in front of a video with a removed background, or a video with an unremoved background, or a user 3D virtual cutout, or a user real-time 3D virtual cutout. In one embodiment, one or more virtual cameras can be directed outward from eye level. In another embodiment, two (one for each eye) virtual cameras can be directed outward from both eye levels. In yet another embodiment, one or more virtual cameras can be directed outward from the center of the position of the head of the user graphic representation. In yet another embodiment, one or more virtual cameras can be directed outward from the center of the user graphic representation. In yet another embodiment, one or more virtual cameras can be positioned in front of the user graphic representation, for example, at the height of the user's head, and directed towards the user graphic representation when in a self-viewing perspective. The virtual camera is created by, at least, using computer vision to associate the captured user viewing perspective data with the viewing perspective of the user graphic representation within the virtual environment.

[0061] In some embodiments, the method further includes embedding an embedded clickable link in the user graphic representation, and the embedded clickable link, in response to a click, directs to a third-party source that includes profile information regarding the corresponding user.

[0062] In another aspect of the present disclosure, a user graphic representation-based user authentication system includes one or more cloud server computers comprising at least one processor and a memory storing data and instructions, the memory storing a user database storing user data associated with user accounts and one or more corresponding user graphic representations, and a face scan and authentication module connected to the database. The one or more cloud server computers authenticate a user by performing a face scan of the user, which includes extracting face feature data from camera data received from a client device through the face scan and authentication module, and checking the extracted face feature data for a match with a user graphic representation associated with a user account in the user database. If a matching user graphic representation is found in the checking step, providing the user with access to the corresponding user account. If no matching user graphic representation is found in the checking step, generating a new user graphic representation together with a new user account stored in the user database from the camera data and providing access to the user account.

[0063] In one embodiment, the system includes at least one camera configured to obtain data from a user of at least one client device that requests access to a user account. The at least one camera is connected to the at least one client device and one or more cloud server computers. The one or more cloud server computers perform a face scan of the user through a face scan and authentication module, check a user database for a match with the user graphic representation, and if the user account is verified and available, provide the user with the corresponding user graphic representation along with access to the user account. If the user account is not available, the user is authenticated by generating a new user account stored in the user database along with a new user graphic representation from the data, along with access to the user account.

[0064] The user account can be used, for example, to access a virtual environment platform or any other application such as an interactive application, game, email account, university profile account, work account, etc. (e.g., an application that can be linked to an environmental platform). The user graphic representation-based user authentication system of the present disclosure provides a higher security level than a standard camera-based face detection authentication system, especially when additional authentication steps of generating a user graphic representation or obtaining an existing user graphic representation from a user database are provided.

[0065] In some embodiments, the user graphic representation is a user 3D virtual cutout, or a user real-time 3D virtual cutout with the background removed, or a video with the background removed, or a video without the background removed. In one embodiment, the user graphic representation is a user 3D virtual cutout constructed from a photo uploaded by the user or provided by a third party, or a user real-time 3D virtual cutout with the background removed, or a video with the background removed, or a video without the background removed, generated based on real-time 2D or 3D live video stream data feeds obtained from a camera. In some embodiments, one or more cloud server computers are further configured to animate a matching user graphic representation or a new user graphic representation. Animating a matching user graphic representation includes applying a machine vision algorithm by a client device or at least one cloud server computer to each user graphic representation to recognize the facial expression of the user and graphically simulate the facial expression in the user graphic representation. In a further embodiment, updating a user 3D virtual cutout constructed from a photo uploaded by the user or provided by a third party includes applying a machine vision algorithm by a client device or at least one cloud server computer to the generated user 3D virtual cutout to recognize the facial expression of the user and graphically simulate the facial expression on the user 3D virtual cutout.

[0066] In some embodiments, one or more cloud server computers are further configured to check the date of a matching user graphic representation and determine whether an update to the matching user graphic representation is needed. In one embodiment, in response to checking the date of the user graphic representation available when the user account is available, one or more cloud server computers determine whether an update to the existing user graphic representation is needed by comparing it to a corresponding threshold or security requirement. For example, if there has been a system security update, it may be necessary to update all user graphic representations or at least those created before a specified date. If a user graphic representation is needed, one or more cloud server computers generate a request to update the user graphic representation for the corresponding client device. If the user approves the request, one or more cloud server computers or the client device proceed to generate the user graphic representation based on the live camera feed. If an update is not needed, one or more cloud server computers proceed to retrieve the existing user graphic representation from the user database.

[0067] In some embodiments, the user graphic representation is inserted into a two - dimensional or three - dimensional virtual environment or linked to a third - party source inserted into the virtual environment (e.g., by overlaying it on the screen of a third - party application or website integrated or coupled with the system of the present disclosure) and graphically combined with the two - dimensional or three - dimensional virtual environment.

[0068] In some embodiments, the process of generating user graphic representations is performed asynchronously with respect to user access to the user account. For example, if the system determines that the user is already authenticated after performing user graphic representation-based face scanning and detection, the system may allow the user to access the user account during the generation of a new user graphic representation so that it can be provided to the user and inserted into and combined with the virtual environment as soon as it is ready.

[0069] In some embodiments, one or more cloud server computers are further configured to authenticate the user through a login authentication credential that includes a personal identification number (PIN), or a username and password, or a combination thereof.

[0070] In some embodiments, authentication is triggered in response to the activation of an invitation link or deep link sent from one client device to another client device. In one embodiment, clicking on the invitation link or deep link triggers at least one cloud server computer to request user authentication. For example, the invitation link or deep link is for an invitation to a phone call, a conference call, or a video game session, and the invited user can be authenticated through the user graphic representation-based authentication system of the present disclosure.

[0071] In another embodiment, the face scan uses 3D authentication that includes guiding the user to execute a head movement pattern and extracting 3D face data based on the head movement pattern. This can be done using application instructions stored in at least one server computer that guides the user to execute a head movement pattern, such as executing one or more head gestures, tilting the head horizontally or vertically, or rotating it in a circular motion, or a gesture pattern generated by the user, or a specific head movement pattern, or a combination thereof, to perform 3D authentication. 3D authentication recognizes not only one view or image for comparison and analysis, but also further features from the data obtained from the live video data feed of the camera. In this embodiment of 3D authentication, the face scan process can recognize further features from data that may include face data including a head movement pattern, face volume, height, depth of face features, face injuries, tattoos, eye color, face skin parameters (e.g., skin color, wrinkles, pore structure, etc.), reflectance parameters, and also the position of such features on the face topology, as in the case of other types of face detection systems. Thus, such capture of face data can increase the capture of a realistic face that can be useful for generating a realistic user graphic representation. The face scan using 3D authentication can be performed using a high-resolution 3D camera, depth camera (e.g., LIDAR), light field camera, etc. The face scan process and 3D authentication can use deep neural networks, convolutional neural networks, and other deep learning techniques to obtain, process, and evaluate user authentication by using face data.

[0072] In another aspect of the present disclosure, a user authentication method based on user graphic representation stores user data associated with a user account and one or more corresponding user graphic representations in the memory of one or more cloud server computers, provides a face scan and authentication module connected to the user database, receives an access request to the user account from a client device, performs a face scan of the user of the client device through the face scan and authentication module by extracting face feature data from camera data captured by at least one camera communicating with the client device, checks the extracted face feature data for a match with the user graphic representation associated with the user account in the user database, provides access to the user account to the user if a matching user graphic representation is found in the check step, and if no matching user graphic representation is found in the check step, generates a new user graphic representation together with a new user account stored in the user database from the camera data and provides access to the user account.

[0073] In one embodiment, the method performs a face scan of the user of at least one client device through a face scan and authentication module by using images and / or media data received from at least one client device and at least one camera connected to one or more cloud server computers, checks the user database for a match of user face data associated with the user account, provides the user with access to the user account and the corresponding user graphic representation if the user account is available, and generates a new user graphic representation together with a new user account stored in the user database and access to the user account from the face data if the user account is not available.

[0074] In some embodiments, the user graphic representation includes a user 3D virtual cutout constructed from a photograph uploaded by the user or provided by a third party, or a user real-time 3D virtual cutout with a background-removed user real-time video stream generated based on real-time 2D or 3D live video stream data feeds obtained from a camera, or a video with a removed background, or a video without a removed background. In further embodiments, the method includes animating a matching user graphic representation or a new user graphic representation, which may include applying a machine vision algorithm by the client device or at least one cloud server computer to each user graphic representation to recognize the user's facial expression and graphically simulate the facial expression in the user graphic representation. In one embodiment, updating the user 3D virtual cutout includes applying a machine vision algorithm by the client device or at least one cloud server computer to the generated user 3D virtual cutout to recognize the user's facial expression and graphically simulate the facial expression on the user 3D virtual cutout.

[0075] In some embodiments, the method further includes, if a matching user graphic representation is found in the checking step, checking the date of the matching user graphic representation, determining whether an update of the matching user graphic representation is necessary based at least in part on the date, and, in the affirmative case that an update of the matching user graphic representation is necessary, generating an update request for the user graphic representation. In one embodiment, the method includes, if a user account is available, checking the date of the available user graphic representation, determining whether an update of the existing user graphic representation is necessary by comparing it with a corresponding threshold or security requirement, and, in the affirmative case that an update of the user graphic representation is required, generating an update request for the user graphic representation and sending it to the corresponding client device. If the user approves the request, one or more cloud server computers or client devices proceed to generate a user graphic representation based on the live camera feed. If an update is not required, one or more cloud server computers proceed to obtain the existing user graphic representation from the user database.

[0076] In some embodiments, the method further includes inserting the user graphic representation into a two-dimensional or three-dimensional virtual environment or into a third-party source linked to the virtual environment (e.g., by overlaying it on the screen of a third-party application or website integrated or combined with the system of the present disclosure), and combining the user graphic representation with the two-dimensional or three-dimensional virtual environment.

[0077] In some embodiments, the process of generating a new user graphic representation is performed asynchronously with respect to user access to the user account.

[0078] In some embodiments, the method further includes authenticating the user through login credentials including at least a username and password.

[0079] In some embodiments, authentication is triggered in response to activation of an invitation link. In one embodiment, the method further includes providing an invitation link or deep link from one client device to another client device, and clicking the invitation link triggers at least one cloud server to request user authentication.

[0080] In another aspect of the present disclosure, a system for virtual broadcasting from within a virtual environment is provided. The system includes a server computer system including one or more server computers, each server computer including at least one processor and a memory, the server computer system including data and instructions implementing a data exchange management module configured to manage data exchange between client devices, and including at least one virtual environment disposed within at least one virtual environment and including at least one virtual broadcast camera configured to capture a multimedia stream from within the at least one virtual environment. The server computer system is configured to receive live feed data captured by at least one camera from at least one client device and to broadcast a multimedia stream to at least one client device based on data exchange management, and the broadcast multimedia stream is configured to be displayed in a corresponding user graphic representation generated from a live data feed of a user from at least one client device. The management of data exchange between client devices by the data exchange management module includes analyzing an incoming multimedia stream and evaluating the transfer of an outgoing multimedia stream based on the analysis of the incoming media stream.

[0081] In one embodiment, the multimedia stream is sent to at least one media server computer for broadcasting to at least one client device. In one embodiment, the system includes at least one camera that obtains live feed data from a user of at least one client device and sends the live feed data from the user via the at least one client device to at least one media computer. The multimedia stream is broadcast to at least one client device based on data exchange management from at least one media server computer, is displayed in a corresponding user graphic representation generated from the user's live data feed through the at least one client device, and the data exchange management between client devices by the data exchange management module includes analyzing and optimizing the incoming multimedia stream and evaluating and optimizing the transfer of the outgoing multimedia stream.

[0082] In some embodiments, when transferring an outgoing multimedia stream, the server computer system uses a routing topology that includes a Selective Forwarding Unit (SFU), Traversal Using Relay NAT (TURN), SAMS, or other suitable multimedia server routing topology, or a media processing and transfer server topology, or other suitable server topology.

[0083] In some embodiments, the server computer system processes an outgoing multimedia stream using a media processing topology to display user graphic representations within at least one virtual environment through a client device. In one embodiment, at least one media server computer is configured to decode, combine, improve, mix, enhance, extend, compute, manipulate, and encode a multimedia stream to a related client device to display user graphic representations within at least one virtual environment when using the media processing topology.

[0084] In some embodiments, the server computer system uses one or more of an MCU, a cloud media mixer, and a cloud 3D renderer when using a transfer server topology.

[0085] In some embodiments, an incoming multimedia stream includes user priority data and distance relationship data, and the user priority data includes a higher priority score for a user graphic representation closer to the source of the incoming multimedia stream and a lower priority score for a user graphic representation farther from the source of the incoming multimedia stream. In one embodiment, the multimedia stream includes data related to user priority and the distance relationship between the corresponding user graphic representation and the multimedia stream, and the data includes metadata, or priority data, or data class, or spatial structure data, or three-dimensional position, or orientation or movement information, or image data, or media data, and scalable video codec-based video data, or a combination thereof. In a further embodiment, the priority data includes a higher priority score for a user closer to the multimedia stream source and a lower priority score for a user farther from the multimedia stream source. In yet another embodiment, the transfer of an outgoing multimedia stream is based on user priority data and distance relationship data. In one embodiment, the transfer of an outgoing multimedia stream performed by a media server based on user priority data and distance relationship data includes bandwidth optimization and calculation of the resource utilization rate of one or more receiving client devices. In yet another embodiment, the transfer of an outgoing multimedia stream further includes modifying, upscaling, or downscaling the multimedia stream for temporal characteristics, spatial characteristics, quality characteristics, and / or color characteristics.

[0086] In some embodiments, a virtual broadcast camera is managed through a client device accessing a virtual environment. In one embodiment, the virtual broadcast camera is configured to manipulate the viewpoint of the camera updated in the virtual environment and broadcast the updated viewpoint to at least one client device.

[0087] In some embodiments, at least one virtual environment comprises a plurality of virtual broadcast cameras, each virtual broadcast camera providing a multimedia stream from a corresponding perspective within the at least one virtual environment. In one embodiment, each virtual broadcast camera is selected by a user of at least one client device and switched between one another to provide a corresponding perspective to a corresponding at least one user graphical representation, providing a multimedia stream from a corresponding perspective within the virtual environment.

[0088] In some embodiments, at least one virtual environment is hosted by at least one dedicated server computer connected via a network to at least one media server computer or is hosted in a peer-to-peer infrastructure and relayed through at least one media server computer.

[0089] In another aspect of the present disclosure, a method for virtual broadcasting from within a virtual environment includes providing data and instructions in the memory of at least one media server that implement a client device data exchange management module for managing data exchange between client devices; capturing a multimedia stream with a virtual broadcast camera disposed within at least one virtual environment connected to at least one media server; sending the multimedia stream to at least one media server for broadcasting to at least one client device; obtaining live feed data from at least one client device (e.g., from at least one camera via at least one client device); performing data exchange management including analyzing an incoming multimedia stream and live feed data from within at least one virtual environment and evaluating the transfer of an outgoing multimedia stream; and broadcasting a corresponding multimedia stream to a client device based on the data exchange management, wherein the multimedia stream is displayed in a user graphic representation of a user of at least one client device. In this context, this refers to what can be "seen" based on their positions in the virtual environment by the user graphic representations, which corresponds to what is displayed to the user (via the client device) when viewing the virtual environment from the perspective of their own user graphic representation.

[0090] In some embodiments, when transferring an outgoing multimedia stream, the method uses a routing topology including a SFU, TURN, SAMS, or other suitable multimedia server routing topology, or a media processing and transfer server topology, or other suitable server topology.

[0091] In some embodiments, when using a media processing topology, the method further includes decoding, combining, improving, mixing, enhancing, expanding, computing, manipulating, and encoding the multimedia stream.

[0092] In some embodiments, when using a transfer server topology, the method further includes using one or more of a multipoint control unit (MCU), a cloud media mixer, and a cloud 3D renderer.

[0093] In some embodiments, an incoming multimedia stream includes user priority data and distance relationship data, and the user priority data includes a higher priority score for a user graphic representation closer to the source of the incoming multimedia stream and a lower priority score for a user graphic representation farther from the source of the incoming multimedia stream. In one embodiment, the method further includes optimizing the transfer of an outgoing multimedia stream by a media server based on the user priority data and the distance relationship data, which may include bandwidth optimization and calculation of the resource usage rate of one or more receiving client devices. In a further embodiment, the optimization of the transfer of the outgoing multimedia stream by the media server further includes modifying, upscaling, or downscaling the multimedia stream for time characteristics, spatial characteristics, quality characteristics, and / or color characteristics.

[0094] In some embodiments, at least one virtual environment includes a plurality of virtual broadcast cameras, and each virtual broadcast camera provides a multimedia stream from a corresponding perspective within at least one virtual environment. In one embodiment, the method further includes providing a plurality of virtual broadcast cameras that are each selected by a user of at least one client device, switched between each other, and provide corresponding perspectives from corresponding perspectives within the virtual environment that can provide corresponding perspectives to at least one corresponding user graphic representation.

[0095] In another aspect of the present disclosure, a system for delivering applications within a virtual environment, comprising at least one cloud server computer having at least one processor and a memory including data and instructions for implementing at least one virtual environment linked to an application module including one or more installed applications and application rules for corresponding multi-user interactions, in response to a selection by a virtual environment host through a client device, during a session of the virtual environment, one or more installed applications are displayed and activated, enabling interaction between a user graphic representation of the virtual environment host and any participating user graphic representations within the virtual environment and one or more installed applications through a corresponding client device, and the at least one cloud server computer manages and processes received user interactions with one or more installed applications according to application rules for multi-user interactions in the application module, and transfers the processed interactions accordingly (e.g., to each client device) to establish a multi-user session enabling a shared experience according to the multi-user interaction application rules. A system is provided.

[0096] In some embodiments, the application rules for multi-user interactions are stored and managed on one or more separate application servers.

[0097] In some embodiments, one or more applications are installed from application install packages available from an application library and provision application services through corresponding application programming interfaces.

[0098] In some embodiments, the application library is filtered by context. In one embodiment, the context filtering is designed to provide applications related to a particular context.

[0099] In some embodiments, one or more installed applications are shared with and displayed through a virtual display application installed on the corresponding client device. In one embodiment, when installed and activated, one or more installed applications are shared with and displayed through a virtual display application installed on the corresponding client device, and the virtual display application receives one or more installed applications from the application library and is configured to publish one or more selected applications to display the conference host user graphic representation and other participant user graphic representations in the virtual environment through their corresponding client devices. In a further embodiment, the application module is represented as a 2D screen or 3D volume application module graphic representation that displays the content from the installed application as a user graphic representation in the virtual environment, and the virtual display application is represented as a 2D screen or 3D volume that displays the content from the installed application as a user graphic representation in the virtual environment.

[0100] In some embodiments, one or more applications are installed directly into the virtual environment before or at the same time as a multi-user session is conducted.

[0101] In some embodiments, one or more applications are installed through the use of a virtual environment setup tool before starting a multi-user session.

[0102] In some embodiments, one or more of the application rules for multi-user interaction define synchronous interaction, or asynchronous interaction, or a combination thereof. Thus, in one embodiment, such rules are used to update user interaction and the respective updated views of one or more applications.

[0103] In some embodiments, asynchronous interaction is enabled through at least one server computer, or through a separate server computer dedicated to processing individual user interactions with at least one installed application.

[0104] In some embodiments, the virtual environment is a classroom, or an office space, or a conference room, or a reception room, or a theater, or a cinema.

[0105] In another aspect of the present disclosure, a method for delivering an application within a virtual environment, comprising providing in the memory of at least one cloud server computer at least one virtual environment and an application module including one or more installed applications linked to the virtual environment and displayed within the virtual environment and application rules for corresponding multi-user interactions; receiving a selection command from a virtual environment host; enabling, during a session of the virtual environment, a user graphic representation of the virtual environment host and one or more participant user graphic representations to interact with one or more installed applications through corresponding client devices that display and activate the one or more installed applications; receiving user interactions with the one or more installed applications; managing and processing the user interactions with the one or more installed applications according to application rules for multi-user interactions in the application module; and transferring the processed interactions to the client devices to establish a multi-user session that enables a shared experience according to the application rules.

[0106] In some embodiments, the method further comprises storing and managing application rules for multi-user interactions at one or more separate application servers.

[0107] In some embodiments, the method further includes installing one or more applications from an application installation package available from an application library, and provisioning application services through corresponding application programming interfaces. In yet another embodiment, the application library is filtered contextually to provide related applications. In yet another embodiment, one or more installed applications are shared with and displayed through a virtual display application installed on a corresponding client device. In one embodiment, the method includes, when activated, sharing and displaying one or more installed applications through a virtual display application installed on a corresponding client device, the virtual display application receiving one or more installed applications from the application library and configured to expose one or more selected applications to display graphical representations of a conference host user and other participant users in a virtual environment through their corresponding client devices.

[0108] In some embodiments, the method further includes installing one or more applications directly into the virtual environment before or at the same time a multi-user session is conducted. In other embodiments, the method further includes installing one or more applications through the use of a virtual environment setup tool before starting a multi-user session.

[0109] In some embodiments, the method further includes defining one or more of application rules for multi-user interaction to include synchronous interaction, or asynchronous interaction, or a combination thereof. In one embodiment, the method further includes user interaction and appropriately updating respective updated views of one or more applications.

[0110] In another aspect of the present disclosure, a system for provisioning virtual computing resources within a virtual environment comprises one or more server computers including at least one server computer system having at least one processor, a memory including data and instructions for implementing at least one virtual environment, and at least one virtual computer associated with the at least one virtual environment, wherein the at least one virtual computer receives virtual computing resources from the server computer system. The association may include connecting the virtual computer to the virtual environment. In one embodiment, the at least one virtual computer has a corresponding graphic representation within the virtual environment. The graphic representation may provide additional advantages such as facilitating interaction between the user and the virtual computer and enhancing the sense of presence of the user experience (e.g., in the case of a home office experience). Thus, in one embodiment, the at least one virtual computer that receives virtual computing resources from the at least one cloud server computer comprises at least one corresponding associated graphic representation disposed within the virtual environment and at least one client device connected to the at least one server computer through a network, and in response to the at least one client device accessing the one or more virtual computers (e.g., by interacting with the corresponding graphic representation), the at least one cloud server computer provisions at least one portion of the available virtual computing resources to the at least one client device.

[0111] In some embodiments, the server computer system is configured to provision at least one portion of virtual computing resources to at least one client device in response to a user graphic representation that interacts with at least one corresponding graphic representation of at least one virtual computer within at least one virtual environment. In further embodiments, one or more virtual computer graphic representations are spatially arranged within the virtual environment for access by the user graphic representation. In one embodiment, the configuration of the virtual environment is associated with a context theme of the virtual environment, such as the arrangement of virtual items, furniture, floor plans, etc. for use respectively in education, meetings, work, shopping, services, socializing, or entertainment. In further embodiments, one or more virtual computer graphic representations are arranged within the configuration of the virtual environment for access by one or more user graphic representations. For example, a virtual computer may be arranged in a virtual room that a user graphic representation will access when engaging in an activity (such as working on a project in a virtual classroom, laboratory, or office) that requires or may benefit from the ability to use the resources associated with the virtual computer.

[0112] In some embodiments, the server computer system is configured to provision at least one portion of virtual computing resources to at least one client device in response to a user who accesses at least one cloud server computer by logging in to at least one client device without accessing the virtual environment. In an exemplary scenario, the virtual computing resources are accessed by a user who accesses at least one cloud server computer by physically logging in to a client device connected to at least one cloud server computer through a network, triggering the provisioning of the virtual computing resources to the client device without accessing the virtual environment.

[0113] In some embodiments, at least one portion of the virtual computing resources is assigned to the client device by the management tool. In further embodiments, the provisioning of at least a portion of the virtual computing resources is performed based on the stored user profile. In one embodiment, the resource allocation is based on a stored user profile that includes one or more of the parameters associated with and assigned to the user profile, including priority data, security data, QOS, bandwidth, memory space, or computing power, or a combination thereof.

[0114] In some embodiments, at least one virtual computer comprises a downloadable application available from an application library. In an exemplary scenario including a plurality of virtual computers, each virtual computer is a downloadable application available from the application library.

[0115] In another aspect of the present disclosure, a method for provisioning virtual computing resources within a virtual environment includes providing, in the memory of at least one cloud server computer, at least one virtual computer and a virtual environment associated with the at least one virtual computer, associating virtual computing resources with the at least one virtual computer, receiving an access request for accessing one or more virtual computers, and in response to the access request received from at least one client device, provisioning at least a portion of the available virtual computing resources associated with the at least one virtual computer to the at least one client device. In one embodiment, associating virtual computing resources with the at least one virtual computer may include the virtual computer receiving the virtual computing resources from at least one cloud server computer.

[0116] In some embodiments, the access request includes a request that enables a user graphic representation to interact with one or more graphic representations representing at least one virtual computer. In one embodiment, the method further includes receiving, from a user graphic representation, an access request for accessing one or more graphic representations of a virtual computer within at least one virtual environment, and providing at least a portion of the available virtual computing resources to a corresponding client device. In a further embodiment, the configuration of the virtual environment is associated with a context theme of the virtual environment, each including a configuration for use in education, conference, work, shopping, service, social, or entertainment, and one or more virtual computers are disposed within the configuration of the virtual environment for access by one or more user graphic representations.

[0117] In some embodiments, the access request is triggered by a user logging in to at least one client device. In one embodiment, the method further includes receiving an access request from a user physically logging in to a client device connected to at least one cloud server computer through a network, and provisioning virtual computing resources to the client device without accessing the virtual environment.

[0118] In some embodiments, the method further includes allocating at least a portion of the virtual computing resources to the client device with a management tool. In yet another embodiment, the allocation is made based on a stored user profile including one or more of the parameters associated with and assigned to a user profile including priority data, security data, QOS, bandwidth, memory space, computing power, or a combination thereof.

[0119] In another aspect of the present disclosure, a system that enables ad-hoc virtual communication between user graphic representations includes one or more cloud server computers comprising at least one processor, and a memory storing data and instructions for implementing a virtual environment. The virtual environment is configured to enable at least one proximate user graphic representation and at least one target user graphic representation in the virtual environment to open an ad-hoc communication channel and to enable an ad-hoc conversation via the ad-hoc communication channel between user graphic representations within the virtual environment. In one embodiment, the system further comprises two or more client devices connected to the one or more cloud server computers via a network and accessing at least one virtual environment through corresponding user graphic representations, and the virtual environment is configured to enable at least one proximate user graphic representation and at least one target user graphic representation to open an ad-hoc communication channel and to enable an ad-hoc conversation between user graphic representations within the virtual environment.

[0120] In some embodiments, opening an ad-hoc communication channel is based on the distance, position, and orientation between user graphic representations, or the current availability status, privacy settings, or status configuration of the ad-hoc communication, or a combination thereof.

[0121] In some embodiments, the ad-hoc conversation is conducted at a location within the virtual environment where both user graphic representation areas exist. In other embodiments, the ad-hoc conversation is conducted using the current viewing perspective in the virtual environment.

[0122] In some embodiments, the ad-hoc conversation enables an optional change in the viewing perspective, location, or a combination thereof within the same or another connected virtual environment in which the ad-hoc conversation is conducted.

[0123] In some embodiments, one or more cloud server computers are further configured to generate current visual feedback in a virtual environment signaling that ad-hoc communication is possible. In one embodiment, a user graphic representation receives visual feedback signaling that ad-hoc communication is possible, thereby triggering the opening of an ad-hoc communication channel and signaling the start of an ad-hoc conversation between user graphic representations.

[0124] In some embodiments, an ad-hoc conversation includes transmitting and receiving real-time audio and video. In an exemplary scenario, such video can be displayed from a user graphic representation.

[0125] In some embodiments, a user corresponding to an approaching user graphic representation selects and clicks on a target user graphic representation before opening an ad-hoc communication channel. In yet another embodiment, one or more cloud server computers are further configured to open an ad-hoc communication channel in response to an invitation acceptance. For example, a user corresponding to an approaching user graphic representation sends an ad-hoc communication participation invitation to a target user graphic representation and receives approval of the invitation from the target user graphic representation before opening an ad-hoc communication channel.

[0126] In some embodiments, an ad-hoc communication channel is enabled through at least one cloud server computer or as a P2P communication channel.

[0127] In another aspect of the present disclosure, a method for enabling ad-hoc virtual communication between user graphic representations includes providing a virtual environment in the memory of one or more cloud server computers comprising at least one processor, detecting two or more client devices connected to the one or more cloud server computers via a network and accessing at least one virtual environment through corresponding graphic representations, and opening an ad-hoc communication channel and enabling an ad-hoc conversation between user graphic representations in the virtual environment in response to at least one user graphic representation approaching another user graphic representation.

[0128] In some embodiments, the method further includes detecting and evaluating one or more of the distance, position, and orientation between user graphic representations, or the current availability status, privacy settings, or status configuration of the ad-hoc communication, or a combination thereof, before opening the ad-hoc communication channel.

[0129] In some embodiments, the method enables the ad-hoc conversation to occur at a location within the virtual environment where both user graphic representation areas exist. In other embodiments, the ad-hoc conversation is conducted using the current viewing perspective in the virtual environment.

[0130] In some embodiments, the method includes enabling an optional change in the viewing perspective, location, or a combination thereof, within the same or another connected virtual environment in which the ad-hoc conversation can occur.

[0131] In some embodiments, the method further includes generating current visual feedback in a virtual environment that signals that ad-hoc communication is possible. The method may further include transmitting the visual feedback signaling that ad-hoc communication is possible to a target user graphic representation, thereby triggering the opening of an ad-hoc communication channel and signaling the start of a conversation between user graphic representations.

[0132] In some embodiments, the conversation includes transmitting and receiving real-time audio and video displayed from a user graphic representation.

[0133] In some embodiments, the method further includes selecting and clicking on a target user graphic representation as the user graphic representation approaches the target user graphic representation. In yet another embodiment, the ad-hoc communication channel is opened in response to an invitation acceptance. In one embodiment, the method further includes transmitting or receiving an ad-hoc virtual communication participation invitation to or from another user graphic representation before opening the ad-hoc communication channel.

[0134] A computer-readable medium storing instructions configured to cause one or more computers to perform any of the methods described herein is also described.

[0135] The above summary is not intended to include an exhaustive list of all aspects of the present disclosure. The present disclosure is considered to include all systems and methods that can be implemented from all suitable combinations of the various aspects summarized above, as well as those disclosed in the following detailed description and particularly pointed out in the claims filed in this application. Such combinations have advantages not specifically recited in the above summary. Other features and advantages will be apparent from the accompanying drawings and the detailed description that follows.

[0136] Certain features, aspects, and advantages of the present disclosure will be better understood in relation to the following description and the accompanying drawings.

Brief Description of the Drawings

[0137]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

DETAILED DESCRIPTION OF THE INVENTION

[0138] In the following description, reference is made to the drawings which show various embodiments by way of example. Also, various embodiments are described below by referring to some examples. It should be understood that the embodiments may include design and structural changes without departing from the scope of the claimed subject matter.

[0139] The systems and methods of the present disclosure solve at least some of the aforementioned drawbacks, inter alia, by providing a virtual environment platform comprising one or more virtual environments that enable real-time multi-user collaboration and interaction similar to those available in real life, which can be used for meetings, work, education, shopping, and services. The virtual environment can be selected from a plurality of different vertical virtual environments available on the virtual environment platform. A virtual environment cluster can be formed by a combination of virtual environments from the same vertical and / or different verticals, which can include hundreds or thousands of virtual environments. The virtual environment can be a 2D or 3D virtual environment including configurations and appearances associated with a vertical of the virtual environment that can be customized by a user according to preferences or needs. The user can access the virtual environment through a graphic representation that is inserted into the virtual environment and graphically combined with the two-dimensional or three-dimensional virtual environment.

[0140] The user graphic representation can be a user 3D virtual cutout with the background removed constructed from a photo uploaded by the user or provided by a third party, or a user real-time 3D virtual cutout, or a video with the background removed, or a video with the background not removed, and any of these can be switched between each other at any time according to the user's request. The user graphic representation can include a user status that provides current availability related to other users or further details regarding other data. In addition to interactions with objects within the virtual environment, interactions such as conversations and collaborations between users within the virtual environment are enabled. The present disclosure further provides a data processing system and method including a combination of multiple image processes that can be used for generating the user graphic representation. The present disclosure further provides a user graphic representation-based user authentication system and method that can be used to access a virtual environment platform or other applications linked from the virtual environment platform to the virtual environment, a system and method for virtual broadcasting from within the virtual environment, a system and method for distributing applications within the virtual environment, a system and method for provisioning cloud computing-based virtual computing resources within a virtual environment cloud server computer, and a system and method for enabling ad-hoc virtual communication between approaching user graphic representations.

[0141] By enabling virtual presence in the virtual environment and realistic interaction and collaboration between users, it is possible to enhance the sense of presence of remote activities as required, for example, in situations where movement is restricted such as during a pandemic or otherwise. The systems and methods of the present disclosure further enable access to various virtual environments on client devices such as mobile devices or computers without the need for more expensive immersive devices such as extended reality head-mounted displays or expensive new system infrastructures. The client devices or peer devices of the present disclosure may include, for example, computers, headsets, mobile phones, glasses, transparent screens, tablets, and general input devices that incorporate a camera or are connected to a camera and can receive a data feed from the camera.

[0142] FIG. 1 is a schematic diagram of a system 100 that enables social interaction in a virtual environment according to one embodiment.

[0143] The system 100 of the present disclosure that enables interaction in a virtual environment includes at least one processor 104 and a memory 106 that stores data and instructions for implementing a virtual environment platform 108 with at least one virtual environment 110 such as virtual environments A - C. The system includes one or more cloud server computers 102. The one or more cloud server computers are configured to insert a user graphic representation generated from a live data feed obtained by a camera into the three - dimensional coordinate positions of at least one virtual environment, update the user graphic representation in at least one virtual environment, and enable real - time multi - user collaboration and interaction in the virtual environment. In the described embodiment, inserting the user graphic representation into the virtual environment involves graphically combining the user graphic representation with the virtual environment such that the user graphic representation appears in the virtual environment (e.g., at a specified 3D coordinate position). In the example shown in FIG. 1, the system 100 further includes at least one camera 112 that obtains a live data feed 114 from a user 116 of a client device 118. One or more client devices 118 are communicatively connected to the one or more cloud server computers 102 and at least one camera 112 via a network. The user graphic representation 120 generated from the live data feed 114 is inserted into the three - dimensional coordinate positions of the virtual environment 110 (e.g., virtual environment A), graphically combined with the virtual environment, and updated using the live data feed 114. The updated virtual environment is provided to the client device either directly via P2P communication or indirectly through the use of the one or more cloud servers 102. The system 100 enables real - time multi - user collaboration and interaction in the virtual environment 110 by accessing the graphical user interface through the client device 118.

[0144] In FIG. 1, two users 116 (e.g., user A and B respectively) access virtual environment A and interact with the elements therein and with each other through corresponding user graphic representations 120 (e.g., user graphic representations A and B respectively) accessed through corresponding client devices 118 (client devices A and B respectively). Although only two users 116, client devices 118, and user graphic representations 120 are shown in FIG. 1, those skilled in the art will understand that the system may enable multiple users 116 to interact with each other through their corresponding graphic representations 120 via the corresponding client devices 118.

[0145] In some embodiments, virtual environment platform 108 and corresponding virtual environment 110 may enable sharing multiple experiences such as live performances, concerts, webinars, keynote speeches, etc. in real time with multiple (e.g., thousands or even millions) of user graphic representations 120. These virtual performances may be presented and / or multicast to multiple instances of virtual environment 110 to accommodate a large number of users 116 from various regions of the world.

[0146] In some embodiments, client device 118 may be, among other things, one or more of a mobile device, a personal computer, a game console, a media center, and a head-mounted display. Camera 112 may be, among other things, one or more of a 2D or 3D camera, a 360-degree camera, a web camera, an RGBD camera, a CCTV camera, a professional camera, a mobile phone camera, a depth camera (e.g., LIDAR), or a light field camera.

[0147] In some embodiments, virtual environment 110 refers to a virtual structure (e.g., a virtual model) designed through any suitable 3D modeling technique by the CAD (computer assisted drawing) method. In further embodiments, virtual environment 110 refers to a virtual structure scanned from an actual structure (e.g., a physical room) through any suitable scanning tool including an image scan pipeline that inputs various photographs, videos, depth measurements, and / or SLAM (simultaneous location and mapping) scans to generate the virtual environment 110. For example, radar imaging such as synthetic aperture radar, real aperture radar, light detection and ranging (LIDAR), inverse aperture radar, monopulse radar, and other types of imaging techniques can be used to map and model real-world structures and convert them into the virtual environment 110. In other embodiments, virtual environment 110 is a virtual structure modeled after an actual structure (e.g., a real-world room, building, or facility).

[0148] In some embodiments, client device 118 and at least one cloud server computer 102 are connected through a wired or wireless network. In some embodiments, the network can include millimeter wave (mmW) or a combination of mmW and sub-6 GHz communication systems such as 5th generation wireless system communication (5G). In other embodiments, the system can be connected through wireless local area networking (Wi-Fi). In other embodiments, the system can be communicatively connected through 4th generation wireless system communication (4G) supported by a 4G communication system, or can include other wired or wireless communication systems.

[0149] In some embodiments, the processing and rendering involved in generating, updating, and inserting and combining the user graphic representation 120 into the selected virtual environment 110 are performed by at least one processor of the client device 118 upon receiving the live data feed 114 of the user 116. One or more cloud server computers 102 can receive the user graphic representation 120 rendered by the client, insert the user graphic representation 120 rendered by the client into the three-dimensional coordinates of the virtual environment 110, combine the inserted user graphic representation 120 with the virtual environment 110, and send the user graphic representation 120 rendered by the client to the receiving client device. For example, as seen in FIG. 1, client device A can receive the live data feed 114 from each camera 112, process and render the data from the live data feed 114 to generate the user graphic representation A, and then send the user graphic representation A rendered by the client to at least one cloud server computer 102. The at least one cloud server computer 102 can place the user graphic representation A in the three-dimensional coordinates of the virtual environment 110 and then send the user graphic representation A to client device B. A similar process applies to the user graphic representation B from client device B and user B. Thus, both user graphic representations A and B can see and interact with each other within virtual environment A. However, various other combinations of image processing can be enabled through the systems and methods of the present disclosure illustrated and described in connection with FIGS. 6A - 7C.

[0150] In some embodiments, the processing and rendering involved in generating, updating, and inserting and combining the user graphic representation 120 into the virtual environment are performed by at least one processor 104 of one or more cloud server computers 102 when the client device 118 transmits the unprocessed live data feed 114 of the user 116. Thus, one or more cloud server computers 102 receive the unprocessed live data feed 114 of the user 116 from the client device 118, and then generate, process, and render the user graphic representation 120 placed within the three-dimensional coordinates of the virtual environment 110 from the unprocessed live data feed, and then transmit the cloud-rendered user graphic representation within the virtual environment to other client devices 118. For example, as seen in FIG. 1, client device A can receive the live data feed 114 from its respective camera 112, and then can transmit the unprocessed live data feed 114 of the user to at least one cloud server computer 102. The at least one cloud server computer 102 can generate, process, and render the user graphic representation A, place the user graphic representation A in the three-dimensional coordinates of the virtual environment 118, and then can transmit the user graphic representation A to client device B. A similar process is applied to the user graphic representation B from client device B and user B. Thus, both the user graphic representations A and B can see and interact with each other within the virtual environment A.

[0151] In some embodiments, the virtual environment platform 108 is configured to enable embedding clickable links that redirect to a virtual environment in one or more third-party sources, including third-party websites, applications, or video games. The link can be, for example, an HTML link. The linked virtual environment 110 can be associated with the content of the website in which the link is embedded. For example, the link can be embedded in the website of an automobile dealer or manufacturer, and the clickable link is redirected to a virtual environment 110 representing the showroom of the automobile dealer that the user can visit through the user graphic representation 120.

[0152] In some embodiments, the user graphic representation 120 includes embedded clickable links, such as links that direct to third-party sources containing profile information about the corresponding user. For example, the clickable link can be an HTML link embedded in the source code of the user graphic representation 120 that permits access to a social media (e.g., a professional social media website such as LinkedIn (trademark)) and can provide additional information about the corresponding user. In some embodiments, when the user permits, if another user clicks on or hovers over the corresponding user graphic representation, at least a portion of the basic information of the user is displayed, which can be obtained by accessing user data from a database or from a third-party source.

[0153] In some embodiments, the user graphic representation is a user 3D virtual cutout constructed from a photograph uploaded by the user or provided by a third party (e.g., from a social media website), or a user real-time 3D virtual cutout including a real-time video stream of user 116 with the background removed, or a video with the background removed, or a video with the background not removed. In further embodiments, client device 118 processes and analyzes the live camera feed 114 of user 116 and generates user graphic representation 120 by generating animation data that is transmitted to other peer client devices 118 via a peer-to-peer (P2P) system architecture or a hybrid system architecture. The receiving peer client device 118 uses the animation data to locally construct and update the user graphic representation.

[0154] A user 3D virtual cutout may include a virtual replica of the user constructed from 2D photos uploaded by the user or provided by a third party. In one embodiment, the user 3D virtual cutout is created by a virtual reconstruction process through machine vision technology that uses as input data 2D photos uploaded by the user or provided by a third party to generate a 3D mesh or 3D point cloud of the user with the background removed. In one embodiment, the user 3D virtual cutout may have a static facial expression. In another embodiment, the user 3D virtual cutout may include a facial expression updated through a camera feed. In yet another embodiment, the user 3D virtual cutout may include an expression that can be changed through buttons on a user graphical interface, such as buttons that enable the user 3D virtual cutout to smile, frown, have a neutral face, etc. In yet another embodiment, the user 3D virtual cutout uses a combination of the aforementioned techniques to display a facial expression. After generating the user 3D virtual cutout, the status and / or facial expression of the user 3D virtual cutout can be continuously updated, for example, by processing a camera feed from the user. However, even when the camera is off, the user 3D virtual cutout can still be displayed to other users with an away status and a static facial expression. For example, the user may be concentrating on the current task and not want to be disturbed (e.g., in a "do not disturb" or "busy" status) and may have their camera turned off. At this point, the user 3D virtual cutout can simply be shown sitting at their desk, either stationary or making pre-configured movements such as typing. However, when the user's camera is turned on again, the user 3D virtual cutout can be updated again in real time with respect to the user's facial expression and / or movement. Standard 3D face model reconstruction (e.g., 3D face fitting and texture fusion) techniques may be used for creating the user 3D virtual cutout so that the resulting user graphic representation can be clearly recognized as being the user.

[0155] A user real-time 3D virtual cutout may include a virtual replica of a user after the user's background has been removed, based on real-time 2D or 3D live video stream data feeds obtained from a camera. In one embodiment, the user real-time 3D virtual cutout is created by a virtual reconstruction process through machine vision technology that uses the user's live data feed as input data by generating a 3D mesh or 3D point cloud of the user with the background removed. For example, the user real-time 3D virtual cutout can be generated from 2D video from a camera (e.g., a webcam) that can be processed to create a holographic 3D mesh or 3D point cloud. In another example, the user real-time 3D virtual cutout can be generated from 3D video from a depth camera (e.g., LIDAR or any depth camera) that can be processed to create a holographic 3D mesh or 3D point cloud. Thus, the user real-time 3D virtual cutout graphically represents the user in three dimensions in real time.

[0156] The video with the background removed includes the video streamed to the client device, and the background removal process is performed such that only the user is visible and is displayed using a polygon structure on the receiving client device. The video without the background removed includes the video streamed to the client device, and the video faithfully represents the camera capture, and thus the user and the user's background are visible and are displayed using a polygon structure on the receiving client device. The polygon structure can be a cutout structure or a more complex 3D structure used as a virtual frame for corresponding to the video.

[0157] Videos without the background removed include videos streamed to the client device, which faithfully represent the camera capture, and thus the user and the user's background are visible and are displayed on the receiving client device using a polygon structure. The polygon structure can be a cut structure or a more complex 3D structure used as a virtual frame for the video.

[0158] In some embodiments, the data used as input data included in the live data feed and / or 2D photos uploaded by the user or provided by a third party may include 2D or 3D image data, 3D geometry, video data, media data, audio data, text data, tactile data, time data, 3D entities, 3D dynamic objects, text data, time data, metadata, priority data, security data, location data, lighting data, depth data, and infrared data, among others.

[0159] In some embodiments, the background removal process required to enable user real-time 3D virtual cutout is performed through the use of image segmentation and deep neural networks, which can be made possible through the execution of instructions by one or more processors of the client device 118 or at least one cloud server computer 102. Image segmentation is a process of partitioning a digital image into multiple objects and helps find objects and boundaries that can separate the foreground (e.g., user real-time 3D virtual cutout) obtained from the live data feed 114 of the user 116 from the background. Sample image segmentation that can be used in embodiments of the present disclosure may include, for example, the Watershed transformation algorithm available from OpenCV.

[0160] A suitable process for image segmentation that can be used for background removal in the present disclosure enables such background removal by using artificial intelligence (AI) technologies such as computer vision, and may include instance segmentation and / or semantic segmentation. Instance segmentation assigns individual labels to individual instances of one or more object classes. In some examples, instance segmentation is performed through MASK R-CNN, which, in addition to adding a branch for predicting object masks in parallel with an existing branch for bounding box recognition, detects objects in an image from a user's live data feed 114, etc., and at the same time generates high-quality segmentation masks for each instance. Then, the segmented masks created for the user and the background can be extracted to remove the background. Semantic segmentation uses deep learning or deep neural network (DNN) technologies to enable an automated background removal process. Semantic segmentation partitions an image into semantically meaningful parts by assigning class labels from one or more categories, such as color, texture, and smoothness, to each pixel according to predefined rules. In some examples, semantic segmentation can use a fully convolutional network (FCN) trained end-to-end, pixel-to-pixel for semantic segmentation, as disclosed in the document "Fully Convolutional Networks for Semantic Segmentation", Evan Shelhamer, Jonathan Long,, and Trevor Darrell, IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 39, No. 4 (April 2017), which is incorporated herein by reference.After the foregoing background removal process, a point cloud within the boundaries of the user's face and body remains, and one or more processors of the client device 118 or at least one cloud server computer 102 can process to generate a 3D mesh or 3D point cloud of the user that can be used to construct a user real-time 3D virtual cutout. The user real-time 3D virtual cutout is then updated from the live data feed 114 from the camera 112.

[0161] In some embodiments, updating the user graphic representation 120 includes applying a machine vision algorithm by the client device 118 or at least one cloud server computer 102 to the generated user graphic representation 120 to recognize the facial expression of the user 116 and graphically simulate the facial expression in the user graphic representation 120 within the virtual environment 110. Generally, such recognition of facial expressions can be performed through the principles of affective computing that deal with the recognition, interpretation, processing, and simulation of human emotions. A review of conventional facial expression recognition (FER) techniques is provided in "Facial Expression Recognition Using Computer Vision: A Systematic Review", Daniel Canedo and Antonio J. R. Neves, Applied Sciences, Vol. 9, No. 21 (2019), which is incorporated herein by reference.

[0162] Conventional FER techniques include an image acquisition step, a preprocessing step, a feature extraction step, and a classification or regression step. In some embodiments of the present disclosure, image acquisition is performed by supplying image data from a camera feed 114 to one or more processors. The preprocessing step may be necessary to provide the data most relevant to the feature classifier and typically includes face detection techniques that can create a bounding box that demarcates the face of the target user, which is the desired region of interest (ROI). The ROI is preprocessed, inter alia, through intensity normalization for illumination change, noise filtering for image smoothing, data augmentation to increase training data, rotation correction for rotated faces, image resizing for different ROI sizes, and image cropping for better background filtering. After preprocessing, the algorithm obtains relevant features from the preprocessed ROI, including action units (AUs), the movement of specific face landmarks, the distances between face landmarks, the texture of the face, gradient features, and the like. These features can then be supplied to a classifier, which can be, for example, a support vector machine (SVM) or a convolutional neural network (CNN). After training the classifier, the user's emotion can be detected in real time and constructed in the user graphic representation 120, for example, by concatenating all face feature relationships.

[0163] In some embodiments, the user graphic representation is associated with a top - view viewing perspective, or a third - person viewing perspective, or a first - person viewing perspective, or a self - view viewing perspective. In one embodiment, the viewing perspective of user 116 when accessing the virtual environment through the user graphic representation is a top - view viewing perspective, or a third - person viewing perspective, or a first - person viewing perspective, or a self - view viewing perspective, or a broadcast camera perspective. The self - view viewing perspective may include the user graphic representation as seen by another user graphic representation and optionally the virtual background of the user graphic representation.

[0164] In some embodiments, the viewing perspective is updated when user 116 manually navigates the virtual environment 110 via the graphical user interface.

[0165] In yet another embodiment, the viewing perspective is automatically established and updated using a virtual camera, and the viewing perspective of the live data feed is associated with the viewing perspective of the user graphic representation and the virtual camera, and the virtual camera is automatically updated by tracking and analyzing the tilt data of the user's eyes and head, or the rotation data of the head, or a combination thereof. In one embodiment, the viewing perspective is automatically established and updated using one or more virtual cameras that are virtually placed and aligned in front of the user graphic representation 120, for example, in front of a video with the background removed, or a video with the background not removed, or a user 3D virtual cutout, or a user real-time 3D virtual cutout. In one embodiment, one or more virtual cameras can be directed outward from eye level. In another embodiment, two (one for each eye) virtual cameras can be directed outward from the level of both eyes. In yet another embodiment, one or more virtual cameras can be directed outward from the center of the position of the head of the user graphic representation. The viewing perspective of the user 116 captured by the camera 112 is associated with the viewing perspective of the user graphic representation 120 and the associated virtual camera that uses computer vision to operate the virtual camera.

[0166] The virtual camera provides a virtual representation of the viewing perspective of the user graphic representation 120 associated with the viewing perspective of the user 116, enabling the user 116 to view the area of the virtual environment 110 that the user graphic representation 120 is looking at in one of many viewing perspectives. The virtual camera is automatically updated by tracking and analyzing the tilt data of the user's eyes and head, or the rotation data of the head, or a combination thereof. The position of the virtual camera can also be manually changed by the user 116 according to the viewing perspective selected by the user 116.

[0167] The self-viewing perspective is a viewing perspective of the user graphic representation 120 as seen by another user graphic representation 120 with the background removed (e.g., like the "selfie mode" of a phone camera). The self-viewing perspective can alternatively include a virtual background of the user graphic representation 120 to understand the perception of the user 116 as seen by other participants. When the self-viewing perspective includes a virtual background of the user graphic representation, it can be set as the area around the user graphic representation that can be captured by a virtual camera, and as a result, it can be circular, square, rectangular, or any other shape suitable for framing the self-viewing perspective. For example, in a scenario where the user graphic representation 120 is virtually placed in a house with a window where trees can be seen behind the user, the self-viewing perspective displays the user graphic representation and, alternatively, the background including the window and the trees.

[0168] In yet another embodiment, tracking and analyzing the tilt data of the user's eyes and head, or the rotation data of the head, or a combination thereof, includes capturing and analyzing the view position and orientation captured by at least one camera 112 using computer vision, and thus operating a virtual camera in the virtual environment 110. For example, such an operation can include receiving and processing the tilt data of the eyes and head captured by at least one camera through a computer vision method, extracting the view position and orientation from the tilt data of the eyes and head, identifying one or more coordinates of the virtual environment included in the position and orientation from the tilt data of the eyes, and operating the virtual camera based on the identified coordinates.

[0169] In some embodiments, instructions within the memory 106 of at least one cloud server computer 102 further enable performing data analysis of user activity within at least one virtual environment 110. The data analysis can be used for interactions including participation in conversations with other users, interaction with objects within the virtual environment 110, purchases, downloads, engagement with content, etc. The data analysis can use multiple known machine learning techniques to collect and analyze data from the interactions in order to perform recommendations, optimizations, predictions, and automations. For example, the data analysis can be used for marketing purposes.

[0170] In some embodiments, at least one processor 104 of one or more cloud server computers 102 is further configured to enable transactions and monetization of content added to at least one virtual environment 110. At least one cloud server computer 102 can be communicatively connected to an application and object library that enables a user to find, select, and insert content in at least one virtual environment through an appropriate application programming interface (API). One or more cloud server computers 102 can further be connected to one or more payment gateways that enable execution of corresponding transactions. The content can include, for example, interactive applications or static or interactive 3D assets, animations, or 2D textures, etc.

[0171] Figures 2A - 2B are schematic diagrams of the deployment 200a and 200b of a system that enables interactions in a virtual environment including multiple verticals of a virtual environment platform.

[0172] Figure 2A is a schematic diagram of a deployment 200a of a system that enables interaction in a virtual environment including a plurality of verticals 202 of a virtual environment platform 108 according to one embodiment. Some elements of FIG. 2A may refer to the same or similar elements of FIG. 1, and thus, the same reference numbers may be used.

[0173] The verticals 202 are associated with context themes of the virtual environment, including, for example, a meeting 204 as a virtual conference room, work 206 as a virtual office space, learning 208 as a virtual classroom, and shopping 210 as a virtual store. Other verticals not shown in FIG. 2A include, for example, services such as banking, reservations (e.g., hotels, tour agencies, or restaurants), and government agency services (e.g., inquiries for setting up a new company for a fee), and entertainment (e.g., karaoke, event halls or arenas, movie theaters, nightclubs, stadiums, art galleries, cruise ships, etc.).

[0174] Each of the virtual environment verticals 202 may include a plurality of available virtual environments 110 (e.g., virtual environments A - L), each having one or more available arrangements and appearances associated with the context of the corresponding vertical 202. For example, the virtual environment A of the vertical 202 for the meeting 204 may include a seated meeting desk, a whiteboard, and a projector. Each of the virtual environments 110 may be provided with corresponding resources (e.g., memory, network, and computing power) by at least one cloud server computer. The vertical 202 can be utilized from the virtual environment platform 108, and one or more users 116 can access the virtual environment platform 108 through the graphical user interface 212 via the client device 118. The graphical user interface 212 is included in a downloadable client application or a web browser application and provides the application data and instructions required to execute the selected virtual environment 110, enabling multiple interactions within the selected virtual environment 110. Further, each of the virtual environments 110 may include one or more human or artificial intelligence (AI) hosts or assistants that can assist users within the virtual environment by providing the required data and / or services through the corresponding user graphic representation. For example, a human or AI bank service representative can assist users of a virtual bank by providing the required information in the form of presentations, forms, lists, etc., in response to user requests.

[0175] In some embodiments, each virtual environment 110 is a persistent virtual environment that records changes made therein, including customizations, and the changes are stored in the persistent memory storage of at least one cloud server computer 102. For example, returning to the example of virtual environment A, the arrangement of seats around the desk, the color of the walls, or even the size and occupancy capacity of the room can be changed according to the needs or preferences of the user. The changes made are saved in the persistent memory and are then available during subsequent sessions in the same virtual environment A. In some examples, enabling persistent storage of changes in the virtual environment 110 may require payment of a subscription fee to the host or owner of the room (e.g., through the virtual environment platform 108 connected to a payment gateway).

[0176] In other embodiments, the virtual environment 110 is a temporary virtual environment stored in the temporary memory storage of at least one cloud server computer 102. In these embodiments, changes made in the virtual environment 110 may not be stored and, thus, may not be available in future sessions. For example, a temporary virtual environment may be selected from the virtual environment platform 108 from pre-defined available virtual environments from different verticals 202. Changes such as decoration or arrangement changes may or may not be possible, but if changes are possible, the changes may be lost after the session ends.

[0177] In some embodiments, the complex of virtual environments 110 within one or more verticals 202 can represent a virtual environment cluster 214. For example, some virtual environment clusters 214 can include hundreds or thousands of virtual environments 110. To the user, the virtual environment cluster 214 may appear to be part of the same system, where the user can interact with each other or seamlessly access other virtual environments within the same virtual environment cluster 214. For example, virtual environments D and E from the work virtual environment vertical 206 and virtual environment B from the conference virtual environment vertical 204 can form a virtual environment cluster 214 representing a company. The user in this example can have two different work areas, such as a game development room and a business development room, in addition to a conference room for video conferencing. A user from either the game development room or the business development room can hold a meeting in the conference room and open a private virtual meeting, and the remaining staff can continue to perform their current activities in the original work areas.

[0178] In other examples, the virtual environment cluster 214 can represent a movie theater or an event facility, and each virtual environment 110 represents an indoor or outdoor event area (e.g., a theater or an event arena) where a live performance is performed by one or more performers through corresponding user graphic representations. For example, an orchestra and / or a singer can hold a music concert through a live recording of the performance by a camera and through user graphic representations, such as through a user live 3D virtual cutout. The user graphic representation of each performance can be inserted into the corresponding three-dimensional coordinates of the stage from which the performance can be given. The audience can watch the performance from the theater through the corresponding user graphic representation and perform a plurality of interactions such as virtually clapping, singing together, dancing virtually, jumping virtually, or cheering.

[0179] In other examples, the virtual environment cluster 214 can represent a casino that includes a plurality of gaming areas (e.g., blackjack area, poker area, roulette area, and slot machine area), a token purchase area, and an event room. The machines in each gaming area can be configured as casino applications configured to provide a user experience associated with each game. The casino operator can have corresponding user graphic representations 120 or user real-time 3D virtual cutouts. The casino operator represented by s can be an actual human operator or an artificial intelligence program that assists users of the virtual casino. Each casino game is coupled to a payment gateway from the casino company operating the virtual casino to enable payments with users.

[0180] In other examples, the virtual environment cluster 214 represents a shopping mall with multiple floors, and each floor includes a plurality of virtual environments such as stores, showrooms, common areas, and food courts. Each virtual room can be managed by a corresponding virtual room manager. For example, each store can be managed by a corresponding store manager. Sales clerks can be utilized in each area as 3D live virtual avatars or user real-time 3D virtual cutouts and can be actual humans or AI assistants. In the current example, each virtual store and the restaurant in the sample food court can be configured to enable online purchase of goods and delivery to the user's address through corresponding payment gateways and delivery systems.

[0181] In another example, the virtual environment cluster 214 includes a plurality of virtual party areas of a virtual nightclub where users can meet and communicate through corresponding user graphic representations. For example, each virtual party area can have a different theme and associated music and / or decorations. In addition to conversations and text message sending, some other interactions in the virtual nightclub can include, for example, virtually dancing or drinking, sitting in various seating areas (such as a lounge or a bar), etc. Further, in this example, an indoor music concert can be held in the virtual nightclub. For example, a disc jockey (DJ) playing behind a virtual table on the stage can play an electronic music concert. In this case, the DJ can be represented by a 3D live virtual avatar or a user real-time 3D virtual cutout. When the DJ is represented by a user real-time 3D virtual cutout, the real-time movement of the DJ playing the audio mixing console can be projected in real-time 3D virtual cutout from live data feeds obtained by a camera capturing images from the location of the DJ (such as from the DJ's home or recording studio). Further, each member of the audience can also be represented by their own user graphic representation. In this case, some users can be represented by 3D live virtual avatars, and other users can be represented by user real-time 3D virtual cutouts according to the user's preference.

[0182] In other examples, the virtual environment cluster 214 can represent a virtual karaoke entertainment facility with multiple private or public karaoke rooms. Each private karaoke room can be equipped with a virtual karaoke machine, a virtual screen, a stage, microphones, speakers, decorations, a couch, a table, and drinks and / or food. By selecting a song through a virtual karaoke machine that the user can connect to the song database, the virtual karaoke machine can play the song for the user and trigger the system to project the lyrics onto the virtual screen for the user to sing through the user graphic representation. The public karaoke room can further have a human or AI DJ who selects songs for the user, calls the user onto the stage, and mutes or unmutes the user as needed to listen to the performance. The user can sing remotely from a client device through a microphone.

[0183] In other examples, the virtual environment cluster 214 can represent a virtual cruise ship with multiple areas such as a bedroom, an engine room, an event room, a bow, a stern, a port side, a starboard side, a bridge, and multiple decks. Some of the areas can have a human or AI assistant who attends to the user through the corresponding user graphic representation, such as providing additional information or services. When available, the virtual environment outside the cruise ship or a simple graphic representation required to depict the scenery of an island, town, or city that can be visited when arriving at a specific destination can be utilized. Thus, the user can experience an ocean voyage through the user graphic representation, discover new places, and interact with each other virtually at the same time.

[0184] In other examples, the virtual environment cluster 214 can represent an e-sports stadium or a gymnasium that includes a plurality of virtual environments representing arenas, courts, or rooms where users can play through user graphic representations via appropriate input / output devices (e.g., a computer keyboard, a game controller, etc.). The mechanics of each e-sport can vary depending on the sport being played. The e-sports stadium or gymnasium can include a shared area where users can select the sports areas they access. A sports schedule that notifies users of which sports activities are available at what times may also be available.

[0185] Figure 2B shows the deployment 200b of a virtual school 216 that combines multiple virtual environments from various verticals 202. The virtual school 216 includes four classrooms (e.g., Classrooms A - D 218 - 224), a lecture hall 226, a sports area 228, a cafeteria 230, a faculty room 232, a library 234, and a bookstore 236. Each virtual environment can include virtual objects represented by corresponding graphic representations associated with the corresponding environment.

[0186] For example, a virtual classroom (e.g., any one of virtual classrooms A to D 218 to 224) enables students to attend classes and participate in the class through various interactions (e.g., raising hands, content projection, presentation, asking questions or posting through speech or text, etc.), and provides special administrative privileges to teachers (e.g., giving someone the right to speak, muting one or more students during class, sharing content through a digital whiteboard, etc.). The lecture hall enables a speaker to give a speech or hold multiple events. The sports area 228 can be configured to enable students to play multiple e-sports through corresponding user graphic representations. The cafeteria 230 can enable students to order food online and socialize through user graphic representations. The staff room 232 can be configured for teachers to meet, discuss topics, and give progress reports on students, etc. through corresponding teacher user graphic representations. The library 234 can enable students to borrow e-books for learning assignments or leisure reading. Finally, the bookstore 236 can be configured to enable students to purchase books (e.g., e-books or physical books) and / or other school teaching materials.

[0187] FIG. 3 is a schematic diagram of a sample hybrid system architecture 300 that can be employed in a system that enables interaction in a virtual environment. The hybrid system architecture 300, in some embodiments, is a hybrid model of communication for interacting with other peer clients (e.g., other attendees in a virtual conference, classroom, etc.), comprising a client-server side 304 and a P2P side 306, each delimited by a dotted region in FIG. 3. By using such a hybrid model of communication, rapid P2P communication between users becomes possible, while reducing latency issues while providing web services, data, and resources for each session, enabling multiple interactions between users and with content in the virtual environment. Some elements in FIG. 3 may refer to the same or similar elements in FIGS. 1-2A, and thus the same reference numbers may be used.

[0188] In various embodiments, the usage level and ratio of the client - server side 304 to the P2P side 306 depend on the amount of data to be processed, the latency allowed to maintain a smooth user experience, the desired quality of service (QOS), the required services, etc. In one embodiment, the P2P side 306 is used for video and data processing, streaming, and rendering. This mode of adopting the hybrid system architecture 300 may be suitable, for example, when there is a need to process a small amount of data with low latency and when there are "heavy" clients, meaning that the client devices have sufficient computing power to perform such operations. In another embodiment, a combination of the client - server side 304 and the P2P side 306 is adopted, and such a P2P side 306 is used for video streaming and rendering, while the client - server side 304 is used for data processing. This mode of adopting the hybrid system architecture 300 may be suitable, for example, when there is a large amount of data to be processed or when other microservices may be required. In yet another embodiment, the client - server side 304 is used for video streaming as well as data processing, while the P2P side 306 is used for video rendering. This mode of adopting the hybrid system architecture 300 may be suitable, for example, when the amount of data to be processed is even larger and / or when only sink clients are available. In yet another embodiment, the client - server side 304 is used for video streaming, rendering, and data processing. This mode of adopting the hybrid system architecture 300 may be suitable when berry sink clients are available. The hybrid system architecture 300 can be configured to be able to switch between different usage modalities of both the client - server side 304 and the P2P side 306 within the same session, if necessary.

[0189] In some embodiments, at least one cloud server from the client-server side 304 is an intermediate server, which means that the server is used to facilitate and / or optimize the exchange of data between client devices. In such embodiments, at least one cloud server manages, analyzes, processes, and optimizes incoming images and multimedia streams, and manages, evaluates, and / or optimizes the transfer of outgoing streams as a router topology (e.g., but not limited to, SFU (Selective Forwarding Units), SAMS (Spatially Analyzed Media Server), multimedia router, etc.), or an image and media processing server topology (e.g., for tasks including but not limited to decoding, combining, improving, mixing, enhancing, expanding, computing, operating, encoding), or a transfer server topology (including but not limited to MCU, cloud media mixer, cloud 3D renderer, media server), or other server topologies.

[0190] In such embodiments where the intermediate server is a SAMS, such a media server manages, analyzes, and processes the incoming data (e.g., metadata, priority data, data class, spatial structure data, three-dimensional position, orientation, or movement information, image, media, scalable video codec-based video, or a combination thereof) of each sending client device, and in such analysis, manages and / or optimizes the transfer of the outgoing data stream to each receiving client device. This includes changing, upscaling, or downscaling the media in terms of time (e.g., various frame rates), space (e.g., different image sizes), quality (e.g., quality based on different compression or encoding), and color (e.g., color resolution and range), and may be based on factors such as the priority relationship to such incoming data to achieve the optimal bandwidth and computing resource utilization for receiving the spatial three-dimensional orientation, distance, and one or more user client devices of a specific receiving client device user.

[0191] In some embodiments, media, video, and / or data processing tasks include one or more of encoding, transcoding, decoding space or 3D analysis and processing including, for example, image filtering, computer vision processing, image sharpening, background improvement, background removal, foreground blurring, eye covering, face pixelation, voice distortion, image upscaling, image cleansing, skeleton analysis, face or head counting, object recognition, marker or QR code tracking, visual target tracking, feature analysis, 3D mesh or volume generation, feature tracking, face recognition, SLAM tracking, and face expression recognition, or other modular plugins in the form of microservices executed on such a media router or server.

[0192] The client-server side 304 uses a secure communication protocol 308 to enable secure end-to-end communication between the client device 118 and the web / application server 310 via the network. Examples of suitable secure communication protocols 308 include, for example, Datagram Transport Layer Security (DTLS), which is itself a secure UDP (user datagram protocol), Secure Realtime Transport Protocol (SRTP), Hypertext Transfer Protocol Secure (https: / / ), and WebSocket Secure (wss: / / ), which are compatible and can provide fully authenticated application access, privacy protection, and integrity of the exchanged data during transfer. Examples of suitable web / application servers 310 include, for example, the Jetty web application server, which is a Java HTTP web server and Java Servlet container and enables proper deployment of machine-to-machine communication and web application services.

[0193] Although the web / application server 310 is shown as a single element in FIG. 3, those skilled in the art will understand that the web server and the application server can be separate elements. For example, the web server can be configured to receive client requests through a secure communication protocol 308 and route the requests to the application server. Thus, the web / application server 310 can receive client requests using the secure communication protocol 308 and process the requests, which can include requesting one or more microservices 312 (e.g., Java-based microservices) and / or retrieving data from the database 314 using the corresponding database management system 316. The application / web server 310 can provide session management, a number of other services such as 3D content and application logic, and session state persistence (e.g., for persistent storage of shared documents, interaction and synchronization of changes in a virtual environment, or persistence of the visual state and modifications of a virtual environment). A suitable database management system 316 can be, for example, an object-relational mapping (ORM) database management system, which is suitable for database management using open-source commercial (e.g., proprietary) services provided with ORM capabilities to convert data between incompatible types of systems using an object-oriented programming language. In further embodiments, a distributed spatial data bus 318 can be further used as a distributed message and resource delivery platform between microservices and client devices by using a publish-subscribe model.

[0194] The P2P side 306 enables real-time communication between peer client devices 118 in a virtual environment through an appropriate application programming interface (API) by using an appropriate P2P communication protocol 320, thereby enabling their real-time interaction and synchronization and resulting in a multi-user collaboration environment. For example, through the P2P side 306, the contributions of one or more users can be directly sent to other users, and other users can observe the executed changes in real time. An example of an appropriate P2P communication protocol 320 is the WebRTC (Web Real-Time Communication) communication protocol, which is a collection of standards, protocols, and JavaScript APIs that, when combined, enable P2P voice, video, and data sharing between peer client devices 118. The client devices 118 on the P2P side 306 can use one or more rendering engines 322 to perform real-time 3D rendering of live sessions. An example of an appropriate rendering engine 322 is a 3D engine based on WebGL, a JavaScript API for rendering 2D and 3D graphics within any compatible web browser without using a plugin, which accelerates the use of physics and image processing and effects by one or more processors (e.g., one or more graphics processing units (GPUs)) of the client device 118. Further, the client devices 118 on the P2P side 306 can perform image and video processing and machine learning computer vision techniques through one or more appropriate computer vision libraries 324. In one embodiment, the image and video processing performed by the client devices on the P2P side 306 includes a background removal process used to create user graphic representations before inserting the user graphic representations into the virtual environment, which can be performed either in real time or near real time on the received media stream or, for example, non-real time on a photograph.An example of a suitable computer vision library 324 can be OpenCV, which is a library of programming functions mainly configured for real-time computer vision tasks.

[0195] FIG. 4 is a schematic diagram of a graphical user interface 400 of a virtual environment live session module 402 according to an embodiment, in which a user can interact in a virtual environment.

[0196] Before the user can access the graphical user interface 400 of the virtual environment live session module 402, the user first receives an invitation from a peer client device to participate in a conversation with the peer user, whereby a P2P communication channel is opened between the user client devices when processing and rendering are performed by the client devices, or alternatively, an indirect communication channel is opened through a cloud server computer when processing and rendering are performed by at least one cloud server computer. Further, as shown in the latter part of the description with reference to FIG. 5, a transition can be made from the user 3D virtual cutout to the user real-time 3D virtual cutout, or a video with the background removed, or a video with the background not removed.

[0197] The virtual environment live session module 402 can include a virtual environment screen 404 that includes a graphical user interface showing the selected virtual environment, which can include the configuration of the virtual environment associated with the selected vertical context of the virtual environment, and corresponding virtual objects, applications, other user graphic representations, and the like. The graphical user interface 400 of the virtual environment live session module 402 can enable and display, for example, a plurality of interactions 406 configured such that users engage with each other through a user real-time 3D virtual cutout. The virtual environment live session module 402 can include one or more data models associated with corresponding tasks enabling each interaction 406, and computer instructions necessary to perform the tasks. Each interaction 406 can be represented in various ways, and in the example shown in FIG. 4, each individual interaction 406 is represented as a button on the graphical user interface 400 from the virtual environment live session module 402, and by clicking on each interaction button, a corresponding service can be requested to perform the task associated with the interaction 406. The virtual environment live session module 402 is enabled, for example, through the hybrid system architecture 300 disclosed with reference to FIG. 3.

[0198] Interactions 406 can include, for example, chat 408, screen sharing 410, host options 412, remote sensing 414, recording 416, voting 418, document sharing 420, emoji sending 422, sharing and editing of topics 424, or other interactions 426. Other interactions 426 can include, for example, virtual hugs, raising hands, handshakes, walking, adding content, preparing a summary of a meeting, moving an object, projection, laser pointing, game play, purchasing, and other social interactions that facilitate exchange, competition, cooperation, and conflict resolution among users. The various interactions 406 are described in more detail below.

[0199] Chat 408 can open a chat window that enables sending and receiving text comments and on-the-fly resources.

[0200] Screen sharing 410 may enable sharing the user's screen with other participants in real time.

[0201] Host options 412 are configured to provide additional options to the conversation host, such as muting one or more users, inviting or removing one or more users, and ending the conversation.

[0202] Remote sensing 414 can display the current status of a user, such as away, busy, available, offline, in a phone call, or in a meeting. The user status can be updated manually through a graphical user interface or automatically through a machine vision algorithm based on a data feed obtained from a camera.

[0203] Recording 416 enables recording audio and / or video from a conversation.

[0204] Voting 418 allows voting on one or more proposals submitted by other participants. Through voting 418, a host or other participant who has obtained such permission can start a voting session at any time. The subject and options can be displayed for each participant. Depending on the configuration of the interaction of voting 418, the results can be shown to all participants at the end of a timeout period or at the end of all responses.

[0205] Document sharing 420 enables sharing documents with other participants in any suitable format. These documents can also be persistently retained by storing them in the persistent memory of one or more cloud server computers and can be associated with the virtual environment in which the virtual communication takes place.

[0206] Emoji sending 422 enables sending emojis to other participants.

[0207] Agenda sharing and editing 424 enables sharing and editing of agendas prepared by any of the participants. In some embodiments, a checklist of agenda items can be set by the host prior to the meeting. The agenda can be brought forward at any time by the host or other participants who have obtained such permission. Through the agenda editing option, items can be unchecked or postponed when an agreement is reached.

[0208] Other interactions 426 provides a non-exhaustive list of possible interactions that can be provided in the virtual environment depending on the virtual environment vertical. Raising a hand enables raising a hand during virtual communication or meetings, whereby the host or other participants who have obtained such qualification enable the user to speak. Walking enables moving within the virtual environment through the user's real-time 3D virtual cutout. Content addition enables the user to add interactive applications or static or interactive 3D assets, animations, or 2D textures to the virtual environment. Preparation of a meeting summary enables automatic preparation of the results of the virtual meeting and distribution of such results to the participants at the end of the session. Object movement enables moving objects within the virtual environment. Projection enables projecting content from the participants' screens onto a screen or wall available in the virtual environment. Laser pointing enables pointing a laser to highlight desired content on a presentation. Gameplay enables playing one or more games or other types of applications that can be shared during a live session. Purchase enables purchases during a content session. Other interactions not described herein can also be configured according to the specific use of the virtual environment platform.

[0209] In some embodiments, the system further enables the creation of ad hoc virtual communications, which may include creating an ad hoc voice communication channel between user graphic representations without the need to change the current viewing perspective or position in the virtual environment. For example, a user graphic representation can approach another user graphic representation and conduct an ad hoc voice conversation at a location within the virtual environment where both user graphic representation areas exist. Such communication can be enabled, for example, by taking into account the distance, position, and orientation between user graphic representations, and / or their current availability status (e.g., available or unavailable), or the status configuration of such ad hoc communication, or a combination thereof. The approaching user graphic representations, in this example, signal that ad hoc communication is possible, thus setting the start of a conversation between both user graphic representations, and the other user graphic representation will see visual feedback regarding the approaching user. In this case, the approaching user can speak, and the other user can listen and respond. In another example, a user graphic representation approaches another user graphic representation, clicks on the user graphic representation to send a conversation invitation, and after approval by the inviter, an ad hoc voice conversation can be conducted at a location within the virtual environment where both user graphic representation areas exist. The other user can see the interaction, expression, hand movements, etc. between the user graphic representations, regardless of whether they can hear the conversation, according to the privacy settings between the two user graphic representations. Any of the aforementioned interactions 406 or other interactions 426 can also be performed directly within the virtual environment screen 404.

[0210] Figure 5 shows a method 500 according to an embodiment that enables transitioning from one type of user graphic representation to another type of user graphic representation, for example, from a user 3D virtual cutout to a user real-time 3D virtual cutout, or to a video with a removed background, or to a video without a removed background.

[0211] Transitions are enabled when the user is conversing with another user graphic representation. For example, the user may currently be sitting in an office chair and working on a computer in a virtual office. The user's current graphic representation may be a representation of a user 3D virtual cutout. At that point, the camera may not be on because live data feeds from the user may not be required. However, if the user decides to turn the camera on, the user 3D virtual cutout may include facial expressions provided through the user's face analysis captured from the user's live data feed, as described in more detail herein.

[0212] When the user starts a live session by conversing with another user graphic representation and the user's camera is not activated, the camera can be activated, the capture of a live data feed that can provide the user's live stream can be started, and the user 3D virtual cutout can be transitioned to a user real-time 3D virtual cutout, or a video with the background removed, or a video with the background not removed. Further, as described in FIG. 1, the live stream of the user real-time 3D virtual cutout 504 can be processed and rendered on the client or server, or transmitted to other peer client devices in a P2P system architecture or a hybrid system architecture for real-time independent processing and rendering (e.g., through the hybrid system architecture 300 described with reference to FIG. 3).

[0213] Method 500 of FIG. 5 starts in step 502 by approaching a user graphic representation. Next, in step 504, method 500 selects and clicks on the user graphic representation. Subsequently, in step 506, method 500 sends or receives an invitation to participate in a conversation with another user graphic representation through a client device. Subsequently, in step 508, method 500 accepts the received invitation by the corresponding client device. Next, method 500 migrates in step 510 the user 3D virtual cutout to a user real-time 3D virtual cutout, or a video with the background removed, or a video with the background not removed. Finally, in step 512, method 500 ends by opening a P2P communication channel between user client devices when processing and rendering are performed by a client device, or by opening an indirect communication channel through a cloud server computer when processing and rendering are performed by at least one cloud server computer. In some embodiments, the conversation includes sending and receiving real-time audio and video displayed from the participant's user real-time 3D virtual cutout.

[0214] 6A-6C are schematic diagrams of a combination of a plurality of image processes performed on the client-server side 304 by the corresponding client device 118 and cloud server 102. The client-server side can be part of a hybrid system architecture, such as the hybrid system architecture 300 as shown in FIG. 3, for example.

[0215] In one embodiment of FIGS. 6A-6C, at least one cloud server 102 can be configured as a Traversal Using Relay Network Address Translation (NAT) (sometimes called TURN) server, which may be suitable for situations where the server cannot establish a connection between client devices 118. TURN is an extension of the Session Traversal Utilities of NAT (STUN).

[0216] NAT is a method of remapping an Internet Protocol (IP) address space to another address space by changing the network address information within the IP header of a packet while the packet is passing through a traffic routing device. Thus, NAT can provide private IP addresses for accessing networks such as the Internet, enabling a single device such as a routing device to function as an agent between the Internet and a private network. NAT can be symmetric or asymmetric. A framework called Interactive Connectivity Establishment (ICE), which is configured to find the optimal path for connecting client devices, can determine whether a symmetric or asymmetric NAT is required. Symmetric NAT performs the job of not only converting IP addresses from private to public or vice versa but also converting ports. On the other hand, asymmetric NAT enables the use of a STUN server to detect the public IP address and the type of the underlying NAT that a client can use for connection establishment. In many cases, STUN can only be used during the setup of a connection, and once the session is established, the data flow can be started between client devices.

[0217] TURN can be used in the case of symmetric NAT and can remain in the media path even after a connection is established while processed and / or unprocessed data is relayed between client devices.

[0218] FIG. 6A shows a client-server side 304 including a client device A, a cloud server 102, and a client device B. In FIG. 6A, the client device A is the transmission side of the data to be processed, and the client device B is the reception side of the data. A plurality of image processing tasks are drawn and classified based on which of the client device A, the cloud server 102, and / or the client device B performs them, and thus are classified as client device A processing 602, server image processing 604, and client device B processing 606.

[0219] The image processing tasks include background removal 608, further processing or improvement 610, and insertion and combination into a virtual environment 612. As will be apparent from FIGS. 6B and 6C and also from FIG. 7B, the combination of the three image processing tasks shown in this specification can be used for the generation, improvement, and insertion / combination into a virtual environment of a user graphic representation. Further, for simplicity, in FIGS. 6B-6C and FIGS. 7B-7C, the background removal 608 is shown as "BG" 608, the further processing or improvement 610 is shown as "++" 610, and the insertion and combination into a virtual environment 612 is shown as "3D" 612.

[0220] In some embodiments, inserting and combining user graphic representations into a virtual environment includes generating one or more virtual cameras that are virtually placed and aligned in front of the user graphic representation, for example, in front of a video with the background removed, or a video with the background not removed, or a user 3D virtual cutout, or a user real-time 3D virtual cutout. In one embodiment, the one or more virtual cameras can be directed outward from eye level. In another embodiment, two (one for each eye) virtual cameras can be directed outward from both eye levels. In yet another embodiment, the one or more virtual cameras can be directed outward from the center of the position of the head of the user graphic representation. In yet another embodiment, the one or more virtual cameras can be directed outward from the center of the user graphic representation. In yet another embodiment, when in a self-viewing perspective, the one or more virtual cameras may be placed in front of the user graphic representation, for example, at the height of the head of the user graphic representation, and directed towards the user graphic representation. The one or more virtual cameras are created by associating the captured user viewing perspective data with the viewing perspective of the user graphic representation within the virtual environment, at least using computer vision. The one or more virtual cameras are automatically updated by tracking and analyzing the tilt data of the user's eyes and head, or the rotation data of the head, or a combination thereof, and can also be manually changed by the user according to the viewing perspective selected by the user.

[0221] The combination of image processing and the corresponding usage levels of client device A processing 602, server image processing 604, and client device B processing 606 depend on the amount of data to be processed, the latency allowed to maintain a smooth user experience, the desired quality of service (QoS), the required services, and the like.

[0222] FIG. 6B shows combinations 1-4 of image processing.

[0223] In image processing combination 1, client device A generates a user graphic representation including background deletion 608 and transmits the user graphic representation with the background deleted to at least one cloud server 102 for further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted. The at least one cloud server transmits the enhanced user graphic representation with the background deleted to client device B, and client device B inserts and combines the enhanced user graphic representation with the background deleted into a virtual environment.

[0224] In image processing combination 2, client device A generates a user graphic representation including background deletion 608, performs further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted, and transmits it to at least one cloud server 102. The at least one cloud server 102 transmits the enhanced user graphic representation with the background deleted to client device B, and client device B inserts and combines the enhanced user graphic representation with the background deleted into a virtual environment.

[0225] In image processing combination 3, client device A generates a user graphic representation including background deletion 608, performs further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted, inserts and combines the enhanced user graphic representation with the background deleted into a virtual environment. Then, client device A transmits the enhanced user graphic representation with the background deleted inserted and combined into the virtual environment to a cloud server for relaying to client device B.

[0226] In image processing combination 4, client device A generates a user graphic representation including background deletion 608 and transmits the user graphic representation with the background deleted to at least one cloud server 102 for further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted. Next, at least one cloud server inserts and combines the enhanced user graphic representation with the background deleted into a virtual environment and transmits it to client device B.

[0227] Figure 6C shows image processing combinations 5 - 8.

[0228] In image processing combination 5, client device A generates a user graphic representation including background deletion 608 and transmits the user graphic representation with the background deleted to at least one cloud server 102 to relay it to client device B. Client device B performs further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted for the user graphic representation with the background deleted, and inserts and combines the enhanced user graphic representation with the background deleted into a virtual environment.

[0229] In image processing combination 6, client device A transmits a camera live data feed received from at least one camera, transmits the unprocessed data to at least one cloud server 102, and at least one cloud server 102 generates a user graphic representation including background deletion 608, performs further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted for the user graphic representation with the background deleted, and transmits the enhanced user graphic representation with the background deleted to client device B. Client device B inserts and combines the enhanced user graphic representation with the background deleted into a virtual environment.

[0230] In image processing combination 7, the client device transmits a camera live data feed received from at least one camera and transmits the unprocessed data to at least one cloud server 102. The at least one cloud server 102 generates a user graphic representation including background deletion 608, and performs further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted for the user graphic representation with the background deleted. Then, the enhanced user graphic representation with the background deleted is inserted into and combined with the virtual environment and transmitted to client device B.

[0231] In image processing combination 8, client device A transmits a camera live data feed received from at least one camera and transmits the unprocessed data to at least one cloud server 102 for relaying to client device B. Client device B uses the data to generate a user graphic representation including background deletion 608, and performs further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted for the user graphic representation with the background deleted. Then, the enhanced user graphic representation with the background deleted is inserted into and combined with the virtual environment. As will be appreciated, in some embodiments, the at least one cloud server 102 is an intermediate server, which means that the server uses an intermediate server topology to facilitate and / or optimize the exchange of data between client devices.

[0232] In such an embodiment, at least one cloud server is an intermediate server, which means that the server is used to facilitate and / or optimize the exchange of data between client devices. In such an embodiment, at least one cloud server can manage, analyze, and optimize an incoming multimedia stream and manage, evaluate, and optimize the transfer of an outgoing stream as a router topology (e.g., SFU, SAMS, multimedia server router, etc.), or media processing (e.g., performing tasks including decoding, combining, improving, mixing, enhancing, extending, computing, operating, or encoding), and a transfer server topology (e.g., but not limited to, multipoint control unit, cloud media mixer, cloud 3D renderer), or other server topologies.

[0233] In such an embodiment where the intermediate server is a SAMS, such a media server manages, analyzes, and processes the incoming data (e.g., metadata, priority data, data class, spatial structure data, three-dimensional position, orientation, or movement information, image, media, or scalable video codec-based video) of the sending client device, and in such analysis, manages or optimizes the transfer of the outgoing data stream to the receiving client device. This may include changing, upscaling, or downscaling the media with respect to time (e.g., various frame rates), space (e.g., different image sizes), quality (e.g., quality based on different compression or encoding), and color (e.g., color resolution and range) based on one or more factors such as the priority relationship for such incoming data to achieve the optimal bandwidth and computing resource utilization for the spatial three-dimensional orientation, distance, and one or more user client devices of a particular receiving client device user.

[0234] The intermediate server topology may be suitable for combinations 1-8 of image processing in which at least one cloud server 102 processes between client devices A and B, for example, as shown in FIGS. 6A-6C.

[0235] FIGS. 7A-7C are schematic diagrams of a plurality of combinations of image processing performed on the P2P side 306 by corresponding client devices shown in FIGS. 7A-7B as peer devices A-B to distinguish the case where communication and processing are performed through the client server side. The P2P side 306 can be part of a hybrid system architecture such as the hybrid system architecture 300 shown in FIG. 3, for example.

[0236] FIG. 7A shows the P2P side 306 comprising peer device A and peer device B, where peer device A is the transmitting side of the data to be processed and peer device B is the receiving side of the data. A plurality of image and media processing tasks are drawn and classified based on which of peer device A or peer device B performs them, and are thus classified as peer device A processing 702 and peer device B processing 704. The image and media processing tasks can include, but are not limited to, background removal 608, further processing or improvement 610, and insertion and combination 612 into a virtual environment.

[0237] FIG. 7B shows combinations 1-3 of image processing.

[0238] In combination 1 of image processing, peer device A generates a user graphic representation including background removal 608, performs further processing or improvement 610 on the user graphic representation with the background removed, generates an enhanced user graphic representation with the background removed, and inserts and combines the enhanced user graphic representation with the background removed into a virtual environment having three-dimensional coordinates. Then, peer device A transmits the enhanced user graphic representation with the background removed, inserted and combined into the virtual environment, to peer device B.

[0239] In image processing combination 2, peer device A generates a user graphic representation including background deletion 608, and transmits the user graphic representation with the background deleted to peer device B. Peer device B performs further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted for the user graphic representation with the background deleted, and inserts and combines it into the virtual environment.

[0240] In image processing combination 3, peer device A transmits a camera live data feed received from at least one camera, and transmits the encoded data to peer device B. Peer device B decodes the data, uses the data to generate a user graphic representation including background deletion 608, performs further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted for the user graphic representation with the background deleted, and then inserts and combines the enhanced user graphic representation with the background deleted into the virtual environment.

[0241] Figure 7C shows image processing combinations 4 to 6.

[0242] In one embodiment of Figure 7C, at least one cloud server 102 can be configured as a STUN server, whereby the peer devices can detect those public IP addresses and the type and information of the NAT behind them that can be used to establish a data connection and data exchange between the peer devices. In another embodiment of Figure 7C, at least one cloud server 102 can be configured for signaling, which can be used for the peer devices to identify and connect to each other's locations and to exchange data through communication coordination performed by at least one cloud server.

[0243] In all of image processing combinations 4 to 6, since at least one cloud server 102 provides services between peer devices A and B, at least one cloud server 102 can use SAMS, SFU, MCU, or other functional server topologies.

[0244] In image processing combination 4, peer device A generates a user graphic representation including background deletion 608, performs further processing or improvement 610 on the user graphic representation with the background deleted to generate an enhanced user graphic representation with the background deleted, and inserts and combines the enhanced user graphic representation with the background deleted into a virtual environment. Then, peer device A transmits the enhanced user graphic representation with the background deleted, which has been inserted and combined into the virtual environment, to peer device B through at least one cloud server acting as a STUN or signaling server.

[0245] In image processing combination 5, peer device A generates a user graphic representation including background deletion 608 and transmits the user graphic representation with the background deleted to peer device B through at least one cloud server acting as a media router server. Peer device B performs further processing or improvement 610 to generate an enhanced user graphic representation with the background deleted on the user graphic representation with the background deleted, and client device B inserts and combines the enhanced user graphic representation with the background deleted into a virtual environment.

[0246] In image processing combination 6, peer device A transmits a camera live data feed received from at least one camera and transmits the unprocessed data to peer device B through at least one cloud server acting as a STUN or signaling server. Peer device B uses the data to generate a user graphic representation including background removal 608, performs further processing or improvement 610 on the user graphic representation with the background removed, generates an enhanced user graphic representation with the background removed, and then inserts and combines the enhanced user graphic representation with the background removed into the virtual environment.

[0247] FIG. 8 shows a user graphic representation-based user authentication system 800 that can be used in an embodiment of the present disclosure. For example, the user graphic representation-based user authentication system 800 can be used to access a user account that can permit access to a virtual environment platform such as the virtual environment platform 108 of FIGS. 1 and 2A.

[0248] The user graphic representation-based user authentication system 800 includes one or more cloud server computers 802 comprising at least one processor 804 and a memory 806 storing data and instructions including a user database 808 storing user data associated with a user account 810 and one or more corresponding user graphic representations 812. The user graphic representation-based user authentication system 800 further comprises a face scan and authentication module 814 connected to a database 808 storing data associated with the user account 810. The one or more cloud server computers 802 are configured to authenticate the user by performing a face scan of the user through the face scan and authentication module 814. The face scan includes extracting face feature data from camera data received from the client device 822 and checking the extracted face feature data for a match with the user graphic representation within the user database 808.

[0249] In the example shown in FIG. 8, the system 800 further includes at least one camera 816 configured to obtain image data 818 from a user 820 of at least one client device 822 that requests access to the user account 810. The at least one camera 816 is connected to at least one client device 822 configured to transmit data captured by the camera 816 to one or more cloud server computers 802 for further processing. Alternatively, the camera 816 can be directly connected to one or more cloud server computers 802. The one or more cloud server computers 802 are configured to authenticate the user by performing a face scan of the user through the face scan and authentication module 814, checking the user database 808 for a match with an existing user graphic representation, and providing the user with access to the user account 810 and the corresponding user graphic representation 812 if the user account 810 is verified and available. Alternatively, if the user account 810 is not available, the one or more cloud server computers 802 are configured to authenticate the user by generating a new user graphic representation 812 along with a new user account 810 to be stored in the user database 808 from the data 818 obtained from the live data feed.

[0250] The user account 810 can be used, for example, to access a virtual environment platform or any other application such as an interactive application, game, email account, university profile account, work account, etc. (e.g., an application that can be linked to the environment platform). The graphic representation-based user authentication system 800 of the present disclosure provides a higher level of convenience and security than a standard camera-based face detection authentication system when, for example, steps for generating the user graphic representation 812 or obtaining the existing user graphic representation 812 from the user database 808 are provided.

[0251] In some embodiments, one or more cloud server computers are further configured to check the date of a matching user graphic representation and determine whether an update to the matching user graphic representation is needed. In one embodiment, in response to checking the date of an available user graphic representation 812 when a user account 810 is available, one or more cloud server computers 802 determine whether an update to an existing user graphic representation 812 is needed by comparing it to a corresponding threshold or security requirement. For example, if there has been a system security update, it may be necessary to update all user graphic representations or at least those created before a specified date. If a user graphic representation 812 is needed, one or more cloud server computers 802 generate an update request for the user graphic representation to the corresponding client device 822. If the user 820 approves the request, one or more cloud server computers 802 or the client device 822 generate the user graphic representation 812 based on data 818 from the live camera feed. If an update is not needed, one or more cloud server computers 802 retrieve the existing user graphic representation 812 from the user database 808 after authentication.

[0252] In some embodiments, the user graphic representation 812 is inserted into a two - dimensional or three - dimensional virtual environment or inserted into a third - party source linked to the virtual environment and combined with the two - dimensional or three - dimensional virtual environment. For example, the user graphic representation 812 can be inserted into a third - party source linked to the virtual environment by overlaying it on the screen of a third - party application or website integrated or coupled with the system of the present disclosure.

[0253] In one example, an overlay of the user graphic representation 812 on the screen of a third-party source is made over a 2D website or application linked to a virtual environment. For example, two or more friends going together on a shopping website can overlay their user graphic representations on the shopping website to explore and / or interact with the website's content. In another example, an overlay of the user graphic representation 812 on the screen of a third-party source is made over a 3D game session linked to a virtual environment. For example, a user can access an e-sports game session linked to a virtual environment through their user graphic representation 812 that can be overlaid on the e-sports game session along with the user graphic representations 812 of other team members. In these examples, such an overlay of the user graphic representation 812 can enable a coherent multicast view of the representations and communications of all users during a visit to a 2D website or during an experience of a 3D game session.

[0254] In some embodiments, the process of generating the user graphic representation 812 is performed asynchronously with respect to the user 820's access to the user account 810. For example, after a face scan, if the user graphic representation-based authentication system 800 determines that the user 820 is already authenticated, the user graphic representation-based authentication system 800 can enable the user 820 to access the user account 810 during the generation of a new user graphic representation 812 and provide it to the user 812 and insert and combine it into the virtual environment as soon as it is ready.

[0255] In some embodiments, one or more cloud server computers 802 further authenticate the user 802 through a login authentication credential that includes a personal identification number (PIN), or a username and password, or a combination of camera authentication and a PIN or username and password.

[0256] In some embodiments, authentication of the user graphic representation-based authentication system 800 is triggered in response to the activation of an invitation link or a deep link sent from one client device 822 to another client device. Clicking on the invitation link or the deep link triggers at least one cloud server computer 802 to request user authentication. For example, the invitation link or the deep link is for an invitation to a phone call, a teleconference, or a video game session, and the invited user can be authenticated through the user graphic representation-based authentication system 800.

[0257] In another embodiment, the face scan uses 3D authentication that includes guiding the user to perform a head movement pattern and extracting 3D face data based on the head movement pattern. This can be done using application instructions stored in at least one server computer that guides the user to perform a head movement pattern, such as performing one or more head gestures, tilting the head horizontally or vertically, or rotating it in a circular motion, generating a gesture pattern by the user, or performing a specific head movement pattern, or a combination thereof, to perform 3D authentication. 3D authentication not only compares and analyzes one view or image, but also recognizes additional features from the data obtained from the live video data feed of the camera. In this embodiment of 3D authentication, the face scan process can recognize additional features from data that may include face data including a head movement pattern, face volume, height, depth of face features, face scars, tattoos, eye color, face skin parameters (e.g., skin color, wrinkles, pore structure, etc.), reflectance parameters, and also the position of such features on the face topology, as in the case of other types of face detection systems. Thus, capturing such face data can increase the capture of realistic faces that can be useful for generating realistic user graphic representations. The face scan using 3D authentication can be performed using a high-resolution 3D camera, a depth camera (e.g., LIDAR), a light field camera, etc. The face scan process and 3D authentication can use deep neural networks, convolutional neural networks, and other deep learning techniques to obtain, process, and evaluate user authentication by using face data.

[0258] FIG. 9 is a schematic diagram of a third-person viewing perspective 900 of the virtual environment 110 through the user graphic representation 120 when the virtual environment 110 is a virtual office.

[0259] The virtual office includes one or more office desks 902, office chairs 904, office computers 906, a projection surface 908 for projecting content 910, and a plurality of user graphic representations 120 representing corresponding users accessing the virtual environment 110 through their client devices.

[0260] The user graphic representation 120 initially is a user 3D virtual cutout and, after the invitation approval process, transitions to a user real-time 3D virtual cutout that includes a user real-time video stream with the background removed, which is generated based on real-time 2D or 3D live video stream data feeds obtained from a camera, or a video with the background removed, or a video with the background not removed. The process can include opening a communication channel as described with reference to FIG. 5 to enable a plurality of interactions within a live session as described with reference to FIG. 4. For example, a user is initially working at the corresponding office computer 906 while sitting in the office chair 904, which can represent an actual action performed by the user in real life. Other users can view the current user status (e.g., through remote sensing 414 of FIG. 4) such that the user is either absent, busy, available, offline, in a conference call, or in a meeting. If the user is available, another user graphic representation can approach the user and send an invitation to participate in a conversation. Both users can decide to, for example, move to a private conference room in the virtual office and start a live session that enables a plurality of interactions. The user can also project desired content (e.g., through screen sharing) onto the projection surface 908.

[0261] In some embodiments, the virtual office further comprises a virtual computer including virtual resources, the virtual resources being from one or more cloud computer resources accessed through a client device and assigned to the virtual computer resources by a management tool. The virtual computer may be associated with an office computer 906. However, the virtual computer may also be associated with a personal home computer or any other computer from where cloud-computer-based virtual computing resources can be accessed. The resources may include memory, network, and processing capabilities required to perform various tasks. Further, in an example of an office space, the virtual computer associated with the virtual office computer 906 can then be coupled to the user's actual office computer, and thus, for example, when the user logs in to such a virtual computer, data stored in the virtual office computer 906 can be utilized from the actual office computer in a physical office or any other space having a physical computer. The virtual infrastructure including all virtual computers associated with the virtual office computer 906 can be managed through a virtual environment platform by using administrative options based on exclusive administrative rights (e.g., provided to the IT team of an organization using the virtual environment 110). Thus, the virtual environment platform of the present disclosure enables virtual office management, expands the possibilities of typical virtual meetings and conferencing applications, enhances the sense of presence of collaboration and interaction, and provides multiple options for streamlining the way collaboration is conducted.

[0262] Figures 10A to 10B are schematic diagrams of a virtual environment as seen through corresponding user graphic representations when the virtual environment is a virtual classroom 1000 according to an embodiment. The user graphic representations of the students and the teacher in Figures 10A to 10B are either user 3D virtual cutouts constructed from photos uploaded by the user or provided by a third party, or user real-time 3D virtual cutouts with the background removed generated based on real-time 2D or 3D live video stream data feeds obtained from a camera, or videos with the background removed, or videos without the background removed.

[0263] In Figure 10A, user graphic representations of a plurality of students 1002 are attending a class remotely provided by the user graphic representation of teacher 1004. Teacher 1004 can project class content 1006 onto one or more projection surfaces 1008 such as a virtual classroom whiteboard. The virtual classroom 1000 can further include a plurality of virtual classroom desks 1010 that can be supported for the user to learn. The students 1002 can be provided with a plurality of interaction options such as raising their hands, screen sharing (e.g., on the projection surface 1008), and laser pointing to specific content as appropriate according to the situation as disclosed with reference to Figure 4. In Figure 10A, the user graphic representation of teacher 1004 is graphically projected onto the projection surface.

[0264] Figure 10B shows an embodiment similar to Figure 10A, the difference being that the user graphic representation of teacher 1004 is sitting behind the virtual desk 1012 and only the content 1006 is shared or projected on the virtual classroom whiteboard projection surface 1008. Since teacher 1004 shares the same virtual space as students 1002 and can move around within classroom 1000, a more realistic and interactive experience is created for both students 1002 and teacher 1004.

[0265] Figure 11 is a schematic diagram of a plurality of virtual camera positions 1100 according to an embodiment.

[0266] In FIG. 11, two user graphic representations 1102, user 3D virtual cutout 1104, and user real-time 3D virtual cutout 1106 have one or more virtual camera positions 1100 for one or more virtual cameras, and each virtual camera position includes a line-of-sight direction, an angle, and a field of view that generate a viewing perspective of the user graphic representation.

[0267] In one embodiment, one or more virtual cameras are at eye level 1108 and are arranged to face outward from the eye level of the user graphic representation 1102. In another embodiment, two (one for each eye) virtual cameras can face outward from the eye levels 1110 of both eyes of the user graphic representation 1102. In yet another embodiment, one or more virtual cameras can face outward from the center 1112 of the position of the head of the user graphic representation 1102. In yet another embodiment, one or more virtual cameras can face outward from the center 1114 of the user graphic representation 1102. In yet another embodiment, one or more virtual cameras can be in front of the user graphic representation 1102, for example, at the height of the head of the user graphic representation 1102, and arranged to face the user graphic representation 1102 when in the self-viewing perspective 1116. One or more virtual cameras can be created while inserting and combining the user graphic representation into the virtual environment as described with reference to FIGS. 6A-7C.

[0268] In one embodiment, the viewing perspective of the user captured by the camera is associated with the viewing perspective of the user graphic representation and the associated virtual camera that uses computer vision to operate the virtual camera. Further, the virtual camera can be automatically updated by tracking and analyzing, for example, the tilt data of the user's eyes and head, or the rotation data of the head, or a combination thereof.

[0269] FIG. 12 is a schematic diagram of a system 1200 for virtual broadcasting from within a virtual environment.

[0270] System 1200 may include one or more server computers. The exemplary system 1200 shown in FIG. 12 includes at least one media server computer 1202 comprising data and instructions implementing a data exchange management module 1208 that manages data exchange between at least one processor 1204 and client devices 1210, and a memory 1206. System 1200 further includes at least one virtual environment 1212 connected to at least one media server computer 1202, and at least one media server computer 1202 is disposed within at least one virtual environment 1212 and includes a virtual broadcast camera 1214 configured to capture a multimedia stream from within at least one virtual environment 1212. At least one virtual environment 1212 is hosted by at least one dedicated server computer connected to at least one media server computer 1202 via a network, or is hosted in a peer-to-peer infrastructure and can be relayed through at least one media server computer 1202. The multimedia stream is transmitted to at least one media server computer 1202 for broadcasting to at least one client device 1210. System 1200 further includes at least one camera 1216 that obtains live feed data from a user 1218 of at least one client device 1210 and transmits the live feed data from the user through at least one client device 1210 to at least one media computer 1202. The live feed data received by at least one media computer 1202 can be generated through a plurality of combinations of image processing disclosed with reference to FIGS. 6A-7C.

[0271] At least one virtual broadcast camera 1214 transmits a multimedia stream to at least one media server computer 1202 to broadcast a corresponding multimedia stream to a receiving client device 1210 based on data exchange management from the at least one media server computer 1202. The multimedia stream is displayed on a corresponding user graphic representation 1222 of a user 1218 of the at least one client device 1210 through a corresponding display. Data exchange management between client devices 1210 by the data exchange management module 1208 includes analyzing an incoming multimedia stream and evaluating the transfer of an outgoing multimedia stream.

[0272] In some embodiments, when transferring an outgoing multimedia stream, the at least one media server computer 1202 uses a routing topology including a Selective Forwarding Unit (SFU), Traversal Using Relay NAT (TURN), Spatially Analyzed Media Server (SAMS), or other suitable multimedia server routing topology, or media processing and transfer server topology, or other suitable server topology. In yet another embodiment, when using a media processing topology, the at least one media server computer 1202 is configured to decode, combine, improve, mix, enhance, extend, compute, manipulate, and encode the multimedia stream. In yet another embodiment, when using a transfer server topology, the at least one media server computer 1202 uses one or more of a Multipoint Control Unit (MCU), a cloud media mixer, and a cloud 3D renderer.

[0273] In some embodiments, the incoming multimedia stream includes user priority data and distance relationship data, and the user priority data includes a higher priority score for user graphic representations closer to the source of the incoming multimedia stream and a lower priority score for user graphic representations farther from the source of the incoming multimedia stream. In one embodiment, the multimedia stream transmitted to at least one media server by at least one client device 1210 and / or broadcast camera 1214 includes data related to user priority and the distance relationship between the corresponding user graphic representation 1222 and the multimedia stream, and the data includes metadata, or priority data, or data class, or spatial structure data, or three-dimensional position, or orientation or movement information, or image data, or media data, and scalable video codec-based video data, or combinations thereof. In yet another embodiment, the priority data includes a higher priority score for users closer to the virtual multimedia stream source 1224 and a lower priority score for users farther from the virtual multimedia stream source 1224. In yet another embodiment, the transfer of the outgoing multimedia stream is based on user priority data and distance relationship data. In one embodiment, the transfer of the outgoing multimedia stream performed by the media server based on user priority data and distance relationship data includes bandwidth optimization and calculation of the resource utilization rate of one or more receiving client devices.

[0274] In some embodiments, at least one virtual broadcast camera 1214 is virtually seen in at least one virtual environment 1212 as a virtual broadcast camera 1214 configured to broadcast a multimedia stream within the at least one virtual environment 1212. The virtual broadcast camera 1214 can be placed near a virtual multimedia stream source 1224 and can also move around within the virtual environment 1212. In further embodiments, the virtual broadcast camera 1214 can be managed through a client device 1210 accessing the virtual environment, operate on the camera perspective updated in the virtual environment, and be configured to broadcast the updated perspective to at least one client device associated with the virtual broadcast camera 1214.

[0275] In some embodiments, the virtual multimedia stream source 1224 includes a live virtual event including one or more of a panel, speech, conference, presentation, webinar, entertainment show, sports event, and performance, and a plurality of user graphic representations of actual speakers speaking remotely (e.g., from their home while recorded by a corresponding camera 1216) are placed within the virtual environment 1212.

[0276] In some embodiments, the multimedia stream can be displayed as a real-time 3D view in a web browser rendered on a client or a cloud computer, or streamed for live viewing on a suitable video platform (e.g., YouTube (trademark) Live, Twitter (trademark), Facebook (trademark) Live, Zoom (trademark), etc.).

[0277] In the example shown in FIG. 12, users A to C access the virtual environment 1212 through the corresponding client devices, and each of users A to C has a camera 1216 that transmits a corresponding multimedia stream to each of users A to C, which can be used for the generation, insertion, and combination of user graphic representations A to C into the virtual environment 1212 as described in embodiments of the present disclosure. Thus, in the virtual environment 1212, each of users A to C has a corresponding user graphic representation A to C. The multimedia stream transmitted by at least one camera 1216 through at least one client device 1210, and the multimedia stream transmitted by at least one broadcast camera 1214 to at least one media server computer 1202, include data related to user priority and the distance relationship between the corresponding user graphic representation and the multimedia stream. This data includes, for example, metadata, priority data, data classes, spatial structure data, 3D position, orientation, or movement information, image data, media data, scalable video codec-based video data, and the like. The data can be used by the data exchange management module 1208 to manage data exchange between client devices 1210, including analyzing and optimizing the incoming multimedia stream and evaluating and optimizing the transfer of the outgoing multimedia stream.

[0278] Thus, for example, since user graphic representation A is close to virtual multimedia stream source 1224 in virtual environment 1212, the transmission of the outgoing multimedia stream can be optimized, for example, to include an image with a higher resolution than the resolutions provided to user graphic representations B and C by user graphic representation A. The multimedia stream can be displayed in first person within virtual environment 1212 by the user, for example, through user graphic representation 1222 via client device 1210. In some examples, the multimedia stream is displayed as a real-time 3D view in a web browser rendered on a client or cloud computer. The user can view the multimedia stream of an event (e.g., webinar, conference, panel, speech, etc.) as a real-time 3D view in a web browser rendered on a client or cloud computer, or stream it for live viewing on a suitable video platform and / or social media.

[0279] FIG. 13 is a schematic diagram of a system 1300 for delivering an application within a virtual environment.

[0280] System 1300 includes at least one cloud server computer 1302 having at least one processor 1304 and a memory 1306 storing data and instructions implementing at least one virtual environment 1308 linked to an application module 1310. The application module 1310 includes one or more installed applications 1312 and application rules 1314 for corresponding multi-user interactions. In response to a selection by a virtual environment host 1316 through a client device 1318, one or more installed applications 1312 are displayed and activated during a session of the virtual environment 1308, and a virtual environment host user graphic representation 1320 and any participant user graphic representation 1322 within the virtual environment 1308 can interact with one or more installed applications 1312 through the corresponding client device 1318. The at least one cloud server computer 1302 manages and processes received user interactions with one or more installed applications 1312 according to the application rules 1314 for multi-user interactions in the application module 1310. The at least one cloud server computer 1302 further transfers the processed interactions to the respective client devices 1318 as appropriate to establish a multi-user session in the virtual environment 1308 and enable a shared experience according to the multi-user interaction application rules 1314.

[0281] In some embodiments, the multi-user interaction application rules 1314 are stored and managed on one or more separate application servers that can be connected to the at least one cloud server computer 1302 through a network.

[0282] In some embodiments, one or more applications are installed from an application installation package available from an application library and provision application services through corresponding application programming interfaces. In yet another embodiment, the application library is filtered by context. In one embodiment, the context filtering is designed to provide only applications relevant to a particular context. For example, host 1316 can contextually filter an application library (e.g., an app store) to find applications related to a particular context (e.g., learning, entertainment, sports, reading, shopping, weather, work, etc.) and select one interesting application to install within application module 1310. In yet another embodiment, the application library is hosted on one or more third-party server computers or on at least one cloud server computer 1302.

[0283] In some embodiments, one or more installed applications are shared with and displayed through a virtual display application installed on a corresponding client device. In one embodiment, when installed and activated, one or more installed applications 1312 are shared with and displayed through a virtual display application 1324 installed on a corresponding client device 1318. The virtual display application 1324 may be configured to receive one or more installed applications 1312 from an application library and publish the one or more selected installed applications 1312 for display through their corresponding client devices 1318 to a conference host user graphic representation 1320 and other participant user graphic representations 1322 in the virtual environment 1308. The virtual display application 1324 may be a type of online or installed file viewer application configured to receive and display the installed applications 1312.

[0284] In some embodiments, the application module 1310 is represented as a 2D screen or 3D volume application module graphic representation 1326 within the virtual environment that displays content from the installed applications 1312 to a user graphic representation 1322 within the virtual environment. In further embodiments, the virtual display application 1324 is represented as a 2D screen or 3D volume that displays content from the installed applications to a user graphic representation within the virtual environment 1308.

[0285] In some embodiments, one or more applications 1312 are installed directly within the virtual environment 1308 before or at the same time as a multi-user session takes place. In other embodiments, one or more applications 1312 are installed through the use of a virtual environment setup tool before starting a multi-user session.

[0286] In some embodiments, some of the application rules for multi-user interaction can define synchronous interactions, or asynchronous interactions, or combinations thereof, and update the user interactions and the respective updated views of one or more applications as appropriate. Both synchronous and asynchronous interactions can be configured through the multi-user interaction application rules 1314 and can be made possible through parallel processing by at least one server computer 1302 or by a separate server computer dedicated to processing individual user interactions with at least one installed application 1312.

[0287] For example, when host 1316 is the teacher, the teacher can select a workbook application that displays the content of the book to the user. The teacher can edit the workbook, and when the student selects to use synchronous interaction and the respective updated views, the student can view the same workbook with or without edits from the teacher through virtual display application 1324. In another example, in a presentation application that includes a presentation file with multiple slides, asynchronous interaction may enable each user to view individual slides asynchronously. In another example, in the case of an educational application, while a student is taking a test, the anatomical structure of the heart is presented, and in this case, the interaction of the student is synchronized for other students to witness and observe the interaction performed by the student. In another example, the teacher can write on a whiteboard, and the student can synchronously view the text written on the whiteboard through a virtual display application. In another example, a video player application can synchronously display a video to all students.

[0288] In some exemplary embodiments, virtual environment 1308 is a classroom, or an office space, or a conference room, or a reception room, or a theater, or a cinema.

[0289] FIG. 14 is a schematic diagram of a virtual environment 1308 based on system 1300 for delivering applications within the virtual environment shown in FIG. 13, according to one embodiment.

[0290] The virtual environment 1308 includes an application module graphic representation 1326 that includes at least one installed application 1312 selected by a host 1316 of the virtual environment 1308, and two users A - B who view and interact with the installed application 1312 through a corresponding virtual display application 1324. As will be understood, user A can view a particular page (e.g., page 1) of this application through the virtual display application 1324, which can be the same as that selected by the host 1316 through the application module graphic representation 1326 that represents the synchronous interaction and management of the installed application 1312. On the other hand, user B can view a page different from both the host 1316 and user A through the asynchronous interaction and management of the installed application 1312 through the virtual display application 1324.

[0291] FIG. 15 is a schematic diagram of a system 1500 for provisioning virtual computing resources within a virtual environment, according to one embodiment.

[0292] The server computer system includes one or more server computers, including at least one server computer 1502 comprising at least one processor 1504, a memory 1506 containing data and instructions for implementing at least one virtual environment 1508, and at least one virtual computer 1510 associated with at least one virtual environment 1508. The at least one virtual computer 1510 receives virtual computing resources from the server computer system. In one embodiment, the at least one virtual computer has a corresponding graphic representation 1512 in the virtual environment 1508. The graphic representation 1512 can provide additional advantages such as facilitating interaction between the user and the virtual computer and enhancing the sense of presence of the user experience (e.g., in the case of a home office experience). Thus, in one embodiment, the at least one virtual computer comprises at least one corresponding associated graphic representation 1512 disposed within the virtual environment 1508, and the at least one virtual computer 1510 receives virtual computing resources from the at least one cloud server computer 1502. The system 1500 further comprises at least one client device 1514 connected to the at least one server computer 1510 through a network. In response to the at least one client device 1514 accessing one or more virtual computers 1510 (e.g., by interacting with the corresponding graphic representation), the at least one cloud server computer 1502 provisions at least one portion of the available virtual computing resources to the at least one client device 1514.

[0293] In some embodiments, the virtual computing resources are accessed (e.g., interacted with) by a user graphic representation 1516 of a user 1518 accessing (e.g., interacting with) one or more graphic representations 1512 of a virtual computer within at least one virtual environment 1508 through a corresponding client device 1514, whereby the corresponding client device 1514 is provisioned.

[0294] In some embodiments, the graphic representation 1512 of the virtual computer is spatially disposed within the virtual environment for access by the user graphic representation. In one embodiment, the configuration of the virtual environment 1508 may include an arrangement of virtual items, furniture, floor plans, etc. for use in education, meetings, work, shopping, services, socializing, and entertainment respectively, associated with a context theme 1508 of the virtual environment. In yet another embodiment, one or more virtual computer graphic representations are disposed within the configuration of the virtual environment 1508 for access by one or more user graphic representations 1516. For example, the virtual computer may be disposed in a virtual room that the user graphic representation 1516 will access when engaging in an activity (such as working on a project in a virtual classroom, laboratory, or office) that requires or may benefit from the ability to use resources associated with the virtual computer.

[0295] In some embodiments, the server computer system is configured to provision at least one portion of the virtual computing resources to at least one client device in response to a user accessing at least one cloud server computer by logging in to at least one client device without accessing the virtual environment. In an exemplary scenario, the virtual computing resources are accessed by a user 1518 who physically logs in to a client device 1514 that is connected to at least one cloud server computer 1502 through a network, and triggers the provisioning of the virtual computing resources to the client device 1514 without accessing the virtual environment. For example, the user 1518 can log in to the cloud server computer 1502 from a home computer, access the virtual computer 1510, and thus receive the virtual computing resources. In another example, the user 1518 can log in to the cloud server computer 1502 from a work computer, access the virtual computer 1510, and thus receive the virtual computing resources.

[0296] In some embodiments, at least one portion of the virtual computing resources is assigned to the client device by a management tool. Thus, the virtual infrastructure, including all related virtual computers, can be managed by using administrative options based on exclusive administrative rights (e.g., provided to the IT team of the organization using the virtual environment).

[0297] In some embodiments, the provisioning of virtual computing resources is performed based on a stored user profile. In one embodiment, the allocation of virtual computing resources is based on a stored user profile that includes one or more of the parameters associated with and assigned to the user profile, including priority data, security data, QOS, bandwidth, memory space, or computing power, or a combination thereof. For example, a user accessing a work virtual computer from home may have a personal profile configured to provide the user with specific virtual computing resources associated with the profile.

[0298] In some embodiments, each virtual computer is a downloadable application available from an application library.

[0299] FIG. 16 is a schematic diagram of a system 1600 that enables ad-hoc virtual communication between user graphic representations, according to one embodiment.

[0300] System 1600 includes one or more cloud server computers 1602 comprising at least one processor 1604 and a memory 1606 storing data and instructions for implementing a virtual environment 1608. The virtual environment 1608 is configured to enable at least one approaching user graphic representation and at least one target user graphic representation in the virtual environment 1608 to open an ad-hoc communication channel and enable an ad-hoc conversation via the ad-hoc communication channel between user graphic representations in the virtual environment 1608. In the example shown in FIG. 16, the system further comprises two or more client devices 1610 connected to one or more cloud server computers 1602 via a network 1612 and accessing at least one virtual environment through corresponding user graphic representations. The virtual environment 1608 enables at least one approaching user graphic representation 1614 and at least one target user graphic representation 1616 to open an ad-hoc communication channel 1618 from the corresponding user 1620 and enables an ad-hoc conversation between user graphic representations in the virtual environment 1608.

[0301] In some embodiments, opening the ad-hoc communication channel 1618 is performed based on the distance, position, and orientation between user graphic representations, or the current availability status, privacy settings, or status configuration of the ad-hoc communication, or a combination thereof.

[0302] In some embodiments, the ad-hoc conversation is conducted at a location within the virtual environment 1608 where both user graphic representation areas exist. For example, when an approaching user graphic representation 1614 meets the target user graphic representation 1614 in a specific area of a lounge room or an office space, an ad-hoc communication is opened, enabling both users to have a conversation within a specific area of the lounge room or the office space without the need to change locations. In yet another embodiment, the ad-hoc conversation is conducted using the current viewing perspective in the virtual environment. In the above example, an ad-hoc communication is opened, enabling both users to have a conversation without changing the viewing perspective. In other embodiments, the ad-hoc conversation allows for an optional change of the viewing perspective, location, or a combination thereof within the same or a different connected virtual environment where the ad-hoc conversation is taking place.

[0303] In some embodiments, one or more cloud server computers are further configured to generate current visual feedback in the virtual environment signaling that ad-hoc communication is possible. In one embodiment, the user graphic representation receives visual feedback signaling that ad-hoc communication is possible, thereby triggering the opening of an ad-hoc communication channel and signaling the start of an ad-hoc conversation between the user graphic representations.

[0304] In some embodiments, the ad-hoc conversation includes transmitting and receiving real-time audio and video displayed from the user graphic representation.

[0305] In some embodiments, the user corresponding to the approaching user graphic representation 1614 selects and clicks on the target user graphic representation 1616 before opening the ad-hoc communication channel 1618.

[0306] In some embodiments, one or more cloud server computers are further configured to open an ad hoc communication channel in response to an invitation acceptance. For example, a user corresponding to the approaching user graphic representation 1614 further sends an ad hoc communication participation invitation to the target user graphic representation 1616 and receives approval of the invitation from the target user graphic representation 1614 before opening the ad hoc communication channel 1618.

[0307] In some embodiments, the ad hoc communication channel 1618 is enabled through at least one cloud server computer or as a P2P communication channel.

[0308] FIG. 17 is a diagram illustrating one embodiment of a method 1700 for enabling interaction in a virtual environment, according to one embodiment.

[0309] The method 1700 for enabling interaction in a virtual environment according to the present disclosure begins by providing a virtual environment platform including at least one virtual environment in the memory of one or more cloud server computers including at least one processor, at steps 1702 and 1704.

[0310] The method, as seen at steps 1706 and 1708, receives a live data feed from a user of a client device from at least one camera and then generates a user graphic representation from the live data feed. Next, the method 1700 inserts the user graphic representation into the three-dimensional coordinates of the virtual environment, as seen at step 1710.

[0311] Thereafter, at step 1712, the method updates the user graphic representation within the virtual environment from the live data feed. Finally, at step 1714, the method processes data generated from interaction in at least one virtual environment through the corresponding graphic representation present within the virtual environment and ends at step 1716.

[0312] FIG. 18 is a diagram showing an embodiment of an image processing method 1800 according to an embodiment.

[0313] Method 1800 begins at steps 1802 and 1804 by providing data and instructions implementing an image processing function to the memory of at least one cloud server computer. Subsequently, at step 1806, method 1800 obtains a live data feed from at least one camera from at least one user of at least one corresponding client device. Next, at step 1808, method 1800 generates a user graphic representation by one or more combinations of image processing of one or more cloud server computers and at least one client device (e.g., the combinations of image processing of FIGS. 6A-7C), and then ends the process at step 1810. The one or more cloud server computers and at least one client device may interact through a hybrid system architecture from the present disclosure comprising a P2P side and a client-server side (e.g., the hybrid system architecture 300 of FIG. 3).

[0314] FIG. 19 is a diagram showing a user authentication method 1900 based on a user graphic representation according to an embodiment.

[0315] Method 1900 begins in steps 1902 and 1904 by providing, in the memory of one or more cloud server computers, a user database that stores user data associated with a user account and a corresponding user graphic representation, and a face scan and authentication module connected to the user database. Method 1900 continues in step 1906 by receiving, from a client device, a request to access a user account, and then, in step 1908, performing a face scan of a user of the at least one client device through the face scan and authentication module by using an image received from at least one camera that can be connected to the at least one client device and / or the one or more cloud server computers. Subsequently, in check 1910, method 1900 checks the user database for a match of user data associated with the user account. If the user account is available, method 1900 provides, in step 1912, to the user, a corresponding user graphic representation along with access to the user account. If NO, i.e., if the user account is not available, method 1900 subsequently generates, in step 1914, from the data, a new user account and a new user graphic representation to be stored in the user database along with access to the user account. The process ends in step 1916.

[0316] FIG. 20 is a block diagram of a method 2000 for virtual broadcasting from within a virtual environment, according to one embodiment.

[0317] Method 2000 begins in step 2002 by providing, in the memory of at least one media server, data and instructions that implement a client device data exchange management module that manages data exchange between client devices. Subsequently, method 2000 captures, in step 2004, a multimedia stream with a virtual broadcast camera disposed within at least one virtual environment connected to the at least one media server.

[0318] Subsequently, in step 2006, method 2000 transmits a multimedia stream to at least one media server for broadcasting to at least one client device. Subsequently, in step 2008, method 2000 obtains live feed data from at least one camera through at least one client device from a user of the at least one client device.

[0319] Subsequently, in step 2010, the method includes performing data exchange management that analyzes and optimizes an incoming multimedia stream and live feed data from a user from within at least one virtual environment, and evaluates and optimizes the transfer of an outgoing multimedia stream. Finally, in step 2012, method 2000 ends by broadcasting a corresponding multimedia stream to a client device based on the data exchange management, where the multimedia stream is displayed in a user graphic representation of a user of the at least one client device.

[0320] FIG. 21 is a block diagram of a method 2100 for delivering an application within a virtual environment according to one embodiment.

[0321] Method 2100 begins in step 2102 by providing in the memory of at least one cloud server computer at least one virtual environment and an application module that includes one or more installed applications and application rules for corresponding multi-user interactions, where the application module is linked to the virtual environment and is displayed within the virtual environment. Subsequently, in step 2104, method 2100 receives a selection command from a virtual environment host. Next, in step 2106, method 2100 displays and activates one or more installed applications during a session of the virtual environment, enabling interaction through a client device where a user graphic representation of the virtual environment host within the virtual environment corresponds to any participant user graphic representation.

[0322] Subsequently, at step 2108, method 2100 receives user interactions with one or more installed applications. Thereafter, method 2100 manages and processes the user interactions with one or more installed applications according to the application rules for multi-user interactions in the application module, as seen at step 2110. Finally, method 2100 ends by transferring the processed interactions accordingly to each client device to establish a multi-user session that enables a shared experience according to the application rules.

[0323] FIG. 22 is a block diagram of a method 2200 for provisioning virtual computing resources within a virtual environment, according to one embodiment.

[0324] Method 2200 begins at step 2202 by providing, in the memory of at least one cloud server computer, at least one virtual computer and a virtual environment comprising one or more graphic representations representing the virtual computer. Subsequently, the method receives, at step 2204, virtual computing resources from at least one cloud server computer by the virtual computer. Next, at step 2206, the method receives access requests to one or more virtual computers from at least one client device. Finally, at step 2208, the method ends by provisioning a portion of the available virtual computing resources to at least one client device based on the needs of the client device.

[0325] FIG. 23 is a block diagram of a method 2300 that enables ad-hoc virtual communication between user graphic representations.

[0326] Method 2300 begins, in step 2302, by providing a virtual environment in the memory of one or more cloud server computers comprising at least one processor. Next, in step 2304, the method detects two or more client devices that are connected to the one or more cloud server computers via a network and access at least one virtual environment through corresponding graphical representations. Finally, in step 2306, method 2300 ends by opening an ad hoc communication channel and enabling an ad hoc conversation between user graphical representations in the virtual environment in response to at least one user graphical representation approaching another user graphical representation.

[0327] A computer-readable medium storing instructions configured to cause one or more computers to perform any of the methods described herein is also described. As used herein, the term "computer-readable medium" includes volatile and non-volatile removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Generally, the functionality of the computing devices described herein can be implemented in computing logic embodied in hardware or software instructions that can be written in programming languages such as C, C++, COBOL, JAVA™, PHP, Perl, Python, Ruby, HTML, CSS, JavaScript, VBScript, ASPX, C#, and other Microsoft.NET™ languages. The computing logic can be compiled into an executable program or written in a programming language that is interpreted. Generally, the functionality described herein can be implemented as logical modules that can be replicated, merged with other modules, or divided into sub-modules to provide higher processing capabilities. The computing logic can be stored in any type of computer-readable medium (e.g., a non-transitory medium such as memory or storage media) or a computer storage device and executed by one or more general-purpose or special-purpose processors, thus resulting in a special-purpose computing device configured to provide the functionality described herein.

[0328] Specific embodiments have been described and are shown in the accompanying drawings, but such embodiments are merely illustrative of a broad invention and not limiting, and those skilled in the art will be able to come up with various other modifications, so it should be understood that the present invention is not limited to the specific configurations and arrangements shown and described. Therefore, the description should be regarded as illustrative rather than limiting.

Claims

1. A user authentication system, comprising one or more cloud server computers comprising at least one processor and a memory storing data and instructions, the memory storing a user database storing user data associated with user accounts and one or more user graphic representations, and a face scan and authentication module connected to the database, comprising: The one or more cloud server computers authenticate a user by performing a face scan of the user through the face scan and authentication module; The face scan extracts face feature data from live camera feed data received from a client device and checks for a match with the one or more user graphic representations associated with the user account in the user database, the one or more user graphic representations being configured to visually represent the user within a two-dimensional or three-dimensional virtual environment; if a matching user graphic representation is found in the check step, providing the user with access to the user account; if no matching user graphic representation is found in the check step, generating a new user graphic representation configured to visually represent the user within the virtual environment together with a new user account from the live camera feed data from which the face feature data was extracted, and providing access to the new user account; configured to perform The face scan uses 3D authentication and includes recognizing features for generating the new user graphic representation configured to visually represent the user within the virtual environment from the live camera feed data. system.

2. The system according to claim 1, wherein the one or more user graphic representations are a user 3D virtual cutout, or a user real-time 3D virtual cutout with the background removed, or a video with the background removed, or a video without the background removed.

3. The one or more cloud server computers are further configured to animate a matching user graphic representation or a new user graphic representation, and animating the matching user graphic representation includes applying a machine vision algorithm by a client device or at least one cloud server computer to each user graphic representation to recognize a user's facial expression and graphically simulate the facial expression in the user graphic representation. The system according to claim 2.

4. The one or more cloud server computers are further configured to check a date of a matching user graphic representation and determine whether an update of the matching user graphic representation is necessary. The system according to claim 1.

5. The user graphic representation is inserted into a two-dimensional or three-dimensional virtual environment or into a third-party source linked to the virtual environment. The system according to claim 1.

6. The process of generating the new user graphic representation is performed asynchronously with respect to user access to the user account. The system according to claim 1.

7. The one or more cloud server computers are further configured to authenticate a user through a login authentication credential including a personal identification number (PIN), or a username and password, or a combination thereof. The system according to claim 1.

8. The authentication is triggered in response to activation of an invitation link sent from one client device to another client device. The system according to claim 1.

9. The 3D authentication includes guiding a user to execute a head movement pattern and extracting 3D face data based on the head movement pattern. The system according to claim 1.

10. The 3D authentication uses one or more deep learning techniques to authenticate a user using the 3D face data. The system according to claim 1.

11. A user authentication method, providing a user database stored in a memory of one or more cloud server computers for storing user data associated with a user account and one or more user graphic representations, and a face scan and authentication module connected to the database. Receiving an access request to a user account from a client device, Performing a user face scan through a face scan and authentication module by extracting face feature data from live camera feed data captured by at least one camera communicating with the client device, Checking the extracted face feature data for a match with the one or more user graphic representations configured to visually represent the user within a two-dimensional or three-dimensional virtual environment and associated with the user account in the user database, If a matching user graphic representation is found in the check step, providing the user with access to the user account, If no matching user graphic representation is found in the check step, generating a new user graphic representation configured to visually represent the user within the virtual environment along with a new user account from the live camera feed data from which the face feature data was extracted, and providing access to the new user account, including, The face scan further includes using 3D authentication and recognizing features for generating the new user graphic representation configured to visually represent the user within the virtual environment from data obtained from the live camera feed data. A method.

12. The method according to claim 11, wherein the user graphic representation is a user 3D virtual cutout, or a user real-time 3D virtual cutout, or a video with the background removed, or a video with the background not removed.

13. The one or more cloud server computers are further configured to animate a matching user graphic representation or a new user graphic representation. Animating the matching user graphic representation includes recognizing the user's facial expression and applying a machine vision algorithm by the client device or at least one cloud server computer to the generated user 3D virtual cutout to graphically simulate the facial expression in the user 3D virtual cutout. The method according to claim 12.

14. If a matching user graphic representation is found in the check step, Checking the dates of consistent user graphic representations, and determining whether an update of the consistent user graphic representation is necessary by comparing with corresponding threshold values or security requirements, at least partially based on said date; and in the affirmative case of determining that an update of the consistent user graphic representation is necessary, generating an update request for the user graphic representation; and The method according to claim 11, further comprising.

15. The method according to claim 11, further comprising inserting the user graphic representation into a two-dimensional or three-dimensional virtual environment or into a third-party source linked to the virtual environment.

16. The method according to claim 11, wherein the process of generating the new user graphic representation is performed asynchronously with respect to user access to the user account.

17. The method according to claim 11, further comprising authenticating the user through a login authentication credential including a username and a password.

18. The method according to claim 17, wherein said authentication is triggered in response to activation of an invitation link.

19. On at least one server computer comprising a processor and a memory, providing a user database that stores user data associated with a user account and one or more user graphic representations in the memory of one or more cloud server computers, and a face scan and authentication module connected to the database; and receiving an access request to the user account from a client device; and performing a face scan of the user through the face scan and authentication module by extracting face feature data from live camera feed data captured by at least one camera communicating with the client device; and checking the extracted face feature data for a match with the one or more user graphic representations associated with the user account in the user database, configured to visually represent the user within a two-dimensional or three-dimensional virtual environment; and if a matching user graphic representation is found in the check step, providing access to the user account to the user. If a matching user graphic representation is not found in the check step, a new user graphic representation configured to visually represent the user within the virtual environment is generated from the live camera feed data from which the face feature data was extracted, along with a new user account, and access to the new user account is provided; stores instructions configured to cause; The face scan further includes using 3D authentication to recognize features for generating the new user graphic representation configured to visually represent the user within the virtual environment from the live camera feed data. A computer-readable recording medium.

20. If a matching user graphic representation is found in the check step, checking the date of the matching user graphic representation; determining whether an update of the matching user graphic representation is necessary based at least in part on the date; in the affirmative case of whether an update of the matching user graphic representation is necessary, generating a request to update the user graphic representation; The computer-readable recording medium according to claim 19, further comprising.

Citation Information

Patent Citations

  • Authentication device, program, and recording medium

    JP2007122400A

  • Biometrics device, biometrics method, and its program

    JP2011008529A

  • Information processor and method

    JP2013011933A

  • JPP6645603B

  • Facial recognition authentication system including path parameters

    US20200210562A1