System and method for rendering scene in virtual environment

By optimizing virtual environment rendering through avatar capture modules and server architecture, the problem of computational and network overload in the metaverse is solved, achieving efficient rendering and a smooth user experience on resource-constrained devices, adapting to large-scale participants and dynamic changes.

CN121569264APending Publication Date: 2026-02-24MUXIC LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480023164.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-08
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

In virtual environments, especially in the metaverse, there are issues of computational and network overload, resulting in a choppy user experience. This is particularly true during large-scale events where avatars, scenes, and videos need to be updated frequently, and existing technologies struggle to effectively manage the different perspectives and dynamic elements of a large number of users.

Method used

It adopts an avatar capture module and server architecture, including avatar capture, dynamically allocated streaming servers and a central server. It optimizes scene rendering through rendering configuration files, divides users and object groups, uses a 2D rendering output format, reduces the computational burden on the client, and utilizes a browser application to display the scene.

Benefits of technology

It enables efficient rendering of virtual environments on resource-constrained devices, reduces latency and computational load, ensures a smooth user experience, adapts to different user perspectives and dynamic changes, and supports virtual activities with a large number of participants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121569264A_ABST
    Figure CN121569264A_ABST
Patent Text Reader

Abstract

A system and method for rendering a scene in a virtual environment. The system comprises: an avatar capture module for capturing inputs associated with a plurality of users and objects arranged to be presented in the virtual environment; and a server architecture comprising at least one server for rendering the scene comprising a plurality of avatars and virtual objects in the virtual environment based on a rendering profile associated with a configuration of a user display device; wherein the scene comprising a plurality of avatars and virtual objects representing all or part of the captured plurality of users and objects is displayed on the user display device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a system and method for rendering a scene in a virtual environment, and more particularly (but not exclusively) to a method for displaying the rendered scene using a resource-constrained device. Background Technology

[0002] Hosting virtual events in virtual environments such as the “metaverse” is an excellent way to stream large-scale events, such as concerts, talks, presentations, or seminars, to different users. In these environments, users can be represented by mobile avatars within the virtual environment to demonstrate their presence and interact with the host or other users. When these avatars appear in the virtual environment scene, their visual state may need to be precisely updated multiple times per second (e.g., 30 frames per second) to ensure a smooth visual experience for users viewing the scene through a display device such as a head-mounted VR headset.

[0003] Depending on the number of users simultaneously in the virtual environment, if each user views the metaverse from a different angle or position, different scenes may need to be computed for each frame. Other "dynamic" elements also exist in the metaverse, such as streaming video on a large screen, which is quite common in many metaverse-based gatherings today. Similarly, the video needs to be continuously updated according to the frame rate. Avatars, scenes, and video together constitute a significant challenge, including computational and network overload. Summary of the Invention

[0004] According to a first aspect of this disclosure, a system for rendering a scene in a virtual environment is provided. The system includes: an avatar capture module for capturing input associated with a plurality of users and objects set to be presented in the virtual environment; and a server architecture including at least one server for rendering the scene, comprising a plurality of avatars and virtual objects, in the virtual environment based on a rendering profile associated with a configuration of a user display device; wherein the scene, comprising a plurality of avatars and virtual objects representing all or a portion of the captured plurality of users and objects, is displayed on the user display device.

[0005] According to the first aspect, the rendering profile is associated with a plurality of technical limitations, which are associated with at least one of the network, processing, and display specifications of the user display device.

[0006] According to the first aspect, the rendering profile includes: a 2D rendering output format for display using a browser application, and / or an audio playback format for output using the browser application.

[0007] According to the first aspect, the rendering profile also includes the image resolution and / or level of detail of the scene rendered from the server and / or transmitted to the user display device.

[0008] According to the first aspect, the plurality of users and objects are divided into at least a first group of users and objects and a second group of users and objects; wherein the scene displayed on the user display device includes: a main part of the scene, which includes a first group of avatars and virtual objects representing the identified first group of users and objects; and a secondary part of the scene, which includes a second group of avatars and virtual objects representing the identified second group of users and objects.

[0009] According to the first aspect, the avatar and / or virtual object representing the client user of the user display device is included in the secondary part of the scene.

[0010] According to the first aspect, the secondary portion of the scenario is one of a plurality of sub-regions of a virtual space in the virtual environment; wherein each of the plurality of sub-regions includes a maximum number of client users not exceeding a predetermined threshold.

[0011] According to the first aspect, the multiple sub-regions are isolated from each other.

[0012] According to the first aspect, a discrete set of scenes representing different views in the virtual environment is generated and shared by multiple client users who choose to be located at the same location in the virtual environment.

[0013] According to the first aspect, the server architecture includes a streaming server for rendering the scene and transmitting the scene to the user display device to facilitate the corresponding client user to enter the virtual environment; wherein, the streaming server is a dynamically allocated virtual machine.

[0014] According to the first aspect, the server architecture also includes a central server, which is used to dynamically allocate the streaming servers, monitor the operation of the streaming servers, and assist in establishing a secure network connection between the streaming servers and the user display devices after the client users who enter the virtual environment using the corresponding user display devices are authenticated.

[0015] According to the first aspect, the 2D rendering output format includes a plurality of 2D images representing a 3D scene viewed by the client user on the browser application.

[0016] According to a second aspect of this disclosure, a method for rendering a scene in a virtual environment is disclosed. The method includes: capturing input associated with a plurality of users and objects configured to be presented in the virtual environment; rendering a scene comprising a plurality of avatars and virtual objects in the virtual environment based on a rendering profile associated with a configuration of a user display device; and displaying on the user display device the scene comprising all or a portion of the plurality of avatars and virtual objects representing the captured plurality of users and objects.

[0017] According to the second aspect, the rendering profile is associated with a plurality of technical limitations, which are associated with at least one of the network, processing, and display specifications of the user display device.

[0018] According to the second aspect, the rendering configuration file includes: a 2D rendering output format for display using a browser application, and / or an audio playback format for output using the browser application.

[0019] According to the second aspect, the rendering profile also includes the image resolution and / or level of detail of the scene rendered from the server and / or transmitted to the user display device.

[0020] According to the second aspect, the method further includes the step of dividing the plurality of users and objects into at least a first group of users and objects and a second group of users and objects; wherein displaying the scene on the user display device includes: displaying a main portion of the scene, which includes a first group of avatars and virtual objects representing the identified first group of users and objects; and displaying a secondary portion of the scene, which includes a second group of avatars and virtual objects representing the identified second group of users and objects.

[0021] According to the second aspect, the avatar and / or virtual object representing the client user of the user display device is included in the secondary part of the scene.

[0022] According to the second aspect, the secondary part of the scenario is one of a plurality of sub-regions of a virtual space in the virtual environment; wherein each of the plurality of sub-regions includes a maximum number of client users not exceeding a predetermined threshold.

[0023] According to the second aspect, the multiple sub-regions are isolated from each other.

[0024] According to the second aspect, a set of discrete scenes representing different views in the virtual environment is generated and shared by multiple client users who choose to be located at the same location in the virtual environment.

[0025] According to the second aspect, the method further includes the step of dynamically allocating a streaming server, wherein the streaming server is used to render the scene and transmit the scene to the user display device to facilitate the corresponding client user to enter the virtual environment; wherein the streaming server is a virtual machine.

[0026] According to the second aspect, the method further includes: providing a central server, the central server being used to: dynamically allocate the streaming server, monitor the operation of the streaming server, and assist in establishing a secure network connection between the streaming server and the user display device after a client user entering the virtual environment using a corresponding user display device has been authenticated.

[0027] According to the second aspect, the 2D rendering output format includes multiple 2D images representing a 3D scene viewed by the client user on the browser application. Attached Figure Description

[0028] Specific embodiments of this disclosure will now be described by way of example with reference to the accompanying drawings.

[0029] Figure 1 This is a schematic diagram of a computer server for implementing a system for rendering a scene in a virtual environment, according to an embodiment of the present disclosure.

[0030] Figure 2 This is an illustration of an exemplary operation of a front-end and back-end separation design for cloud gaming.

[0031] Figure 3A This is a block diagram of a system for rendering a scene in a virtual environment according to an embodiment of the present disclosure.

[0032] Figure 3B It is shown Figure 3A The system uses a browser-based approach to illustrate the operation, where rendering is done on the server side.

[0033] Figure 3C It is shown Figure 3A The system uses an application-based approach to illustrate operations, where rendering is performed on the terminal device.

[0034] Figure 4 This is an illustration of a flexible server architecture for rendering a scene in a virtual environment according to an embodiment of the present disclosure.

[0035] Figure 5 This is a block diagram illustrating a data transmission design for load balancing and automatic discovery according to an embodiment of the present disclosure.

[0036] Figure 6This is a block diagram illustrating the operation and automation core of a central server according to an embodiment of the present disclosure.

[0037] Figure 7 This is a block diagram illustrating a central server and user authentication process according to an embodiment of the present disclosure.

[0038] Figure 8 This is a flowchart illustrating a multimedia streaming process according to an embodiment of the present disclosure. Detailed Implementation

[0039] Reference Figure 1 An embodiment of this disclosure is illustrated. This embodiment aims to provide a system for rendering a scene in a virtual environment. The system includes: an avatar capture module for capturing input associated with multiple users and objects set to be presented in the virtual environment; and a server architecture including at least one server for rendering a scene including multiple avatars and virtual objects in the virtual environment based on a rendering profile associated with a configuration of a user display device; wherein the scene including multiple avatars and virtual objects representing all or part of the captured multiple users and objects is displayed on the user display device.

[0040] In this exemplary embodiment, the interface and processor are implemented by a computer with a suitable user interface. The computer can be implemented by any computing architecture, including: a portable computer, tablet computer, personal computer (PC), smart device, Internet of Things (IoT) device, edge computing device, client / server architecture, "dumb" terminal / host architecture, cloud computing-based architecture, or any other suitable architecture. The computing device can be appropriately programmed to implement this disclosure.

[0041] For example, the system can support multiple users (e.g., meeting participants) logging into the system hosting the virtual environment and entering a virtual meeting room to participate in a virtual meeting. Each meeting participant can "see" the avatars of other meeting participants in the same meeting room and interact with them in the virtual environment by entering text or voice messages, performing gestures, etc. In addition, virtual objects can also be displayed and even used in the virtual environment.

[0042] Optionally, various types of events, such as conferences, concerts, or musical performances, can be held in the virtual environment, where audiences can log into the system to enter a virtual, computer-generated venue to watch speakers or performers perform on a virtual stage within that virtual venue. In this example, audiences can also interact with other audiences, for example, by controlling their user avatars to make actions or movements within the virtual venue, which can be seen by other users in the same room / venue.

[0043] Conference participants or audience members can choose to enter the virtual environment using a head-mounted display device. This device can extend the display content to include virtual rooms, virtual objects, user avatars, and user interfaces. Alternatively, the virtual environment can also be visualized using other 2D or 3D display devices, such as smartphone displays, personal computer or tablet screens, or larger screens that can be mounted on a room's wall.

[0044] Some exemplary embodiments of this disclosure are particularly applicable to display devices where processing power may be limited, such as handheld devices, low-power devices, or display devices not designed for high-intensity processing (e.g., smart TVs or wall-mounted display panels). Preferably, reference is made to... Figure 2 The cloud-based gaming system 200 can centralize most or all critical computations on the server side 202, while only requiring the user's display device 204 to be able to display the computation results for each frame. This requirement can be easily met by a wide range of devices, including low-end devices. Therefore, participating users only need to use a browser application such as a web browser to participate in the metaverse application; that is, user input is provided on the client side and the output is viewed through a web client.

[0045] This differs from other metaverse implementations (or massively multiplayer online game games with similar architectures). In other metaverse implementations, users may be required to install proprietary or specially designed applications on their devices / phones to allow most of the overall computation to be performed on the client side. In these examples, for a satisfactory user experience, users / clients may need relatively high-end computing devices / phones, and installing proprietary applications and ensuring that the applications function correctly on the device may also be necessary. Advantageously, such an architecture allows players to easily enter virtual environments such as metaverses using any device pre-installed with any browser application.

[0046] In a preferred embodiment, the method for rendering and displaying a virtual environment scene can be considered "web-based" because the act of joining the metaverse is no different from clicking and interacting with a website on the Internet, while leaving all or most of the computational tasks to the server side. Advantageously, the device requirements for the user can be kept very low because the web client on the user side only needs to be able to display images (metaverse scene) (more preferably, 2D images as further described in this disclosure), play back audio via a Web-Audio API, and collect input from human players using comments or gestures (if any). The player's input may affect some objects within the metaverse and thus affect the next scene the player will see.

[0047] like Figure 1 The diagram illustrates a computer system or computer server 100, an exemplary embodiment of which can be used to implement a virtual environment, such as a virtual environment for a concert performance or large event with a large number of participants. In this embodiment, the system includes server 100. Server 100 includes appropriate components for receiving, storing, and executing appropriate computer instructions. These components may include: a processing unit 102, which includes a central processing unit (CPU), a math co-processing unit (Math Processor), a graphics processing unit (GPU), or a tensor processing unit (TPU) for tensor or multidimensional array computation or manipulation operations; read-only memory (ROM) 104; random-access memory (RAM) 106; input / output (I / O) devices (e.g., disk drive 108); input devices 110 (e.g., Ethernet port, USB port, etc.); a display 112, such as a liquid crystal display, a light-emitting display, or any other suitable display; and a communication link 114. Server 100 may include instructions that can be stored in ROM 104, RAM 106, or disk drive 108 and executed by processing unit 102. Multiple communication links 114 can be provided, which can be connected differently to one or more computing devices, such as servers, personal computers, terminals, wireless or handheld computing devices, IoT devices, smart devices, and edge computing devices. At least one of the multiple communication links can be connected to an external computing network via a telephone line or other types of communication link.

[0048] Server 100 may include storage devices such as disk drives 108, which may include solid-state drives, hard disk drives, optical drives, tape drives, or remote or cloud-based storage devices. Server 100 may use a single disk drive or multiple disk drives, or remote storage services 120. Server 100 may also have a suitable operating system 116 residing on the disk drives or in the ROM of server 100.

[0049] Computers or computing devices can also provide the necessary computing power to operate or interface with machine learning networks of neural networks to provide various functions and outputs. Neural networks can be implemented locally, or they can be accessed or partially accessed through servers or cloud-based services. Machine learning networks can also be untrained, partially trained, or fully trained, and / or can be retrained, modified, or updated over time.

[0050] Reference Figure 3A This illustration shows an embodiment of a system 300 for rendering a scene in a virtual environment. In this embodiment, a server 100 is used as part of the system 300 for rendering videos, images, or image sequences of the virtual environment. These videos, images, or image sequences can be displayed by a display device, and the display device has at least a minimum capability to display these images / videos. The server can be one of multiple servers in a server architecture, although in some exemplary embodiments, the server can be a virtual machine (VM) based server in such a server architecture, implemented on a single physical server 100.

[0051] In one exemplary application, system 300 can be used to facilitate multiple users entering a virtual environment simultaneously, for example, in a virtual concert held in a virtual environment, where performers, such as singers, can perform on a virtual stage in a virtual hall watched by a large audience. In this example, both the singer and the audience can be considered "client users," represented by avatars in the virtual environment, and the user's input can be captured by an avatar capture module or in input capture processing 302, for example, by capturing the user's actions and gestures using a camera or motion sensor. Furthermore, the avatar capture module can be used to capture other forms of input 304 associated with the user and / or objects, such as cheering props or any supported physical object recognizable by a camera, as well as audio such as speech or sound from the user's physical environment. Preferably, in the virtual environment, the user and the recognized object can be represented by an avatar and a virtual object, respectively.

[0052] In an optional example, certain users or objects captured by the avatar capture module may not be displayed, even if these users or objects have been successfully identified. This is useful for reducing the processing and communication resources that may be required to render and display avatars or virtual objects in a virtual environment. This can be processed / selected in calculation process 306. In this process, the desired 3D scene 308 may be generated based on the input 304 provided by the capture module, or the selection of objects to be displayed may be performed in a subsequent rendering step.

[0053] Preferably, the system includes at least one server in a server architecture for rendering a scene comprising multiple avatars and virtual objects based on a rendering profile by executing a rendering process 310. This rendering profile is associated with a configuration of a specific user display device, which may have one or more technical limitations, such as limitations in network, processing, and display specifications. For example, the user display device may be connected to a streaming server or transport server via a network link with limited bandwidth. In another example, the user display device may be a low-end smartphone with limited processing power and may also have a low-resolution display panel.

[0054] Preferably, the rendering profile includes a 2D rendering output format for display using a browser application, and / or an audio playback format for output using a browser application. Considering the technological limitations of low-end display devices, the potential advantage of using a rendering profile with a lower bitrate to render the scene is that users can participate in the virtual environment using display devices with low-end hardware configurations, and the user experience will not be significantly affected by artifacts such as latency or delay caused by hardware limitations and / or bandwidth overload.

[0055] Reference Figure 3A This illustrates the overall architecture of a system that runs or implements a virtual environment. In this example, two main events can occur at each time step: (1) an update of the metaverse scene (including all moving objects such as avatars) due to input from the client and other roles (e.g., artists in a virtual concert) and changes to the objects themselves; and (2) rendering the 3D scene 308 into a 2D frame 312 to be displayed on the client device. Where (1) and (2) occur corresponds to different types of methods for implementing a metaverse or virtual environment.

[0056] Figure 3BA browser-based approach is illustrated, in which scene updates and rendering on the client device are all performed by the server. In this embodiment, the client only needs to use a browser to display the results (2D frames) of each time step, which is the final process 314 in each time step. This embodiment is applicable to a combination of low-end computing devices or "lightweight" processing devices (e.g., smartphones) and powerful cloud computing servers, and it may be advantageous to utilize the more powerful computing resources in cloud servers to perform most or all of the computing tasks.

[0057] Figure 3C The diagram illustrates an "application-based (APP)" approach, where scene updates and rendering are entirely performed by an application installed on the client device. Alternatively, other methods are possible, where the load on both sides can be balanced: for example, the server handles scene updates, while the client device handles rendering. Rendering profiles, including a balance of "quality" or "complexity" between rendering and data transfer tasks, can be optimized based on the available resources on the client display device and the server.

[0058] Advantageously, a browser-based approach may be preferred over an application-based approach because even low-end smartphones can run browsers, thus enabling a wider range of client devices to run the metaverse. The inventors believe that for a browser-based approach to be feasible, the server-side "response time"—the time from when a client action occurs to when the scene is changed and the required 2D frames 312 are rendered in response to that action—is preferably negligible or short enough to be imperceptible.

[0059] There may still be a delay between the time the server generates a frame after process 310 and the time it is actually displayed on the browser's screen in step 314. This delay may depend on network conditions and, to a lesser extent, on the capabilities of the client device. Some low-end devices may need to negotiate with the server for a lower resolution frame, or the client display device may need to scale down locally, for example, if the display panel can only display lower resolution images, or if it is necessary to reduce the resolution to minimize the utilization of the low-end device's computing resources.

[0060] Preferably, in the worst case, the rendered object is not a single frame, but as many frames as there are clients. This is because each client's perspective may be different.

[0061] Embodiments of this disclosure provide an architecture and associated methods capable of meeting the response time requirements of the aforementioned browser-based methods. This architecture can be specifically tailored for virtual concert scenarios, where there are audience members and artists performing on stage. It should also be applicable to other deployment scenarios. Artists and audience members are represented by avatars in the virtual space of the metaverse, while real objects can be represented by virtual objects within the virtual environment. Furthermore, actions or interactions within the virtual environment can be rendered and observed as if they were performed in reality.

[0062] Furthermore, the rendering profile may further include the image resolution and / or level of detail of the scene from server rendering and / or communication to the user's display device.

[0063] Since the computational burden falls entirely on the server side, the system can be implemented using an architecture with additional features designed to reduce computational load or required resources.

[0064] Preferably, the plurality of users and objects are divided into at least a first group of users and objects and a second group of users and objects; wherein the scene displayed on the user display device includes: a main part of the scene, which includes a first group of avatars and virtual objects representing the identified first group of users and objects; and a secondary part of the scene, which includes a second group of avatars and virtual objects representing the identified second group of users and objects; wherein avatars and / or virtual objects representing client users of the user display device are included in the secondary part of the scene.

[0065] In this example, the primary part of the scene is what is "visible" to all users participating in the virtual environment, such as the stage and performers in a concert; while the secondary part of the scene is the floor where the audience, including the user of the display device, is located. Given that the audience's focus is on enjoying the performance on stage, the actions or interactions of other users may be less important to the participants.

[0066] Preferably, the secondary part of the scene is one of multiple sub-regions of a virtual space within a virtual environment; wherein each of the multiple sub-regions includes a maximum number of client users not exceeding a predetermined threshold. More preferably, the multiple sub-regions are isolated from each other.

[0067] For example, the entire concert hall could be divided into smaller "rooms" to reduce the amount of interaction between audience avatars within each room, thereby reducing the computational load in a controlled manner. The rooms could be configured to accommodate a maximum number of audience members, determined through load experiments. In this example, these rooms differ from rooms in the physical world; they are completely isolated from each other, as if they exist in different "dimensions." However, all rooms have a clear view of the stage—that is, the artists on stage and other activities (e.g., video streams).

[0068] In the previously described example, the system could be implemented for artists performing on stage at a concert, thus the artists, including other performers such as dancers, are represented by a first set of avatars in the virtual environment. Furthermore, if there are 100 rooms, each supporting up to 100 users entering the same room, each user will only "see" 100 avatars in the secondary parts of the scene, which, combined with the number of performers on stage (i.e., the main part of the scene), effectively reduces the computational load on the server. This is because there is no need to update a scene with more than 10,000 avatars in the same scene.

[0069] Preferably, the spatial division can be an artificial one, whereby (e.g., a system administrator) assigns avatars to multiple "rooms," each accommodating no more than a specific number of avatars. These "rooms" are invisible to the players within them; the players perceive the entire metaverse space. The system's strategy can determine who should go to which room and whether avatars are allowed to switch rooms at runtime. Optionally, the system administrator can assign rooms associated with different rendering profiles; for example, there could be rooms that provide a better experience for users with lower-end display devices, such as rooms with a smaller group of avatars and / or fewer virtual objects, while in other rooms, more avatars can coexist, depending on user preferences. However, it should be noted that the main part of the scene, i.e., the stage, can be substantially the same across the different available rooms.

[0070] Note that this division is not for defining the physical space of the metaverse, but rather for the purpose of reducing computational load by distributing avatars across different dimensions. In any dimension, the user can see the entire metaverse, including the concert stage, but will not see the client avatars in other dimensions.

[0071] Alternatively, another natural division of the metaverse space is provided. This division lies between content within the client's viewpoint and content outside that viewpoint. This viewpoint is determined by the system's strategy or by the client, for example, through view settings preferences, along with the setting of "density" (i.e., the number of avatars per unit space in the metaverse). Clearly, a sufficiently narrow viewpoint can save significant overhead when calculating the scene for each client, because the number of avatars within a narrow viewpoint can also be very limited, effectively restricting the computational resources required to update the scene. This is because each player can only see a subset of all avatars existing in the metaverse, but this may come at the cost of each participant only seeing a small portion of the stage.

[0072] View settings are applied to players roaming in the metaverse. The system calculates the view camera's perspective based on the number of avatars within the view (which must be below a preset threshold). If a stage exists, the system also ensures that the stage is within the view. Since adjusting the perspective in every frame may be undesirable (as this would cause continuous view "jitter"), these settings are only applied when the user zooms in or out or makes long pans.

[0073] Optionally or additionally, computational resources can be further reduced by sharing scenes with other users, depending on the different viewpoints available. Preferably, a discrete set of scenes representing different views in the virtual environment can be generated, which can be shared by multiple client users who choose to be located in the same location within the virtual environment.

[0074] For example, given M (total number of users / players), up to M views might need to be generated per frame, which can translate to a huge load on the server. Furthermore, the number of different views that players can generate is infinite, since the player's translation angle and zoom level are continuous variables. Players located at the same (or nearby) position in the metaverse can share the same view. Additionally, the number of discrete views can be limited to restrict the use of computational resources.

[0075] For each such location, a discrete set of views (based on camera pan and zoom) can be generated, from which the relevant player will choose for their upcoming scene. If the number of players sharing this set of views at a location is large enough to exceed the size of the set, the overall computational load is reduced.

[0076] Preferably, players can point their cameras at any angle within a 3D 360-degree space around their position. (If there is no stage, the system administrator can define certain focal points during setup.) The entire virtual space can be divided into discrete locations, which are more densely packed closer to the stage. A progressive layout of possible viewpoints can be performed for each of these locations; similarly, angles pointing closer to the stage (or focal points) will be more densely packed.

[0077] Similarly, it's possible to allocate different rendering profiles based on the perspective assigned to a specific user, thereby allocating cloud computing resources to ensure a "better experience" at the concert for users who have paid more for more expensive tickets. This is analogous to seats in real-world concert halls, which can be significantly more expensive than seats that don't offer such a superior viewing experience in reality.

[0078] Not wanting to be bound by theory, the inventors further conceived that the system for rendering scenes in a virtual environment should also include a server architecture that optimizes server deployment based on the current load, which may vary significantly at any given moment during metaverse operations. This architecture should be able to respond quickly to changes and be dynamically reconfigurable.

[0079] Furthermore, at the heart of all activities is the streaming of data. Unlike primarily one-way video-on-demand, streaming here is designed to support highly frequent interactions between each client and the metaverse—its changing scenes and its moving objects and avatars. Therefore, minimizing latency may be more important than achieving smooth streaming.

[0080] Preferably, the server architecture includes a streaming server. The server renders the scene and communicates the scene to the user's display device to facilitate access to the virtual environment for the corresponding client user. The streaming server is a dynamically allocated virtual machine. The inventors refer to this as a "flexible and dynamic server architecture."

[0081] Unlike physical concerts, virtual concerts may involve a large number of attendees dynamically entering and leaving throughout the event. Therefore, a resilient server cluster may be required to support the operation of the metaverse and its dynamically changing array of avatars. A resilient architecture can be provided to achieve optimal cost, and this architecture can be built and rebuilt as needed. Reference Figure 4 This diagram illustrates a structural diagram of a server architecture 400 according to an embodiment of the present disclosure. In this example, M / P streaming servers 402 can be deployed, where M represents the total number of clients and P represents the capacity of the streaming server 402. The actual value of P depends on the complexity of the metaverse application under discussion and user requirements. Note that the streaming server 402 can be “virtual,” and Q of them can run on a single GPU virtual machine instance. Therefore, M / P / Q (=V) virtual machine instances will need to be obtained from a cloud / edge service provider. A feasible and efficient architecture must attempt to minimize the value of V. The server architecture can utilize the following techniques to achieve this goal.

[0082] It is worth noting that, Figure 4 Only the architecture is shown, not the actual number of components. Multiple clouds can be deployed depending on the number of clients 404. For each cloud, a public transport server 406 with static IPs (external and internal) is required for the cloud to function. An "avatar capture PC" 408 that captures the movements of human performers on stage and a "multimedia streaming PC" 410 that performs real-time video streaming are connected to the "public transport server" 406 designated in each cloud. This is an exemplary setup for a concert, but for more complex applications, there may be many more types of "PCs (computers)".

[0083] Preferably, the streaming server 402, the cluster transport server 412, and the public transport server 406 form a hierarchical structure, wherein the cluster transport server 412 can be added (or removed) as more clients join (or leave the concert). Note that in one exemplary embodiment, multiple streaming servers 402 can run on a single virtual machine called a P-server. Although in other examples, a dedicated virtual machine can be assigned to each user login. In addition to the number of servers, the connections between these servers can also be dynamically changed to achieve load balancing. The relationship between the different transport servers, including the public transport server 406, the cluster transport server 412, and the local transport server 502, in Figure 5 The diagram further illustrates that the local transport server on the GPU virtual machine requests an IP from the public transport server 406, and then the cluster transport server 412 provides the IP accordingly.

[0084] Preferably, the server architecture also includes a central server, which is used to dynamically allocate streaming servers, monitor the operation of streaming servers, and assist in establishing a secure network connection between the streaming servers and the user display devices after the client users who enter the virtual environment using the corresponding user display devices are authenticated.

[0085] The central server is another important component. (See reference...) Figure 6 The central server 602 may play several crucial roles in the monitoring and automation of system 600. Each client joining the concert will be assigned a P-server 404 by the central server 602, which will also monitor the online status of client 604. The central server 602 can track how many P-servers 404 are online. For example, a failed server can be detected by its failure to render frames within a heartbeat cycle, and the central server 602 will attempt to restart the server. The client connection count is preferably kept below a health threshold, with some extra P-servers 404 reserved as backup. If the number of client 604 increases, new P-servers 404 will be automatically started via a simple script.

[0086] In addition, another role of the central server 602 is in implementing HTTPS. Dynamically setting up SSL certificates and HTTPS availability can be challenging. Typically, certificates are granted based on a verified domain and its server IP. Since many P-servers 404 are deployed on demand, pre-determining IP addresses using measures such as static IPs is impractical. In this example, the central server 602 also assists in establishing a secure network connection between the streaming server 404 and the user display device, for example, after responding to a login request and authenticating a client user entering the virtual environment.

[0087] Preferably, the "central server" 602, with a static public IP address, can then use the DNS API to dynamically set up and create subdomains for each P-server 404. To achieve this, an SSL wildcard certificate can be pre-installed in the P-server image. When a new instance starts, the P-server 404 can establish a socket.IO connection with the central server 602, thus making the P-server's public IP address known. The central server 602 can then use this IP address as the string for the subdomain name. This ensures the uniqueness of the subdomain name because it is assumed that the public IP address is unique.

[0088] Central server 602 can also handle user authentication. For commercial metaverse activities, authentication (i.e., ticketing) is crucial for the proper functioning of the system. (See reference...) Figure 7 The diagram illustrates the process of authentication method 700, where participant login information is stored in a "login database" 702. Each record can be as simple as a single string or token. WebSocket connections can only be established with the correct token, whose bytes must be sent in a correct proprietary format. The central server 602 can also control duplicate logins by monitoring WebSocket traffic. Any brute-force attempt against the P-server 404 will alert the central server 602, which will then notify the system administrator.

[0089] As previously stated, the rendering profile includes a 2D rendering output format for display, and preferably, the 2D rendering output format includes multiple 2D images representing a 3D scene viewed by a client user on a browser application. The inventors refer to this as a "low-latency streaming" implementation 800, see reference 800. Figure 8 .

[0090] The inventors propose that two types of video streaming may occur during metaverse operations: streaming the metaverse scene to each player's client device at the current frame rate, and streaming video from external sources within the metaverse. For the former, the size, resolution, and frame rate of the stream can be adapted in real time based on the device (and network) status. For the latter, adaptation can be performed in the streaming server to reduce the required computation. This is slightly more complex than the former adaptation because the video display within the metaverse may appear at different tilt angles and sizes in different players' views.

[0091] Preferably, the system can stream images rather than video, because video encoding introduces latency. Advantageously, image-based streaming has provided near real-time performance in experiments with a latency of only one frame. Preferably, the latest frames and audio blocks are always delivered to the end user.

[0092] Furthermore, the level of detail in the rendered scene can also be a factor in reducing computational load. Preferably, adaptive Levels of Detail (LOD) can also be used. For example, if a player is viewing an object from a distant location in the metaverse, there is no need to calculate that object in full detail for the next scene, because the resolution of the viewing device simply cannot display such detail.

[0093] In one exemplary operation, the server can automatically strip away details to render the scene based on viewing distance and angle. Users do not need to provide any additional object variations, only the original, most detailed object itself.

[0094] Preferably, the server can combine a group of polygons at a certain Level of Degree (LOD) into a single, larger polygon for the next or another LOD. The selection process examines edges and boundaries (formed by the shape, color, or texture of the polygons) to preserve the object's outline or overall features. The resulting polygons do not need to have the same or even similar dimensions, which is an important flexibility that facilitates the conversion (unlike pixel-based algorithms used to reduce image size).

[0095] The advantage of these embodiments lies in providing a novel metaverse architecture in which the majority of the overall computation is exclusively undertaken by servers. These servers include a set of specialized streaming servers at the forefront, whose task is to generate the final display frames for their assigned clients. Advantageously, the clients are exempt from any substantial computational tasks other than displaying freshly computed images and transmitting any user input to the servers.

[0096] Advantageously, the embodiments described address the challenge of designing flexible and efficient architectures to support metaverse operations in applications with a large number of participants (users / players). An exemplary application is a virtual concert, where artists perform on stage and there are more than 10,000 audience members. Physical concert venues capable of accommodating such a large audience are rare.

[0097] Furthermore, to design a computing architecture capable of meeting all computational requirements and thus providing a smooth operational experience for the metaverse and its applications from the participants' perspective, a novel example of a server architecture is presented. This server, in terms of actual computation, draws a clear line between the client (metaverse player) and the server. It assumes that the server side has readily available and ample cloud and edge resources, while the client side uses ordinary handheld devices such as smartphones or tablets. The aforementioned line naturally arises from the significant gap in computing power between the server and the handheld device.

[0098] While not strictly necessary, the embodiments described with reference to the accompanying drawings can be implemented as an Application Programming Interface (API) or a set of libraries used by the developer, or can be included in another software application, such as a terminal or personal computer operating system or a portable computing device operating system. Typically, since program modules include routines, programs, objects, components, and data files that help perform specific functions, those skilled in the art will understand that the functionality of a software application can be distributed among multiple routines, objects, or components to achieve the same functionality required herein.

[0099] It should also be understood that any suitable computing system architecture can be used where the methods and systems of this disclosure are implemented wholly or partially by a computing system. This computing architecture may include tablet computers, wearable devices, smartphones, Internet of Things devices, edge computing devices, standalone computers, network computers, cloud-based computing devices, and dedicated hardware devices. When using the terms "computing system" and "computing device," these terms are intended to cover any suitable configuration of computer hardware capable of implementing the described functions.

[0100] Those skilled in the art will understand that various variations and / or modifications can be made to this disclosure as illustrated in the specific embodiments without departing from the spirit or scope of this disclosure as broadly described. Therefore, this embodiment should be considered illustrative rather than restrictive in all respects.

[0101] Unless otherwise stated, any references to prior art contained herein should not be construed as an admission that the information is common general knowledge.

Claims

1. A system for rendering scenes in a virtual environment, characterized in that, include: An avatar capture module is used to capture input that is set to be associated with multiple users and objects presented in the virtual environment; as well as A server architecture comprising at least one server for rendering a scene, including multiple avatars and virtual objects, in the virtual environment based on a rendering profile associated with a configuration of a user display device; wherein the scene, including multiple avatars and virtual objects representing all or part of the captured multiple users and objects, is displayed on the user display device.

2. The system according to claim 1, characterized in that, in, The rendering profile is associated with multiple technical limitations, which are associated with at least one of the network, processing, and display specifications of the user display device.

3. The system according to claim 2, characterized in that, in, The rendering configuration file includes: a 2D rendering output format for display using a browser application, and / or an audio playback format for output using the browser application.

4. The system according to claim 3, characterized in that, in, The rendering configuration file also includes the image resolution and / or level of detail of the scene rendered from the server and / or transmitted to the user display device.

5. The system according to claim 4, characterized in that, in, The plurality of users and objects are divided into at least a first group of users and objects and a second group of users and objects; wherein the scene displayed on the user display device includes: The main part of the scenario includes a first set of avatars and virtual objects representing the first group of identified users and objects; and A secondary portion of the scenario includes a second set of avatars and virtual objects representing the identified second set of users and objects; wherein the avatars and / or virtual objects representing client users of the user display device are included in the secondary portion of the scenario.

6. The system according to claim 5, characterized in that, in, The secondary part of the scenario is one of a plurality of sub-regions of a virtual space in the virtual environment; wherein each of the plurality of sub-regions includes a maximum number of client users not exceeding a predetermined threshold; wherein the plurality of sub-regions are isolated from each other.

7. The system according to claim 5, characterized in that, in, A set of discrete scenes representing different views in the virtual environment is generated and shared by multiple client users who choose to be located in the same location in the virtual environment.

8. The system according to claim 4, characterized in that, in, The server architecture includes a streaming server, which is used to render the scene and transmit the scene to the user display device to facilitate the corresponding client user to enter the virtual environment; wherein, the streaming server is a dynamically allocated virtual machine.

9. The system according to claim 7, characterized in that, in, The server architecture also includes a central server, which is used to dynamically allocate the streaming servers, monitor the operation of the streaming servers, and assist in establishing a secure network connection between the streaming servers and the user display devices after the client users who enter the virtual environment using the corresponding user display devices are authenticated.

10. The system according to claim 3, characterized in that, in, The 2D rendering output format includes multiple 2D images that represent a 3D scene viewed by the client user on the browser application.

11. A method for rendering a scene in a virtual environment, characterized in that, include: Capture input associated with multiple users and objects set to be presented in the virtual environment; Based on a rendering profile associated with the user's display device configuration, a scene including multiple avatars and virtual objects is rendered in the virtual environment; as well as The scene, comprising all or part of the plurality of avatars and virtual objects representing the captured plurality of users and objects, is displayed on the user display device.

12. The method according to claim 11, characterized in that, in, The rendering profile is associated with multiple technical limitations, which are associated with at least one of the network, processing, and display specifications of the user display device.

13. The method according to claim 12, characterized in that, in, The rendering configuration file includes: a 2D rendering output format for display using a browser application, and / or an audio playback format for output using the browser application.

14. The method according to claim 13, characterized in that, in, The rendering configuration file also includes the image resolution and / or level of detail of the scene rendered from the server and / or transmitted to the user display device.

15. The method according to claim 14, characterized in that, Also includes: The step of dividing the plurality of users and objects into at least a first group of users and objects and a second group of users and objects; wherein, the step of displaying the scene on the user display device includes: The main part of the scene is displayed, including a first set of avatars and virtual objects representing the first group of identified users and objects; and The secondary portion of the scene is displayed, which includes a second set of avatars and virtual objects representing the identified second set of users and objects; wherein, the avatars and / or virtual objects representing the client users of the user display device are included in the secondary portion of the scene.

16. The method according to claim 15, characterized in that, in, The secondary part of the scenario is one of a plurality of sub-regions of a virtual space in the virtual environment; wherein each of the plurality of sub-regions includes a maximum number of client users not exceeding a predetermined threshold; wherein the plurality of sub-regions are isolated from each other.

17. The method according to claim 15, characterized in that, in, A set of discrete scenes representing different views in the virtual environment is generated and shared by multiple client users who choose to be located in the same location in the virtual environment.

18. The method according to claim 14, characterized in that, Also includes: The steps of dynamically allocating streaming servers; the streaming server is used to render the scene and transmit the scene to the user display device to facilitate the corresponding client user to enter the virtual environment; wherein, the streaming server is a virtual machine.

19. The method according to claim 17, characterized in that, Also includes: A central server is provided; the central server is used to: dynamically allocate the streaming servers, monitor the operation of the streaming servers, and assist in establishing a secure network connection between the streaming servers and the user display devices after the client users who enter the virtual environment using the corresponding user display devices are authenticated.

20. The method according to claim 13, characterized in that, in, The 2D rendering output format includes multiple 2D images that represent a 3D scene viewed by the client user on the browser application.