Generative VR world creation from natural language
Through natural language command processing and a generative virtual environment builder, sky box and 3D object models are generated, solving the complex and time-consuming problems of artificial reality environment design in the existing technology, and achieving fast and easy-to-use virtual environment creation.
Patent Information
- Application Number
- CN202380076922.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-27
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art processes are complex and time-consuming when designing and creating new artificial reality environments and virtual objects, especially for non-technical users, limiting their ability to participate in creating virtual worlds.
By receiving natural language commands describing the virtual environment, using the natural language command processor to determine the position and experience parts of the command, combined with the generative virtual environment builder, the sky box and 3D object model are generated based on this information to create a navigable 3D virtual environment.
It realizes the rapid generation of navigable 3D virtual environments, reducing the complexity and time-consuming of the design and creation process, and allowing non-technical users to participate in the creation of the virtual world.
Smart Images

Figure CN120225988A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to artificial reality and machine learning systems, in which natural language commands are used to automatically generate an artificial reality environment. Background Art
[0002] A user interacting with an extra reality (XR) device can view content in an artificial reality environment that includes real-world objects and / or two-dimensional (2D) virtual objects and / or three-dimensional (3D) virtual objects. For example, an XR environment can be a virtual environment depicted by a virtual reality (VR) device that shows a set of virtual objects. As another example, an XR environment can be a mixed reality environment that has real-world objects and virtual objects superimposed on top of these real-world objects. The user can view the objects in the artificial reality environment and modify the content in the artificial reality environment.
[0003] Although XR systems can provide an intuitive way to view objects in an XR environment, navigate those objects, and interact with those objects, the process of designing and creating new XR environments and / or the objects therein can be challenging and time-consuming. Typically, a creator needs to provide or access digital assets that define surface textures, build 3D objects with complex geometries and properties, and use construction tools (e.g., computer-aided design (CAD) modeling software, vector drawing software, etc.) either within the XR application and / or outside the XR application, and these construction tools can be expensive and difficult for non-technical users to learn. Thus, this typical XR world design and construction process can be too difficult for non-technical users, limiting the ability of many users to participate in creating their own virtual worlds. Summary of the Invention
[0004] According to a first aspect of the present disclosure, there is provided a method for generating a navigable 3D virtual environment, the method comprising: receiving a command in a concise language that describes the virtual environment; using a natural language command processor to determine (i) a location part and (ii) an experience part of the command; using a generative virtual environment builder including one or more first machine learning models to generate a skybox based on the location part of the command, wherein the skybox includes at least a shape and an image projected onto the shape; using a generative virtual environment builder including one or more second machine learning models to generate one or more 3D object models based on both the location part and the experience part of the command, wherein each 3D object model includes at least a geometry and a location within the navigable 3D virtual environment; creating a 3D virtual environment model that combines the skybox with the one or more 3D object models, wherein the 3D virtual environment model is navigable using an artificial reality (XR) device; and storing the 3D virtual environment model on a data storage device.
[0005] In some embodiments, determining the location part includes applying a natural language command processor to identify a specific geographical location or type of the environment.
[0006] In some embodiments, the method further includes determining an embedding and metadata associated with the command, wherein the generative virtual environment builder receives the embedding and metadata as part of its input to generate the skybox and the one or more 3D object models.
[0007] In some embodiments, generating the skybox includes identifying elements semantically related to the determined location part and adding representations of these elements to the skybox.
[0008] In some embodiments, determining (i) the location part and (ii) the experience part of the command includes: preprocessing the command by tokenizing and contextualizing phrases of the command using a first language model, the first language model being pre-trained to classify phrases according to location indicators or activity indicators.
[0009] In some embodiments, determining (i) the location part and (ii) the experience part of the command further includes: applying a second language model to the tokenized and contextualized phrases and receiving one or more embeddings from the second language model; wherein the one or more embeddings are provided to the one or more first machine learning models to generate the skybox, and the one or more embeddings are provided to the one or more second machine learning models to generate the one or more 3D object models.
[0010] In some embodiments, the one or more first machine learning models and / or the one or more second machine learning models: A) include a 2D modeling portion that is trained on a combination of images and / or videos that are labeled with metadata describing the content or context of the images and / or videos; and B) generate one or more 2D representations based on one or more input labels, and the one or more first machine learning models and / or the one or more second machine learning models include a generative portion that is trained to take the one or more 2D representations and generate one or more 3D representations.
[0011] In some embodiments, the determining location portion includes identifying a semantic identifier of the location portion; wherein, generating the skybox includes mapping the semantic identifier of the location portion into the latent space of the one or more second machine learning models of the generative virtual environment builder.
[0012] In some embodiments, the method further includes iteratively updating the 3D virtual environment model by: receiving a further natural language command; identifying the target of the further natural language command; and changing an aspect of the target based on the further natural language command.
[0013] In some embodiments, the method further includes: using additional training items created based on pairing the target with an aspect of the changed target to update the training of the one or more first machine learning models and / or the one or more second machine learning models.
[0014] In some embodiments, the one or more generated 3D object models are generated with an attribute that specifies whether each 3D object model in the one or more 3D object models is fixed or movable.
[0015] According to another aspect of the present disclosure, there is also provided a computer-readable storage medium storing a plurality of instructions that, when executed by a computing system, cause the computing system to perform a process for generating a navigable 3D virtual environment, the process comprising: receiving a command in a concise language describing the virtual environment; using a natural language command processor to determine (i) a location part and (ii) an experience part of the command; using a generative virtual environment builder including one or more first machine learning models to generate a skybox based on the location part of the command, wherein the skybox includes at least a shape and an image projected onto the shape; using a generative virtual environment builder including one or more second machine learning models to generate one or more 3D object models based on both the location part and the experience part of the command, wherein each 3D object model includes at least a geometry and a location within the navigable 3D virtual environment; creating a 3D virtual environment model that combines the skybox with the one or more 3D object models, wherein the 3D virtual environment model is navigable using an artificial reality (XR) device; and storing the 3D virtual environment model on a data storage device.
[0016] In some embodiments, the process further includes determining an embedding and metadata associated with the command, wherein the generative virtual environment builder receives the embedding and metadata as part of its input to generate the skybox and the one or more 3D object models.
[0017] In some embodiments, generating the skybox includes identifying elements semantically related to the determined location part and adding representations of these elements to the skybox.
[0018] In some embodiments, determining (i) the location part and (ii) the experience part of the command includes: preprocessing the command by tokenizing and contextualizing phrases of the command using a first language model pre-trained to classify phrases according to location indicators or activity indicators.
[0019] In some embodiments, determining (i) the location part and (ii) the experience part of the command further includes: applying a second language model to the tokenized and contextualized phrases and receiving one or more embeddings from the second language model; wherein the one or more embeddings are provided to the one or more first machine learning models to generate the skybox, and the one or more embeddings are provided to the one or more second machine learning models to generate the one or more 3D object models.
[0020] In some embodiments, the one or more first machine learning models and / or the one or more second machine learning models: A) include a 2D modeling portion that is trained on a combination of images and / or videos, the images and / or videos being labeled with metadata that describes the content or context of the images and / or videos; and B) generate one or more 2D representations based on one or more input labels, and the one or more first machine learning models and / or the one or more second machine learning models include a generative portion that is trained to take the one or more 2D representations and generate one or more 3D representations.
[0021] According to another aspect of the present disclosure, there is provided a computing system for generating a navigable 3D virtual environment, the computer system comprising: one or more processors; and one or more memories storing a plurality of instructions that, when executed by the one or more processors, cause the computing system to perform a process that includes: receiving a command in a concise language that describes the virtual environment; using a natural language command processor to determine (i) a location portion and (ii) an experience portion of the command; using a generative virtual environment builder that includes one or more first machine learning models to generate a skybox based on the location portion of the command, where the skybox includes at least a shape and an image projected onto the shape; using a generative virtual environment builder that includes one or more second machine learning models to generate one or more 3D object models based on both the location portion and the experience portion of the command, where each 3D object model includes at least a geometry and a location within the navigable 3D virtual environment; creating a 3D virtual environment model that combines the skybox with the one or more 3D object models, where the 3D virtual environment model is navigable using an artificial reality (XR) device; and storing the 3D virtual environment model on a data storage device.
[0022] In some embodiments, determining the location portion includes identifying a semantic identifier of the location portion; wherein, generating the skybox includes mapping the semantic identifier of the location portion into the latent space of the one or more second machine learning models of the generative virtual environment builder.
[0023] In some embodiments, the process further includes iteratively updating the 3D virtual environment model by: receiving a further natural language command; identifying a target of the further natural language command; changing an aspect of the target based on the further natural language command; and using an additional training item created by pairing the target with the changed aspect of the target to update the training of the one or more first machine learning models and / or the one or more second machine learning models.
[0024] It will be understood that any feature described herein as being suitable for incorporation into one or more aspects or embodiments of the present disclosure is intended to be generalizable to any and all aspects and embodiments of the present disclosure. Those skilled in the art can understand other aspects of the present disclosure based on the description, claims, and drawings of the present disclosure. The above general description and the following detailed description are merely exemplary and illustrative, and not restrictive of the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a conceptual diagram showing an example output of a VR world generator based on a concise language description of a club.
[0026] Figure 2 is a conceptual diagram showing an example output of a VR world generator based on a concise language description of an automobile display.
[0027] Figure 3 is a block diagram showing an example processing pipeline for generating VR world data based on the spoken concise language commands.
[0028] Figure 4 is a flowchart showing a process for generating a navigable virtual environment based on one or more concise language commands in some embodiments.
[0029] Figure 5 is a block diagram showing an overview of multiple devices on which some embodiments of the present technology can operate.
[0030] Figure 6A is a wireframe diagram showing a virtual reality headset that can be used in some embodiments of the present technology.
[0031] Figure 6B is a wireframe diagram showing a mixed reality headset that can be used in some embodiments of the present technology.
[0032] Figure 6C is a wireframe diagram showing a controller, and in some embodiments, a user can hold the controller with one or both hands to interact with the artificial reality environment.
[0033] Figure 7 is a block diagram showing an overview of an environment in which some embodiments of the present technology can operate.
[0034] Figure 8 is a block diagram showing components that can be used in a system adopting the disclosed technology in some embodiments. DETAILED DESCRIPTION
[0035] The latest advancements in artificial intelligence (AI) have enabled new applications for analyzing natural language and detecting image and video content. Natural language processing (NLP) techniques can parse and / or perform lexical analysis to tokenize words in a sentence or paragraph, generate word embeddings, and model syntactic and semantic relationships between these tokens. Regarding image analysis, convolutional neural networks (CNNs) can be trained to identify one or more objects and / or multiple instances of the same object from image data, enabling the classification of these images and videos semantically (and in some cases, spatially). With the emergence of transformers and the development of "self-attention," the relationships between individual objects in a scene can now be numerically correlated to develop a more contextually aware understanding of what is depicted in the image data.
[0036] Many deep learning models for analyzing image data are initially developed by repeatedly adjusting the weights of the model against training data to accurately predict the existing training data labels or create a statistically significant separation between unlabeled classes in a variable space. In effect, the learned weights of the hidden layers of the model provide a latent space representation between different semantic similarity representations of the object. For example, a deep learning model trained to classify images of dogs with the label "dog" can effectively convert the spatial image data into language, and this conversion is enabled by the latent space representation that correlates the spatial representation of the dog's image with the linguistic representation of the dog. This latent space representation of different forms of information enables a neural network architecture that can classify image data and / or its content (e.g., an encoder), and conversely, can generate image data that probabilistically depicts one or more labels or words (e.g., autoregressive image generation).
[0037] Aspects of the present disclosure combine NLP with generative or autoregressive models to convert plain language words or commands into automatically generated 3D environments. In an example embodiment, a user speaks a command that describes in plain language the environment to be generated, which is captured by a microphone and converted into text data using a text-to-speech model. A natural language command processor can then parse the text data to extract at least one location (e.g., a specific location, environment type, etc.) and activity described in the user's command. A generative virtual environment builder can receive the location and activity (as well as embeddings, metadata, etc.) as inputs and generate a virtual environment (e.g., a skybox) and / or 3D objects within the virtual environment. Generally, the natural language command processor and the generative virtual environment builder can be combined to form a VR world generator that can generate an interactive 3D virtual environment and / or objects within the 3D virtual environment based on a plain language description of a location and / or activity.
[0038] For example, if the user command is "scuba diving in Hawaii", the natural language command processor can identify the location "Hawaii" (which can be semantically related to an island, tropical environment, etc.), and the activity "scuba diving" (which can be semantically related to an ocean, coastline, coral reef, etc.). The generative virtual environment builder can then generate a skybox depicting a distant bay, coastline, island, ocean, and sky, as well as objects in the environment associated with scuba diving (e.g., coral, fish, other marine life, etc.) that are located in the environment in a similar position to where these objects would be in the real world (e.g., underwater). The skybox and objects can form a navigable 3D virtual environment that the user can explore to view the results of the generative world building process.
[0039] In some implementations, the natural language command processor can perform preprocessing operations, where the natural language command processor first tokenizes the words and / or otherwise contextualizes the words based on a pre-trained language model. For example, the language model can be trained to classify words or phrases as locations, activities, or experiences or others, which is initially used to extract relevant parts of the user's spoken command. Another language model (or the same language model) can then infer one or more embeddings for each location word and for each activity word to form a semantic understanding of the words spoken. The natural language command processor can provide the words, tokens, embeddings, or some combination thereof as inputs to the generative virtual environment builder. In other embodiments, the natural language command processor can be integrated into the model architecture of the generative virtual environment builder (i.e., the VR world generator, rather than a cascade of two or more separate models).
[0040] In various embodiments, one or more models forming the VR world generator can be trained on some combination of images and / or videos that are tagged with metadata that describes the content or context of those images and / or videos. For example, the images and / or videos can be captured as a user performs various activities in different locations, and the images and / or videos can be geotagged and / or classified using machine learning to associate the images and / or videos with labels that describe the content of the images and / or videos. Additional training data can be generated (either manually or automatically) that converts the images and / or videos (e.g., using a generative network) into a 3D environment or object, where the fidelity of the 2D-to-3D conversion is measured and adjusted during training (e.g., using a discriminator network). In other words, a combination of manual and automatic tagging can be used to create tagged images and / or videos, supervised learning can be performed on these tagged images and / or videos to associate words or phrases with graphical representations of the labels (and, in some cases, graphical representations of combinations of labels using transformers), and then a generative adversarial network (GAN) can be trained to produce 3D representations of 2D objects and environments.
[0041] Regarding the example of "scuba diving in Hawaii" described above, the model can be trained on images and / or videos (e.g., captured by a smartphone, action camera, etc.) in which people are swimming, snorkeling, and scuba diving (i.e., semantically similar activities) along a tropical beach, coral reef, or other coastal waters (i.e., semantically similar locations). When the natural language command processor receives a command, the natural language command processor can parse the command to semantically identify the location (Hawaii and / or tropical island locations) and the experience (scuba diving or other coastal water-based activities). Mapping these semantic identifiers in the latent space of the generative virtual environment builder, the generative virtual environment builder can generate a skybox (e.g., a 2D image mapped to the inner surface of a 3D object such as a cuboid or ellipsoid). Additionally, the generative virtual environment builder can generate objects and / or structures within the virtual environment, such as beaches, oceans, corals, fish, marine life, and / or other things commonly found in the coastal waters captured in the training data.
[0042] After generating an initial virtual world, a VR world generator can be used to iteratively add objects to or remove objects from the generative world, and / or to modify aspects of the generated skybox and / or the initially generated objects. For example, a user can examine the generated virtual environment to evaluate whether the VR world generator has created an accurate representation of the virtual environment that the user intended to create. If aspects of the virtual environment are inaccurately generated, or if the user wishes to otherwise change aspects of the generated virtual environment, the user can issue subsequent natural language commands to remove objects, add new objects, change details about various objects, and / or change details about the skybox. For example, if the VR world generator creates a virtual environment in the coastal waters of a tropical island with a coral reef in response to the command "scuba diving in Hawaii", the user may wish to update the type of coral reef (e.g., "add staghorn coral", "increase the size of the coral reef", etc.), add or remove fish or other marine life (e.g., "add starfish", "remove eels", etc.), and / or more generally change aspects of the environment (e.g., "increase the depth", "add seaweed plants", etc.). In this way, the user can issue secondary concise language commands to adjust the environment to achieve a specific environment or aesthetic. These secondary concise language commands and adjustments retained by the user can be used as additional training data for one or more VR world generator models, such that the one or more models can better infer the user's intent from the initial concise language commands over time.
[0043] As described herein, the terms "concise language" and "natural language" can be used interchangeably to refer to a language that does not have a formally defined syntactic or semantic structure (such as a programming language). Thus, a "concise language" or "natural language" command can generally refer to a phrase or sentence that does not necessarily follow a strict set of rules and furthermore does not have to be interpreted or translated by a special application (such as a compiler). Additionally, the terms "command" and "description" can be used interchangeably to refer to the substance of a concise language string or a natural language string.
[0044] As described herein, the terms "navigable virtual environment", "interactive 3D virtual environment", and "VR world data" may be used interchangeably to refer to data for rendering 2D assets and / or 3D assets in a 3D virtual environment that can be explored using an avatar or a virtual camera of a game engine or a VR engine. VR world data differs from the generated image data in that VR world data can map 2D images onto 3D surfaces (e.g., in skyboxes, surface textures, etc.), and / or the VR world data can be a geometric model defining the shape of an object in 3D space. The format of VR world data may vary in different embodiments depending on the game engine used to render the generated skybox and / or the generated 3D objects.
[0045] Figure 1 FIG. 100 is a conceptual diagram showing an example VR world generator output based on a plain language description of a club. In this example, the user speaks the command "playing pool in a sports-themed clubhouse" 102, and provides this command as an input (e.g., as audio data, as text data determined using a text-to-speech model, etc.) to the VR world generator 104. The VR world generator 104 processes this command to identify the location as "sports-themed clubhouse" and the experience as "playing pool". In this example, the natural language command processor and / or the VR world generator may use an attention-based model (such as a transformer) to infer the semantic context, where the word "pool" here refers to the sport of billiards, rather than other possible meanings (e.g., other homonyms such as "swimming pool", resource set, etc.). Additionally, the natural language processor may further perform tokenization or other lexical analysis to determine the location as "clubhouse", which is qualified by the adjective "sports-themed" to specify a particular aesthetic or decorative style of the location.
[0046] Then, the VR world generator 104 generates a 3D environment 106 based on the command 102. As Figure 1 shown, the VR world generator 104 can generate a room with wall decorations (e.g., sports equipment, pictures, or other wall-mounted decorations typical of a club) and furniture (e.g., a billiards table, a sofa, a trash can, etc.). In this example, the generated "skybox" includes the geometry and dimensions of the room, as well as any wall textures. In other words, the term "skybox" as used herein may refer to a graphic projected onto a 3D surface - the 3D surface may or may not be reachable depending on the particular situation.
[0047] In addition, each object generated by the VR world generator 104 is associated with an absolute or relative position and orientation within the environment, such that a game engine or the like can render the object within a specific location in the 3D environment. In some embodiments, the VR world generator 104 can generate properties of the objects within the environment, such as whether these objects are fixed or movable, whether these objects are collidable, and / or any behaviors or interactions associated with these objects. For example, with respect to Figure 1 the example shown in
[0048] Figure 2
[0049] Figure 2 a pool table can be fixed and collidable such that the user's avatar cannot move through the pool table, any billiard balls can be collidable and movable, and a billiard cue can be movable and capable of being picked up and swung by the user's player character. In some cases, video training data (e.g., video of a person holding or using a billiard cue) can be used to train a model to identify stationary objects (e.g., a pool table and a couch), as well as movable or interactive objects. These properties of the generated objects can be automatically initialized by the VR world generator 104, and / or manually configured by the user after the virtual environment is generated. Figure 2As shown in [description], the VR world generator 204 can generate a large room that has a support structure (e.g., load-bearing columns), a display stage on which different sports cars are placed, and other furniture (e.g., exhibition stands, tables, etc.) that are typically present at an auto show. As in the previous example, the generated skybox can include the geometry and dimensions of the room and any wall textures - all of these items can be reachable, or only a portion of these items can be reachable.
[0050] In some cases, the VR world generator 204 can generate avatars that represent non-playable characters (NPCs) or act as placeholders for player characters (e.g., potential avatars of other users). In the case where the avatar is an NPC, the VR world generator 204 can define default behaviors for the NPC (e.g., walking, staying, periodically interacting with other NPCs, etc.), which can be modified by the user after the virtual environment is generated.
[0051] Figure 3 FIG. [figure number] is a block diagram showing an example processing pipeline 300 for generating VR world data based on the spoken plain language commands. The pipeline 300 initially receives an audio sample 302 and provides the audio sample as input to a speech recognizer 304. The speech recognizer 304 can perform speech-to-text operations using a machine learning model, etc., and generates a transcription of any words spoken in the audio sample 302. In some embodiments, a tokenizer or lexical analyzer 306 can receive the transcription (e.g., text data representing the spoken command) to parse the command into different syntactic units and / or perform semantic analysis on the location and experience to attach metadata to the command to improve the accuracy and / or robustness of the subsequent operations of the pipeline 300.
[0052] Then, the command and / or the output of the tokenizer or lexical analyzer 306 is provided to a VR world generator 308, which includes a natural language command processor 310 and a generative virtual environment builder 312. As described above, the natural language command processor 310 can identify the location and experience from the command (e.g., using a pre-trained language model), which are used as inputs to the generative virtual environment builder 312. The generative virtual environment builder 312 can use a generative model (e.g., a generative adversarial network (GAN)) to generate a virtual environment that includes a skybox and / or one or more 3D objects (and the corresponding positions and orientations of the one or more 3D objects in the environment). The output of the generative virtual environment builder 312 can be a data payload that combines the skybox data, 3D object data, and / or 3D object metadata into VR world data 314, which can be jointly used to generate a navigable 3D environment.
[0053] In some cases, a combination of automated data tagging, data augmentation, and / or other automated processing can be used to enrich image and video data to improve the robustness of the generative model of the generative virtual environment builder 312. For example, an image segmentation model can be used to delineate pixels associated with different classes of objects and / or distinguish different instances of the same object. A subset of the pixels associated with a particular instance of an object can be used to train the generative model to correlate the shape, size, and geometry of a given type of object. In some embodiments, a depth estimation model can be used to infer the 3D geometry of a surface or object from a 2D image (e.g., infer a depth map from an image or video frame), where the depth map serves as further training data to improve the fidelity of the generated 3D objects. The training data can also include stereoscopic video and / or 3D video (e.g., spatial data captured from a time-of-flight sensor, etc.) to further improve the accuracy of the generated 3D object geometry.
[0054] Overall, the pipeline 300 can make full use of one or more deep learning models to enable a user to naturally describe a 3D environment (i.e., without memorizing a specialized command set) and generate a 3D environment based on those descriptions. The advantage of the VR world generator 308 is that it enables non-technical users to meaningfully participate in the VR world building process, thus promoting more equitable participation in the creation of VR worlds and, more broadly, virtual reality.
[0055] Figure 4 is a flowchart showing a process 400 for generating a navigable virtual environment based on one or more plain language commands in some embodiments. In some embodiments, the process 400 can be executed in response to a user interacting with a virtual button, selecting a menu option, speaking a keyword or wake word, and / or manually triggered in some other way by the user's desire to generate a new virtual environment. The process 400 can be performed by, for example Figure 5The VR world generator 564 executes over a network on a server or a combination of multiple servers, or locally on a user system or device (i.e., an XR device (e.g., a head-mounted display) or a 2D device (such as a computer or other display or processing device)). In some embodiments, some steps of process 400 may be executed on a server, while other steps of process 400 may be executed locally on a user device. Although shown as having only one iteration, process 400 may be executed multiple times, repeatedly, iteratively, continuously, concurrently, in parallel, etc. in response to requests for generating 3D environments and / or modifying aspects of these 3D environments. Some steps of process 400 may be described as "generating" a skybox or 3D object model, which may involve providing input data to one or more pre-trained models and capturing the output from the one or more models.
[0056] At block 402, process 400 may receive a first plain language command. As described above, a plain language command may be a typed transcription or a text-to-speech transcription of a description of a 3D virtual environment from a user.
[0057] At block 404, process 400 may determine at least a location and an experience from the first plain language command. In some embodiments, a natural language command processor may perform some combination of tokenization, lexical analysis, and / or otherwise apply a language model to parse the plain language command to identify the location (and / or semantic category of the location) and the experience or activity (and / or semantic category of the experience or activity). In some cases, the location may be a specific place (which may fit a category of locations), or may be a category of locations (e.g., an island, a mountain, a café, a restaurant, etc.).
[0058] The experience or activity may be semantically related to a specific time of day, a season, the decor in the environment, or the objects expected to be in the environment. For example, the experience of "skiing" in the mountains may imply the presence of snow, which in itself means the season is winter. Conversely, the experience of "hiking" in the mountains may imply summer or fall, when there is little snow in the mountains.
[0059] At block 406, process 400 may generate a skybox based on at least one of the location and the experience. The skybox may be represented as data that defines the dimensions, geometry, texture, and / or images projected onto one or more surfaces of the skybox, all of which may be packed into a data structure interpretable by a game engine to render the skybox in real time. The images projected onto the surfaces of the skybox may initially be generated by a generative virtual environment builder (e.g., a GAN), which are then resized, reshaped, or otherwise projected onto the inner surfaces of the skybox geometry.
[0060] At block 408, process 400 can generate one or more 3D object models based on at least one of location and experience. Each 3D object (e.g., an object identified or labeled in training data) can be associated with a location and / or an experience. In some cases, the object can be a fixed structure forming a terrain shape, a fixed object that a user cannot move when deploying the 3D environment, a movable and / or interactive object within the environment, and / or an autonomously moving object (e.g., an animal, an NPC, a weather element, etc.). Similar to a skybox, each 3D object can be represented as data that defines the 3D object's size, geometry, texture, kinematics, properties, location, orientation, and / or an image projected onto one or more surfaces of the 3D object, and some or all of this data can be packed into a data structure readable by a game engine to render the 3D object in a virtual environment.
[0061] In some embodiments, the user can iteratively modify, remove, or add the generated skybox and / or one or more 3D object models via subsequent plain language commands. At block 410, process 400 can receive a second plain language command. The second plain language command can be semantically related to the first plain language command at least to the extent that the second plain language command is used to qualify or modify the VR environment generated according to the first plain language command. At block 412, process 400 can modify at least one of the skybox and the one or more 3D object models based on the second plain language command. For example, after generating a 3D environment related to tropical scuba diving, a second plain language command (such as "add a starfish", "shallow the coral reef", or "remove the clownfish") can cause the VR world generator to add a 3D object, remove a 3D object, and / or adjust the terrain or skybox.
[0062] At block 414, process 400 can store the skybox and the one or more 3D object models as a navigable virtual environment. For example, the generated skybox and the one or more 3D object models can be stored as data on a server, and a user can subsequently retrieve the data, and a game engine running on these users' devices interprets the data to instantiate a navigable 3D virtual environment. In some cases, the creator of the 3D virtual environment or other creators can access the stored virtual environment in a separate part of the VR application or in a completely separate application to manually modify the virtual environment and / or build on the virtual environment. In this way, the VR world generator can be used to quickly generate an initial environment, and then the creator can use other VR world building tools to manually adjust subsequent aspects of the environment.
[0063] Figure 5is a block diagram showing an overview of a plurality of devices on which some embodiments of the disclosed technology may operate. These devices may include the hardware components of device 500 that generate a navigable 3D virtual environment. In various embodiments, computing system 500 may include a single computing device 503 or multiple computing devices (e.g., computing device 501, computing device 502, and computing device 503) that communicate via a wired or wireless channel to distribute processing and share input data. In some embodiments, computing system 500 may include a stand-alone head-mounted viewer that can provide a computer-created or enhanced experience for a user without external processing or external sensors. In other embodiments, computing system 500 may include multiple computing devices, such as a head-mounted viewer and a core processing component (such as a console, a mobile device, or a server system), where some processing operations are performed on the head-mounted viewer and other processing operations are offloaded to the core processing component. The following describes example head-mounted viewers in conjunction with Figure 2 A and Figure 2 B. In some embodiments, position data and environmental data may be collected only by sensors incorporated in the head-mounted viewer device, while in other embodiments, one or more non-head-mounted viewer computing devices among a plurality of non-head-mounted viewer computing devices may include sensor components capable of tracking environmental data or position data.
[0064] Computing system 500 may include one or more processors 510 (e.g., a central processing unit (CPU), a graphical processing unit (GPU), a holographic processing unit (HPU), etc.). Processor 510 may be a single processing unit or multiple processing units that are located in one device or distributed across multiple devices (e.g., distributed across two or more of computing devices 501 to 503).
[0065] Device 500 may include one or more input devices 520 that provide input to one or more processors 510 (e.g., one or more CPUs, one or more GPUs, one or more HPUs, etc.), thereby notifying the one or more processors of an action. These actions may be relayed by a hardware controller that interprets signals received from the input devices and conveys information to the processor 510 using a communication protocol. For example, the input devices 520 include a mouse, a keyboard, a touch screen, an infrared sensor, a touchpad, wearable input devices (e.g., haptic gloves, bracelets, rings, earrings, necklaces, watches, etc.), camera- or image-based input devices, a microphone, or other user input devices.
[0066] The processor 510 may be coupled to other hardware devices, for example, by using a bus such as a Peripheral Component Interconnect Standard (PCI) bus or a Small Computer System Interface (SCSI) bus. The processor 510 may communicate with the hardware controller of a device such as a display 530. The display 530 may be used to display text and graphics. In some embodiments, the display 530 provides visual feedback of graphics and text to the user. In some embodiments, the display 530 includes an input device, such as when the input device is a touch screen or is equipped with an eye movement direction monitoring system, the input device is part of the display. In some embodiments, the display is separate from the input device. Examples of display devices include: liquid crystal display (LCD) screens, light emitting diode (LED) screens, projection displays, holographic displays, or augmented reality displays (such as, head-up display devices or head-mounted devices), etc. Other input / output (I / O) devices 540 may also be coupled to the processor, and these I / O devices are, for example, network cards, video cards, audio cards, Universal Serial Bus (USB), FireWire, or other external devices, cameras, printers, speakers, compact disc read-only memory (CD-ROM) drives, digital video disc (DVD) drives, disk drives, etc.
[0067] In some embodiments, the computing system 500 may use inputs from I / O devices 540 (such as cameras, depth sensors, inertial motion unit (IMU) sensors, global positioning system (GPS) units, lidar (LiDAR), or other time-of-flight sensors, etc.) to identify and map the user's physical environment while tracking the user's position within that environment. A simultaneous localization and mapping (SLAM) system may generate a map (e.g., topological, grid, etc.) of an area (which may be a room, building, outdoor space, etc.) and / or obtain a map previously generated by the computing system 500 or another computing system that has mapped the area. The SLAM system may track the user within the area based on factors such as GPS data, match the identified objects and structures with the mapped objects and structures, monitor accelerations and other position changes, etc.
[0068] In some embodiments, the device 500 further includes a communication device that is capable of communicating wirelessly or wire-based with a network node. The communication device may communicate with another device or server via a network using, for example, Transmission Control Protocol / Internet Protocol (TCP / IP). The device 500 may utilize the communication device to distribute operations across multiple network devices.
[0069] The processor 510 may access the memory 550, which may be included on one of the multiple computing devices of the computing system 500 or may be distributed across the multiple computing devices of the computing system 500 or other external devices. The memory includes one or more of various hardware devices for volatile and non-volatile storage and may include both read-only memory and writable memory. For example, the memory may include random access memory (RAM), various caches, CPU registers, read-only memory (ROM), and writable non-volatile memory such as flash memory, hard disk drives, floppy disks, compact discs (CDs), DVDs, magnetic storage devices, and tape drives, etc. The memory is not a propagated signal separate from the underlying hardware; thus, the memory is non-transitory. The memory 550 may include a program memory 560 that stores programs and software such as an operating system 562, a VR world generator 564, and other applications 566. The memory 550 may also include a data memory 570, such as image data, video data, image and video data tags, geographical location data, language models, image classification models, object detection models, image segmentation models, configuration data, settings, user options or preferences, etc., which may be provided to the program memory 560 or any element of the device 500.
[0070] Some embodiments may operate in conjunction with many other computing system environments or configurations. Examples of computing systems, environments, and / or configurations suitable for use with the technology include, but are not limited to, personal computers, server computers, handheld or laptop devices, cellular telephones, wearable electronic devices, gaming consoles, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers (PCs), minicomputers, mainframe computers, or distributed computing environments including any of the above systems or devices, etc.
[0071] Figure 6A is a line diagram of a virtual reality head-mounted display (HMD) 600 according to some embodiments. The HMD 600 includes a front rigid body 605 and a strap 610. The front rigid body 605 includes one or more electronic display elements of an electronic display 645, an inertial motion unit (IMU) 615, one or more position sensors 620, a locator 625, and one or more computing units 630. The position sensors 620, the IMU 615, and the computing units 630 may be located inside the HMD 600 and may be invisible to the user. In various embodiments, the IMU 615, the position sensors 620, and the locator 625 may track the movement and position of the HMD 600 in the real world and in an artificial reality environment in three degrees of freedom (3DoF) or six degrees of freedom (6DoF). For example, the locator 625 may emit infrared light beams that produce light spots on real objects around the HMD 600. As another example, the IMU 615 may include, for example: one or more accelerometers; one or more gyroscopes; one or more magnetometers; one or more other non-camera-based sensors of position, force, or orientation; or combinations thereof. One or more cameras (not shown) integrated with the HMD 600 may detect the light spots. The computing unit 630 in the HMD 600 may use the detected light spots to infer the position and movement of the HMD 600, as well as to identify the shape and position of real objects around the HMD 600.
[0072] The electronic display 645 may be integrated with the front rigid body 605 and may provide image light to the user according to the instructions of the computing unit 630. In various embodiments, the electronic display 645 may be a single electronic display or multiple electronic displays (e.g., one display for each user eye). Examples of the electronic display 645 include: a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, an active-matrix organic light-emitting diode (AMOLED) display, a display including one or more quantum dot light-emitting diode (QOLED) sub-pixels, a projector unit (e.g., a micro LED, a LASER, etc.), some other display, or some combination thereof.
[0073] In some embodiments, the HMD 600 may be coupled to a core processing component, such as a personal computer (PC) (not shown) and / or one or more external sensors (not shown). The external sensors may monitor the HMD 600 (e.g., via light emitted from the HMD 600), and the PC may use the HMD in combination with the outputs from the IMU 615 and the position sensor 620 to determine the position and movement of the HMD 600.
[0074] Figure 6B is a line diagram of a mixed reality HMD system 650, which includes a mixed reality HMD 652 and a core processing component 654. The mixed reality HMD 652 and the core processing component 654 may communicate via a wireless connection (e.g., a 60 GHz link) as indicated by the link 656. In other embodiments, the mixed reality system 650 includes only a head-mounted viewer without an external computing device, or includes other wired or wireless connections between the mixed reality HMD 652 and the core processing component 654. The mixed reality HMD 652 includes a see-through display 658 and a frame 660. The frame 660 may house various electronic components (not shown), such as a light projector (e.g., a LASER, an LED, etc.), a camera, an eye-tracking sensor, a microelectromechanical system (MEMS) component, a networking component, etc.
[0075] The projector can be coupled to the see-through display 658, for example, via optical elements to display media to the user. The optical elements can include one or more waveguide assemblies, one or more reflectors, one or more lenses, one or more mirrors, one or more collimators, one or more gratings, etc. for guiding light from the projector to the user's eyes. Image data can be transmitted from the core processing component 654 to the HMD 652 via the link 656. The controller in the HMD 652 can convert the image data into a plurality of light pulses from the projector, and these light pulses can be transmitted to the user's eyes as output light via the optical elements. The output light can be mixed with the light passing through the display 658, thereby allowing the output light to present the following virtual objects: these virtual objects appear as if they exist in the real world.
[0076] Similar to the HMD 600, the HMD system 650 can also include a motion and position tracking unit, a camera, a light source, etc. The above-mentioned motion and position tracking unit, camera, light source, etc. allow the HMD system 650 to track itself, for example, in 3DoF or 6DoF, track multiple parts of the user (such as hands, feet, head, or other body parts), render virtual objects to appear stationary as the HMD 652 moves, and make virtual objects respond to poses and other real-world objects.
[0077] Figure 6C A plurality of controllers 670 (including controllers 676A and 676B) are shown. In some embodiments, the user can hold the plurality of controllers with one or both hands to interact with the artificial reality environment presented by the HMD 600 and / or the HMD 650. The controller 670 can communicate with the HMD directly or via an external device (such as the core processing component 654). The controller can have its own IMU unit, position sensor, and / or can emit additional light points. The sensors in the HMD 600 or 650, external sensors, or controllers can track these controller light points to determine the position and / or orientation of the controller (for example, track the controller in 3DoF or 6DoF). The computing unit 630 in the HMD 600 or the core processing component 654 can combine the IMU output and the position output and use this tracking to monitor the user's hand position and movement. The controller can also include various buttons (such as buttons 672A to 672F) and / or joysticks (such as joysticks 674A and 674B), and the user can actuate these buttons and / or joysticks to provide input and interact with objects.
[0078] In various embodiments, the HMD 600 or 650 may also include additional subsystems for monitoring indications of user interaction and intent, such as an eye tracking unit, an audio system, various network components, and the like. For example, in some embodiments, one or more cameras included in the HMD 600 or 650 instead of or in addition to the controller, or one or more cameras from a plurality of external cameras, may monitor the position and pose of the user's hand to determine the pose and other hand and body movements. As another example, one or more light sources may illuminate one or both of the user's eyes, and the HMD 600 or 650 may use an eye-facing camera to capture the reflection of the light to determine the eye position, model the user's eyes, and determine the gaze direction (e.g., based on a set of reflected light around the user's cornea).
[0079] Figure 7 FIG. is a block diagram showing an overview of an environment 700 in which some embodiments of the disclosed technology may operate. The environment 700 may include one or more client computing devices 705A through 705D, examples of which may include the device 500. The client computing device 705 may operate in a network environment using a logical connection through the network 730 to one or more remote computers, such as server computing devices.
[0080] In some embodiments, the server 710 may be an edge server that receives client requests and coordinates the fulfillment of those requests through other servers, such as servers 720A through 720C. The server computing devices 710 and 720 may include computing systems such as the device 500. Although each server computing device 710 and 720 is logically shown as a single server, multiple server computing devices may each be a distributed computing environment that includes multiple computing devices located at the same physical location or geographically distinct physical locations. In some embodiments, each server 720 corresponds to a group of servers.
[0081] The client computing device 705, and the server computing devices 710 and 720 can each act as a server or a client for one or more other server / client devices. Server 710 can be connected to database 715. Servers 720A through 720C can each be connected to a respective database 725A through 725C. As discussed above, each server 720 can correspond to a group of servers, and each of these servers can share a database or can have its own database. Databases 715 and 725 can warehouse (e.g., store) information. Although databases 715 and 725 are logically shown as a single unit, databases 715 and 725 can each be a distributed computing environment that includes multiple computing devices, can be located within their corresponding servers, or can be located at the same physical location or at geographically distinct physical locations.
[0082] Network 730 can be a local area network (LAN) or a wide area network (WAN), but can also be some other wired or wireless network. Network 730 can be the Internet or some other public or private network. The client computing device 705 can be connected to network 730 through a network interface (such as through wired communication or wireless communication). Although the connections between server 710 and the multiple servers 720 are shown as separate connections, these connections can be any type of local area network, wide area network, wired network, or wireless network that includes network 730 or a separate public or private network.
[0083] Figure 8FIG. 0 is a block diagram showing components 800 that may be used in a system employing the disclosed techniques. Components 800 include hardware 802, general software 820, and specialized components 840. As described above, systems implementing the disclosed techniques may use a variety of hardware, including a processing unit 804 (e.g., CPU, GPU, attached processing unit (APU), etc.), working memory 806, storage memory 808 (local storage or as an interface to remote storage, such as memory 715 or 725), and input and output devices 810. In various embodiments, storage memory 808 may be one or more of the following: a local device, an interface to a remote storage device, or a combination thereof. For example, storage memory 808 may be a collection of one or more hard disk drives accessible via a system bus (e.g., redundant array of independent disk (RAID)), or may be a cloud storage provider or other network memory accessible via one or more communication networks (e.g., network accessible storage (NAS) device, such as memory 215 or memory provided by another server 720). Components 800 may be implemented in a client computing device (e.g., client computing device 705) or on a server computing device (e.g., server computing device 710 or 720).
[0084] General software 820 may include various applications, including an operating system 822, native programs 824, and a basic input output system (BIOS) 826. Specialized components 840 may be sub-components of general software applications 820 (e.g., native programs 824). Specialized components 840 may include a speech recognizer 858, a tokenizer / lexical analyzer 860, a natural language command processor 862, a generative virtual environment builder 864, and components (e.g., interface 842) that may be used to provide a user interface, transfer data, and control the specialized components. In some embodiments, components 800 may be located in a computing system distributed across multiple computing devices, or may be an interface to a server-based application that executes one or more of the specialized components 840.
[0085] The speech recognizer 858 may perform speech-to-text operations using a machine learning model, etc., and generate a transcription of the input speech audio. Additional details regarding language recognition are provided above with reference to, for example Figure 3 box 304 and Figure 4 box 404.
[0086] The tokenizer / lexical analyzer 860 can receive the speech text and parse the speech text into different syntactic units to perform semantic analysis for location and experience determination. Above with reference to, for example Figure 3 box 306 of Figure 4 and box 404 of
[0087] provide additional details regarding tokenization and lexicographic analysis of the text. Figure 3 box 310 of Figure 4 and box 404 of
[0088] provide additional details regarding identifying location and experience from the input language / tokens. Figure 3 box 312 of Figure 4 and box 406 and box 408 of
[0089] provide additional details regarding generating the skybox and / or one or more 3D objects.Embodiments of the disclosed technology may include or be implemented in conjunction with an artificial reality system. Artificial reality or extended reality (XR) is a form of reality that has been adjusted in some manner before being presented to a user. Artificial reality or extended reality may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., a photograph of the real world). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or multiple channels (e.g., stereoscopic video that produces a three-dimensional effect for a viewer). Additionally, in some embodiments, artificial reality may also be associated with an application, product, accessory, service, or some combination thereof, such as for creating content in artificial reality and / or using in artificial reality (e.g., performing an activity in artificial reality). An artificial reality system that provides artificial reality content may be implemented on various platforms, including a head-mounted display (HMD) connected to a host computer system, a stand-alone HMD, a mobile device or computing system, a "cave" environment or other projection system, or any other hardware platform capable of providing artificial reality content to one or more viewers.
[0090] As used herein, "virtual reality" or "VR" refers to an immersive experience in which a user's visual input is controlled by a computing system. "Augmented reality" or "AR" refers to systems in which a user views real-world images after they have passed through a computing system. For example, a tablet computer with a camera on the back can capture real-world images, and then the tablet computer can display those images on a screen on the side of the tablet computer opposite the camera. The tablet computer can process and adjust or "augment" those images as they pass through the system, such as by adding virtual objects. "Mixed reality" or "MR" refers to systems in which the light entering a user's eyes is partially generated by a computing system and partially composed of light reflected from objects in the real world. For example, an MR headset can be shaped like a pair of glasses with a pass-through display that allows light from the real world to pass through a waveguide that simultaneously emits light from a projector in the MR headset, thereby allowing the MR headset to present virtual objects that are mixed with real objects visible to the user. As used herein, "artificial reality", "hyperreality", or "XR" refers to any one of VR, AR, MR, or any combination or mixture thereof.
[0091] Those skilled in the art will recognize that the components and blocks shown above can be changed in various ways. For example, the order of the logic can be rearranged, sub-steps can be performed in parallel, the logic shown can be omitted, other logic can be included, etc. As used herein, the word "or" refers to any possible permutation of a set of items. For example, the phrase "A, B, or C" refers to at least one of A, B, C, or any combination thereof, such as any one of the following items: A; B; C; A and B; A and C; B and C; A, B, and C; or multiple of any item, such as A and A; B, B, and C; A, A, B, C, and C; etc. If necessary, aspects can be modified to adopt the systems, functions, and concepts of the various references above to provide further embodiments. If statements or topics in the various references conflict with the statements or topics of this application, this application shall prevail.
Claims
1. A method for generating a navigable 3D virtual environment, the method comprising: Receiving a command in a concise language that describes the virtual environment; Using a natural language command processor to determine (i) a location part and (ii) an experience part of the command; Using a generative virtual environment builder including one or more first machine learning models to generate a skybox based on the location part of the command, wherein the skybox includes at least a shape and an image projected onto the shape; Using a generative virtual environment builder including one or more second machine learning models to generate one or more 3D object models based on both the location part and the experience part of the command, wherein each 3D object model includes at least a geometry and a location within the navigable 3D virtual environment; Creating a 3D virtual environment model that combines the skybox with the one or more 3D object models, wherein the 3D virtual environment model is navigable using an artificial reality (XR) device; and Storing the 3D virtual environment model on a data storage device.
2. The method according to claim 1, wherein Determining the location part includes applying the natural language command processor to identify a specific geographical location or type of the environment.
3. The method according to claim 1, i. The method further includes determining an embedding and metadata associated with the command, wherein, The generative virtual environment builder receives the embedding and the metadata as part of its input to generate the skybox and the one or more 3D object models; and / or preferably ii. wherein generating the skybox includes: identifying elements semantically related to the determined location part and adding a representation of the elements to the skybox.
4. The method according to claim 1, 2 or 3, wherein, Determining (i) the location part and (ii) the experience part of the command includes: Preprocessing the command by using a first language model to tokenize and contextualize phrases of the command, the first language model being pre-trained to classify phrases according to location indicators or activity indicators; and preferably, wherein determining (i) the location part and (ii) the experience part of the command further includes: Applying a second language model to the tokenized and contextualized phrases and receiving one or more embeddings from the second language model; wherein the one or more embeddings are provided to the one or more first machine learning models to generate the skybox, and the one or more embeddings are provided to the one or more second machine learning models to generate the one or more 3D object models.
5. The method according to any one of the preceding claims, wherein, A) The one or more first machine learning models and / or the one or more second machine learning models include a 2D modeling part that is trained on a combination of images and / or videos labeled with metadata describing the content or context of the images and / or videos; and B) The one or more first machine learning models and / or the one or more second machine learning models generate one or more 2D representations based on one or more input labels, and The one or more first machine learning models and / or the one or more second machine learning models include a generative portion that is trained to receive one or more 2D representations and produce one or more 3D representations.
6. The method according to any one of the preceding claims, Among them, determining the location portion includes identifying a semantic identifier of the location portion; and wherein generating the skybox includes mapping the semantic identifier of the location portion into a latent space of the one or more second machine learning models of the generative virtual environment builder.
7. The method according to any one of the preceding claims, the method further comprising iteratively updating the 3D virtual environment model by: receiving a further natural language command; identifying the target of the further natural language command; and changing an aspect of the target based on the further natural language command; and preferably, the method further comprises using additional training items created based on pairing the target with the changed aspect of the target to update the training of the one or more first machine learning models and / or the one or more second machine learning models.
8. The method according to any one of the preceding claims, wherein, The one or more 3D object models generated are generated with an attribute that specifies whether each 3D object model in the one or more 3D object models is fixed or movable.
9. A computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform a process for generating a navigable 3D virtual environment, the process comprising: receiving a command in a concise language describing a virtual environment; using a natural language command processor to determine (i) a location portion and (ii) an experience portion of the command; using a generative virtual environment builder including one or more first machine learning models to generate a skybox based on the location portion of the command, wherein the skybox includes at least a shape and an image projected onto the shape; using a generative virtual environment builder including one or more second machine learning models to generate one or more 3D object models based on both the location portion and the experience portion of the command, wherein each 3D object model includes at least a geometric shape and a location within the navigable 3D virtual environment; creating a 3D virtual environment model that combines the skybox with the one or more 3D object models, wherein the 3D virtual environment model is navigable using an artificial reality (XR) device; and storing the 3D virtual environment model on a data storage device.
10. The computer-readable storage medium according to claim 9, wherein, The process further includes determining an embedding and metadata associated with the command, wherein the generative virtual environment builder receives the embedding and the metadata as part of its input to generate the skybox and the one or more 3D object models.
11. The computer-readable storage medium according to claim 9 or 10, wherein, Generating the skybox includes identifying elements semantically related to the determined location portion and adding a representation of the elements to the skybox.
12. The computer-readable storage medium according to claim 9, 10 or 11, wherein, Determining (i) the location part and (ii) the experience part of the command includes: Preprocessing the command by tokenizing and contextualizing phrases of the command using a first language model pre-trained to classify phrases according to location indicators or activity indicators; and preferably, wherein determining (i) the location part and (ii) the experience part of the command further includes: Applying a second language model to the tokenized and contextualized phrases and receiving one or more embeddings from the second language model; wherein the one or more embeddings are provided to the one or more first machine learning models to generate the skybox, and the one or more embeddings are provided to the one or more second machine learning models to generate the one or more 3D object models.
13. The computer-readable storage medium according to any one of claims 9 to 12, wherein, A) The one or more first machine learning models and / or the one or more second machine learning models include a 2D modeling part trained on a combination of images and / or videos labeled with metadata describing the content or context of these images and / or videos; and B) the one or more first machine learning models and / or the one or more second machine learning models produce one or more 2D representations based on one or more input labels, and the one or more first machine learning models and / or the one or more second machine learning models include a generative part trained to take one or more 2D representations and produce one or more 3D representations.
14. A computing system for generating a navigable 3D virtual environment, the computing system comprising: One or more processors; And One or more memories storing instructions which, when executed by the one or more processors, cause the computing system to perform a process that includes: Receiving a command in a concise language describing a virtual environment; Using a natural language command processor to determine (i) a location part and (ii) an experience part of the command; Using a generative virtual environment builder including one or more first machine learning models to generate a skybox based on the location part of the command, wherein the skybox includes at least a shape and an image projected onto the shape; Using a generative virtual environment builder including one or more second machine learning models to generate one or more 3D object models based on both the location part and the experience part of the command, wherein each 3D object model includes at least a geometric shape and a location within the navigable 3D virtual environment; Creating a 3D virtual environment model that combines the skybox with the one or more 3D object models, wherein the 3D virtual environment model is navigable using an artificial reality (XR) device; and Storing the 3D virtual environment model on a data storage device.
15. The computing system according to claim 14, i. Among them, Determine that the location portion includes a semantic identifier identifying the location portion; and wherein generating the skybox includes mapping the semantic identifier of the location portion into the latent space of the one or more second machine learning models of the generative virtual environment builder; and / or preferably, ii. wherein the process further includes iteratively updating the 3D virtual environment model by: Receiving a further natural language command; Identifying the target of the further natural language command; Changing one aspect of the target based on the further natural language command; and Updating the training of the one or more first machine learning models and / or the one or more second machine learning models using additional training items created based on pairing the target with the changed aspect of the target.