Generating three-dimensional digital content from natural language requests
Through a three-dimensional modeling system based on natural language requests, natural language processing technology is used to generate three-dimensional scenes, which solves the problem of high user professional knowledge requirements in the existing system, and achieves fast, accurate and flexible three-dimensional scene generation.
Patent Information
- Application Number
- CN202510417473.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-11-13
- Filing Date
- 2019-08-15
- Publication Date
- 2025-07-22
AI Technical Summary
The existing three-dimensional digital content creation system has high requirements for users' professional knowledge and skills, which makes it difficult for new users to learn, the generation process is time-consuming and inefficient, and requires complex operations between multiple interfaces and control menus.
Through a three-dimensional modeling system based on natural language requests, natural language processing technology is used to analyze the phrases entered by users, generate entity-command representations, and map them to the semantic scene graphics of existing three-dimensional scenes, and quickly generate or modify three-dimensional scenes.
It improves the flexibility, speed and accuracy of three-dimensional scene generation, reduces complexity, simplifies user operations, is suitable for new users, and reduces generation time and learning difficulty.
Smart Images

Figure CN120353371A_ABST
Abstract
Description
[0001] Division Application Instructions
[0002] This application is a divisional application of a Chinese patent application with an application date of August 15, 2019, an application number of 201910753563.0, and a title of "Generating 3D Digital Content from Natural Language Requests". Technical Field
[0003] Embodiments of the present disclosure relate to generating 3D digital content from natural language requests. Background Art
[0004] Recent advances in computational design, virtual reality / augmented reality (VR / AR), and robotics have increased the demand for 3D digital content. For example, many computer-aided design or VR / AR systems utilize 3D models. High-quality 3D models can significantly improve the aesthetic design, realism, and immersion in a 3D environment.
[0005] Conventional systems typically utilize software applications that allow content creators to use various tools to create 3D digital content. Conventional software applications provide a high degree of customization and precision, thereby allowing content creators to generate anything from basic 3D shapes to highly detailed and complex 3D scenes with many 3D objects. Although conventional systems provide a great deal of control to content creators, such applications have a large number of tools to perform a large number of operations. Therefore, using a conventional system to create a 3D scene typically requires a great deal of training to learn to use the content creation tools, which may provide a high barrier to entry for new users. Accordingly, creating 3D digital content is limited by the expertise and capabilities of the content creator.
[0006] In addition, even when the user is proficient and knowledgeable, conventional systems typically require navigation between various user interfaces and / or various control menus in order to generate a 3D scene. Therefore, even for proficient and knowledgeable users, using a conventional system to create a 3D scene is both time-consuming and inefficient.
[0007] There are these and other drawbacks with respect to conventional systems for creating 3D digital content. Summary of the Invention
[0008] One or more embodiments provide benefits and / or solve one or more of the foregoing or other problems in the art using systems, methods, and non-transitory computer-readable storage media that intelligently generate three-dimensional digital content based on natural language requests. More particularly, the disclosed systems include a framework that generates a language-based representation of an existing 3D scene, the language-based representation encoding geometric information and semantic scene information about the 3D scene. When a natural language command to generate or modify a 3D scene is received, the disclosed systems also generate a representation of the natural language phrase that encodes geometric information, semantic information, and the relationship to the command. The disclosed systems then map the natural language representation to one or more language-based representations of the 3D scene or sub-scene. The disclosed systems then use the identified 3D scene or sub-scene to generate or modify the 3D scene based on the natural language command.
[0009] For example, in one or more embodiments, the disclosed systems analyze a natural language phrase that requests generation of a three-dimensional scene to determine dependencies involving one or more entities and / or commands in the natural language phrase. Specifically, the disclosed systems utilize the dependencies involving the (multiple) entities and (multiple) commands to generate an entity-command representation of the natural language phrase, the entity-command representation annotated with attributes and relationships of the (multiple) entities and (multiple) commands. Additionally, the disclosed systems generate a three-dimensional scene based on the entity-command representation using at least one three-dimensional scene from a database of previously generated three-dimensional scenes. Specifically, the disclosed systems may select a three-dimensional scene from the database by correlating the entity-command representation with the semantic scene graph of the three-dimensional scene. Thus, the disclosed systems can effectively, flexibly, and accurately generate a three-dimensional scene from a natural language request by determining a representation of the request that allows comparison with a representation of an existing three-dimensional scene.
[0010] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description below, and in part will be obvious from the description, or may be learned by practice of the example embodiments of one or more embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Various embodiments will be described and explained using additional features and details with the accompanying drawings, wherein:
[0012] Figure 1 Illustrates an example environment in which a three-dimensional (3D) modeling system according to one or more implementations may operate;
[0013] Figure 2A schematic diagram illustrating a process of generating a three-dimensional scene from a natural language phrase according to one or more implementations;
[0014] Figures 3A to 3C A schematic diagram illustrating a process of parsing a natural language phrase according to one or more implementations;
[0015] Figures 4A to 4C A schematic diagram illustrating different entity-command representations for generating a three-dimensional scene according to one or more implementations;
[0016] Figure 5 A schematic diagram illustrating the selection of a previously generated three-dimensional scene according to one or more implementations;
[0017] Figures 6A to 6C An embodiment illustrating the generation of a three-dimensional scene from a series of natural language phrases according to one or more implementations;
[0018] Figure 7 Illustrates according to one or more implementations Figure 1 A schematic diagram of the illustrated three-dimensional (3D) modeling system;
[0019] Figure 8 A flowchart illustrating a series of actions for synthesizing a three-dimensional scene using natural language according to one or more implementations; and
[0020] Figure 9 A block diagram of an exemplary computing device according to one or more embodiments. Detailed Description
[0021] One or more embodiments of the present disclosure include a natural language-based three-dimensional modeling system (also referred to as a "natural language-based 3D system" or simply a "3D modeling system"), which generates a three-dimensional scene based on a natural language request. For example, the 3D modeling system uses natural language processing to analyze a natural language phrase and determine the dependencies of the components involved in the natural language phrase. In particular, the 3D modeling system can determine the relationships between various nouns and / or verbs in the phrase and generate an entity-command representation of the phrase based on the determined relationships. In addition, the 3D modeling system also generates semantic scene graphs for existing 3D scenes and sub-scenes. These semantic scene graphs encode geometric information and semantic scene information about the 3D scene. Then, the 3D modeling system maps the entity-command representation of the phrase to the semantic scene graph of the existing three-dimensional scene. The 3D modeling system uses the identified three-dimensional scene to generate a three-dimensional scene to meet the request in the natural language phrase. By generating semantic scene graphs based on the entity-command representation of the natural language request, the 3D modeling system can accelerate and simplify the process of generating a three-dimensional scene by quickly finding an existing three-dimensional scene corresponding to the requested content.
[0022] As mentioned, a 3D modeling system can use natural language processing to analyze natural language phrases including requests for generating 3D scenes. In one or more embodiments, the 3D modeling system uses natural language processing to transform a natural language phrase into a representation that the 3D modeling system can use to compare with a similar representation of an existing 3D scene. Specifically, the 3D modeling system tags one or more entities and one or more commands in the natural language phrase and then determines the dependencies involving the entities and commands. For example, the 3D modeling system can determine the dependencies by parsing the natural language phrase and creating a dependency tree that assigns a parent token and annotation labels to each token in the phrase.
[0023] In one or more embodiments, the 3D modeling system uses the determined dependencies to generate an entity-command representation of the natural language phrase. Specifically, the 3D modeling system converts the dependency representation of the phrase tokens (e.g., the dependency tree) into an entity-command representation that provides a detailed graphical representation of the components of the natural language phrase and their relationships. For illustration, the entity-command representation can include a list of entities annotated with corresponding attributes and relationships and a list of command verbs operating on the entities.
[0024] In one or more additional embodiments, the 3D modeling system determines a canonical entity-command representation of multiple natural language phrases including requests for the same concept. For example, the 3D modeling system can determine that there are multiple different ways to express a request to build the same 3D scene. Since the phrases include different parsing structures, the 3D modeling system creates different entity-command representations for each form. Then, the 3D modeling system can select the entity-command representation of a form as the canonical entity-command representation such that future requests including requests for the same concept utilize the canonical entity-command representation.
[0025] After generating the entity-command representation of the natural language phrase, the 3D modeling system generates a semantic scene graph of the natural language phrase. Specifically, the 3D modeling system converts the entity-command representation into a semantic scene graph for use in generating a 3D scene. For example, the 3D modeling system determines object categories, entity counts, qualifiers, and relationships from the entity-command representation. For illustration, the 3D modeling system includes object nodes, relationship nodes, and edge nodes based on the determined information to represent the relative positioning and relationships of objects within the requested 3D scene.
[0026] Using 3D scene graphics constructed for natural language phrases, a 3D modeling system then generates a three-dimensional scene. In one or more embodiments, the 3D modeling system uses at least one three-dimensional scene from a database of available three-dimensional scenes to generate the three-dimensional scene. For example, the 3D modeling system can compare the semantic scene graph of the natural language phrase with the semantic scene graph of the available three-dimensional scenes to select at least one scene in the scene that best matches the semantic scene graph of the natural language phrase.
[0027] As mentioned, the 3D modeling system offers several advantages over conventional systems. For example, the 3D modeling system improves the flexibility of the three-dimensional scene generation process. In particular, the 3D modeling system improves flexibility by allowing users to generate three-dimensional scenes using natural language phrases. By using natural language requests, the 3D modeling system improves the usability of the three-dimensional modeling program for inexperienced users by improving the input methods available to the user when generating a three-dimensional scene. In contrast, conventional systems typically require users to have a deep understanding of the modeling tools and how to use them.
[0028] In addition, the 3D modeling system improves the speed and efficiency of the three-dimensional scene generation process. Specifically, by constructing an object representation (e.g., a semantic scene graph) of the natural language phrase for generating the three-dimensional scene, the 3D modeling system reduces the time required to generate the three-dimensional scene. For example, in contrast to conventional systems that require the user to use multiple three-dimensional modeling tools to generate each object within the scene, the 3D modeling system described herein can interpret a natural language request for a three-dimensional scene and then quickly generate the three-dimensional scene without using any additional tools. Thus, the 3D modeling system improves the conventional user interface and modeling system by increasing the efficiency of the computing device using natural language-based generation of 3D scenes.
[0029] In addition to the foregoing, when determining a request and generating a three-dimensional scene based on a natural language phrase, the 3D modeling system also improves consistency and reduces complexity. In particular, the 3D modeling system determines a canonical entity-command representation that can be used to represent various different forms (i.e., scene editing / construction concepts) of a conceptual request. For example, the 3D modeling system can determine that a particular form (e.g., a descriptive form) of the entity-command representation is the canonical entity-command representation for representing requests of different forms for generating / editing a three-dimensional scene.
[0030] Additionally, the 3D modeling system improves the accuracy of the three-dimensional scene generation process. In particular, by creating a semantic scene graph of natural language phrases through a phrase-based entity command representation, the 3D modeling system generates a representation of the phrase that the system can use to identify similar existing three-dimensional scenes. The semantic scene graph provides the 3D modeling system with a representation that the 3D modeling system can easily compare with the available three-dimensional scenes in a database of three-dimensional scenes, enabling the 3D modeling system to accurately determine the relationships between objects within a three-dimensional scene based on a request.
[0031] As illustrated by the foregoing discussion, the present disclosure uses various terms to describe the features and advantages of the 3D modeling system. Additional details regarding the meanings of the terms are now provided. For example, as used herein, the term "natural language phrase" refers to text or speech that includes ordinary language employed by a user. Specifically, a natural language phrase can include text or speech that does not have a special syntactic or formal structure specifically configured for interacting with a computing device. For example, a natural language phrase can include a conversational request to generate or modify a scene. By way of illustration, a user can use natural language when speaking a request or otherwise entering the request into a computing device.
[0032] As used herein, the term "three-dimensional scene" refers to a digital representation of one or more objects in a three-dimensional environment. For example, a three-dimensional scene can include any number of digital objects located at one or more positions within a three-dimensional environment according to a set of coordinate axes (e.g., the x-axis, y-axis, and z-axis). By way of illustration, displaying the objects of a three-dimensional scene on a display device of a computing device involves using the mathematical positioning of the objects on the coordinate axes to reconstruct the objects on the display device. Reconstructing a three-dimensional scene takes into account the size and shape of each object based on the coordinates used by the computing device to reconstruct the object for a plurality of vertices. The mathematical positioning also allows the computing device to present the relative positioning of the objects with respect to each other.
[0033] As used herein, the terms "entity" and "command" refer to the spoken or textual components of phrases analyzed by a natural language processor. Specifically, a 3D modeling system uses a natural language processor to identify tokens in a sentence by identifying strings with an assigned / identified meaning. The 3D modeling system then labels the tokens as either entities or commands. As used herein, the term "entity" refers to an identified noun in a sentence that meets a set of conditions, including: determining that the noun does not have a compound dependency with another noun, the noun is not an abstract concept, and the noun does not represent a spatial region. In particular, an entity can include a base noun associated (e.g., annotated) with one or more attributes, one or more counts, one or more relationships, or one or more determiners based on the structure of the sentence. As used herein, the term "command" refers to an identified command verb that operates on one or more entities. Specifically, a command can include a base verb associated with one or more attributes or targets. Examples of entities and commands are described in more detail below. In one or more embodiments, the 3D modeling system labels all verbs in their base form as commands.
[0034] As used herein, the term "entity-command representation" refers to a logical representation of the entities and commands in a natural language phrase and their relationships. Specifically, an entity-command representation includes a list of entities annotated with attributes and relationships and a list of command verbs that operate on the entities. As described in more detail below, an entity-command representation can also be diagrammed as a block diagram indicating the attributes and relationships (e.g., dependencies) between the different components of a natural language phrase.
[0035] As used herein, the term "semantic scene graph" refers to a graphical representation of a three-dimensional scene that includes geometric information for the objects in the scene and semantic information for the scene. In particular, in addition to object relationships (e.g., pairing and grouping) indicating the relative positioning of object instances, a semantic scene graph can also include object instances and object-level attributes. Additionally, a semantic scene graph includes two node types (object nodes and relationship nodes) that have edges connecting the object nodes and relationship nodes. Thus, a semantic scene graph can represent object relationships and composition in both natural language phrases and three-dimensional scene data.
[0036] Additional details regarding the 3D modeling system will now be provided with reference to the illustrative diagrams depicting exemplary implementations. For purposes of illustration, Figure 1Embodiments of an environment 100 in which a natural language 3D modeling system 102 can operate are included. In particular, the environment 100 includes a client device 103 associated with a user, one or more server devices 104, and a content database 106, which communicate via a network 108. Additionally, as shown, the client device 103 includes a client application 110. Further, the one or more server devices 104 include a content creation system 112, which includes the 3D modeling system 102.
[0037] The content database 106 is a database containing a plurality of existing / available three-dimensional scenes or sub-scenes. Specifically, the content database 106 includes three-dimensional scenes with different objects and / or object arrangements. As described below, the content database 106 provides various different three-dimensional scenes that the 3D modeling system 102 can use to generate three-dimensional scenes from natural language phrases. In one or more embodiments, the content database 106 includes content from multiple content creators. Additionally, the content database 106 can be associated with the content creation system 112 (e.g., including content from users of the content creation system 112) and / or a third-party system (e.g., including content from users of the third-party system).
[0038] In one or more embodiments, the 3D modeling system 102 also accesses the content database 106 to obtain semantic scene graph information for a plurality of available three-dimensional scenes. In particular, the 3D modeling system 102 can obtain the semantic scene graphs of a plurality of available three-dimensional scenes in the content database 106. The 3D modeling system 102 can obtain the semantic scene graph by analyzing the three-dimensional scene and then generating the semantic scene graph. Alternatively, the 3D modeling system 102 obtains pre-generated semantic scene graphs from the content database 106 or a third-party system.
[0039] As mentioned, the one or more server devices 104 can host or implement the content creation system 112. The content creation system 112 can manage content creation for multiple users. The content creation system 112 can include or be associated with a plurality of content creation applications that allow multiple users to create various types of digital content and otherwise interact with various types of digital content. The content creation system can also provide content storage and sharing capabilities (e.g., via the content database 106), such that content creators can view the content of other content creators.
[0040] In addition, in one or more embodiments, the 3D modeling system 102 communicates with the client device 103 via the network 108 to receive one or more requests for generating a three-dimensional scene. For example, the client device 103 may include a client application 110 that enables communication with the natural language-based 3D modeling system 102. For example, the client application 110 may include a web browser or a native computing device. In an example implementation, a user of the client device 103 may use the client application 110 to input a natural language phrase that includes a request for generating a three-dimensional scene. For illustration, the client application 110 may allow the user to input the natural language phrase using a keyboard input or a voice input. For voice input, the client application 110 or the 3D modeling system 102 may convert the voice input into text input (e.g., using speech-to-text analysis).
[0041] The 3D modeling system 102 may receive the natural language phrase from the client device 103 and then analyze the natural language phrase to determine the request. Specifically, the 3D modeling system 102 parses the natural language phrase to identify the components and corresponding dependencies of the natural language phrase. The 3D modeling system 102 uses the components and dependencies to generate an entity-command representation of the natural language phrase. Based on the entity-command representation, the 3D modeling system 102 generates a semantic scene graph for the natural language phrase to indicate the objects within the three-dimensional scene and their relative positions.
[0042] The 3D modeling system 102 generates a three-dimensional scene based on the semantic scene graph of the natural language phrase. In one or more embodiments, the 3D modeling system 102 compares the semantic scene graph of the natural language phrase with the semantic scene graphs of multiple available three-dimensional scenes in the content database 106 to select one or more three-dimensional scenes. Then, the 3D modeling system 102 may use the selected scenes to generate a three-dimensional scene that satisfies the natural language request received from the client device 103. The 3D modeling system 102 may also allow the user to modify the generated three-dimensional scene using one or more additional natural language phrases. After generating the three-dimensional scene, the 3D modeling system 102 may store the three-dimensional scene in the content database 106 for use in generating future three-dimensional scenes.
[0043] Although Figure 1The illustrated environment is depicted as having various components, but environment 100 can have any number of additional or alternative components (e.g., any number of server devices, client devices, content databases, or other components that communicate with 3D modeling system 102). For example, 3D modeling system 102 can allow any number of users associated with any number of client devices to generate three-dimensional scenes. Additionally, 3D modeling system 102 can communicate with any number of content databases to access existing three-dimensional scenes for generating three-dimensional scenes based on natural language requests. Additionally, more than one component or entity in environment 100 can implement the operations of 3D modeling system 102 described herein. For example, alternatively, 3D modeling system 102 can be implemented in its entirety (or in part) on client device 103 or on a separate client device.
[0044] As mentioned above, 3D modeling system 102 can parse a natural language phrase to generate a semantic scene graph for the phrase for generating a three-dimensional scene. Figure 2 A schematic diagram illustrating the process of generating a three-dimensional scene from a natural language phrase is shown. In particular, the process includes a series of actions 200, where 3D modeling system 102 processes the natural language phrase to convert the natural language phrase into a form that 3D modeling system 102 can use to compare with an already existing three-dimensional scene.
[0045] In one or more embodiments, the series of actions 200 includes a first action 202 of identifying a natural language phrase. Specifically, 3D modeling system 102 identifies the natural language phrase based on a user's input. The user input can include text input (e.g., via a keyboard or touch screen) or voice input (e.g., via a microphone). 3D modeling system 102 can also determine that the user input includes multiple natural language phrases and then process each phrase individually. Alternatively, 3D modeling system 102 can determine that multiple natural language phrases are combined to form a single request and process the multiple natural language phrases together.
[0046] After identifying the natural language phrase, the series of actions 200 includes an action 204 of parsing the phrase to create a dependency tree. In one or more embodiments, 3D modeling system uses natural language processing to analyze the natural language phrase. For example, 3D modeling system uses natural language processing to parse the natural language phrase to determine the different syntactic components of the phrase and the relationships and dependencies between the components within the phrase. One way the natural language processor can determine these dependencies is by creating a dependency tree.
[0047] A dependency tree (or dependency-based parse tree) is a tree that includes multiple nodes with a dependency grammar. For example, in one or more embodiments, the nodes in the dependency tree are all terminals, such that there is no distinction between terminal categories (formal grammar elements) and non-terminal categories (syntactic variables), as is the case with constituency-based parse trees. Additionally, in one or more embodiments, the dependency tree has fewer nodes than a constituency-based parse tree because the dependency tree lacks phrasal categories (e.g., sentences, noun phrases, and verb phrases). Instead, the dependency tree includes a structure that indicates the dependency relationships of the components of a natural language phrase by setting the verb as the structural center (e.g., the first parent node) of the dependency tree and then constructing the tree based on the noun dependencies (either direct or indirect) with respect to the verb node.
[0048] In one or more embodiments, a 3D modeling system converts a natural language phrase into a dependency tree by first identifying the tokens of the natural language phrase. As previously mentioned, tokens include strings (e.g., words) with an identified meaning. Then, the 3D modeling system assigns annotation labels to each token in the natural language phrase. This includes: identifying nouns, verbs, adjectives, etc. in the natural language phrase and then determining the dependency of each component in the phrase with respect to the other components in the phrase. An example of assigning a parent token and annotation labels to each token in the phrase is described in more detail below. According to at least some embodiments, the 3D modeling system uses an established framework for determining lexical dependencies (e.g., the Universal Dependencies framework). Figure 3A The series of actions 200 includes an action 206 of generating an entity-command representation of a natural language phrase. In one or more embodiments, the 3D modeling system converts the low-level dependency representation of a natural language phrase in a dependency tree into an entity-command representation by tagging different components from the dependency tree. Specifically, generating the entity-command representation involves an action 208a of tagging entities and an action 208b of tagging commands. According to one or more embodiments, the 3D modeling system 102 tags the entities and commands in a natural language phrase by analyzing the natural language phrase to identify specific types of speech components and then providing corresponding labels to the identified types of speech components. For example, as described in more detail for
[0049] and Figure 3A and Figure 3B The 3D modeling system tags all nouns in a sentence as entities, unless one of multiple conditions is met: 1) the noun has a compound dependency relationship with another noun (e.g., "computer" in "computer desk"); 2) the noun is an abstract concept (e.g., "additive", "attraction"); or 3) the noun represents a special region (e.g., "right", "side"). Additionally, the 3D modeling system tags all base form verbs as commands.
[0050] In addition, when tagging entities, the 3D modeling system 102 dereferences pronouns. In particular, the 3D modeling system 102 replaces pronouns in natural language phrases with the corresponding nouns in the phrase to identify the associated relationships involving the underlying nouns, rather than creating new entities for the pronouns themselves in the entity-command representation. For example, the 3D modeling system may use co-reference information from a natural language processing framework (e.g., the Stanford CoreNLP framework). If there is ambiguity for a non-pronoun scenario where two or more different objects in a natural language phrase (or in more than one natural language phrase) can be referred to, the 3D modeling system 102 does not use such co-reference information. In such instances, the 3D modeling system 102 resolves the ambiguity when aligning entities with the objects in the requested three-dimensional scene.
[0051] Generating the entity-command representation may involve performing pattern matching in action 210 to assign additional attributes to the tagged entities and commands. In one or more embodiments, the 3D modeling system 102 uses pattern matching on the dependency tree to assign attributes that are not tagged as entities or commands. Specifically, the 3D modeling system 102 may analyze natural language phrases to identify nouns and verbs that describe or modify the underlying noun or verb. For illustration, the 3D modeling system 102 may identify spatial nouns, counting adjectives, group nouns, and gerunds, and then enhance the underlying noun / verb with an annotation indicating the association. Figure 3B And the accompanying description indicates the result of the pattern matching to assign non-entity / command attributes of the simple natural language phrase, thereby generating the entity-command representation of the natural language phrase.
[0052] Optionally, in one or more embodiments, the series of actions 200 includes an action 212 of creating a canonical entity-command representation. As briefly mentioned previously, requests for generating a particular three-dimensional scene can typically take different forms in natural language phrases. In particular, the flexibility of language generally allows users to use different words, combinations of words, and word orders to express requests in various different ways to describe essentially the same thing. Since there can be different ways of forming the same request, the 3D modeling system 102 can generate multiple different entity-command representations corresponding to the different ways.
[0053] To improve consistency and reduce complexity when determining requests and generating three-dimensional scenes based on natural language phrases, the 3D modeling system 102 can determine a single entity-command representation that can be used to represent all different forms (i.e., scene editing / construction concepts) of the request for a concept. For example, as for Figures 4A to 4CMore specifically described, the 3D modeling system 102 can determine that an entity-command representation of a particular form (e.g., a descriptive form) is a canonical entity-command representation representing requests of different forms for generating / editing a three-dimensional scene. Accordingly, the 3D modeling system 102 can attempt to transform the entity-command representation of a natural language phrase into a canonical entity-command representation via a set of pattern matching rules.
[0054] In one or more embodiments, if a command cannot be applied as a graphical transformation in an entity-command representation, the 3D modeling system leaves the command unchanged and uses a dedicated function to execute the command regarding the scene. For example, for commands such as "delete" or "rotate", the 3D modeling system can use a dedicated function to perform the "delete" or "rotate" function on the corresponding objects within the three-dimensional scene (e.g., after constructing the three-dimensional scene without the command being executed). Alternatively, the 3D modeling system can notify the user that the command was not understood and that the user is required to input a new natural language phrase.
[0055] After determining the entity-command representation for a natural language phrase, the series of actions 200 includes the action 214 of converting the entity-command representation into a semantic scene graph. Specifically, the 3D modeling system 102 converts the entity-command representation into a form that the 3D modeling system 102 can use to generate a three-dimensional scene. As previously mentioned, the semantic scene graph serves as a bridge between the user's language commands and the scene modeling operations, which directly modify the three-dimensional scene. Accordingly, converting the entity-command representation into a semantic scene graph allows the 3D modeling system to easily construct a three-dimensional scene based on the identified positional relationships of the objects in the semantic scene graph.
[0056] In one or more embodiments, converting the entity-command representation includes: determining the layout of the objects within the three-dimensional scene. Specifically, the 3D modeling system 102 uses the entity-command representation to classify the base nouns by object category, determine the entity count, entity qualifiers, and relationships. As described in more detail Figure 3C More specifically described, the 3D modeling system 102 then constructs a semantic scene graph to include object nodes, relationship nodes, and edges indicating the pairwise and grouped positioning of the objects within the three-dimensional scene. The semantic scene graph of the natural language phrase allows the 3D modeling system 102 to determine the spatial representation of the objects for the requested three-dimensional scene of the natural language phrase.
[0057] Finally, the series of operations 200 includes an operation 216 to generate a three-dimensional scene based on a semantic scene graph. As briefly described previously, the 3D modeling system 102 can access a database including a plurality of available three-dimensional scenes (e.g., previously generated three-dimensional scenes), the plurality of available three-dimensional scenes including various objects and object layouts. The 3D modeling system 102 compares the semantic scene graph of the natural language phrase with the semantic scene graphs of the available three-dimensional scenes, and then selects one or more three-dimensional scenes based on the degree of match of the semantic scene graphs. Then, the 3D modeling system uses the selected three-dimensional scene(s) (or part of the three-dimensional scene) to generate a three-dimensional scene based on the user request.
[0058] For illustration, the 3D modeling system 102 obtains the semantic scene graph for the available three-dimensional scenes. Then, the 3D modeling system 102 compares the semantic scene graph of the natural language phrase with the obtained semantic scene graph to determine the similarity of the structure of the graphs. For example, if the two semantic scene graphs match exactly, the semantic scene graphs have the same node and edge label structure. Similarly, a partial match indicates that some of the nodes (or edges) in the semantic scene graphs are the same, but other nodes (or edges) are different. Then, the 3D modeling system 102 can rank the semantic scene graphs based on the degree of similarity of the node structure to determine one or more available three-dimensional scenes that best match the requested three-dimensional scene.
[0059] To generate the requested three-dimensional scene, the 3D modeling system 102 can use one or more available three-dimensional scenes that best match the requested three-dimensional scene. For example, the 3D modeling system can extract one or more objects from the available three-dimensional scenes, and then place the one or more objects in a three-dimensional environment at the user's client device. For illustration, the 3D modeling system 102 can place the object(s) at coordinates that reflect the relative positions of the objects on a set of coordinate axes. For example, as described in more detail with respect to Figures 6A to 6C If the user is modifying an existing scene, the 3D modeling system 102 can insert the object into the existing scene. Alternatively, the 3D modeling system 102 can cause the user client device to directly load a file containing the existing three-dimensional scene into a new project workspace.
[0060] As described with respect to Figure 2 And in the corresponding Figures 3A to 6C As described, the 3D modeling system 102 can perform operations for processing natural language phrases to create an entity-command representation (and corresponding semantic scene graph) of the natural language phrase. These operations allow the 3D modeling system to receive natural language requests to generate new three-dimensional scenes and / or modify existing three-dimensional scenes using existing scenes. Accordingly, as described above with respect to Figure 2 And below with respect toFigures 3A to 6C The illustrated and described actions and operations provide corresponding structures for example steps for generating an entity-command representation that uses natural language processing to address dependencies of one or more entities and one or more commands in a natural language phrase.
[0061] Figure 3A Illustrated is a dependency tree 300 generated by a 3D modeling system 102 for a natural language phrase. Specifically, Figure 3A Illustrated is a natural language phrase that reads "Put some books on the desk." The dependency tree illustrates dependencies (including dependency / relationship types) involving different phrase components in the natural language phrase. The dependency tree provides a consistent lexical structure that the 3D modeling system can use to identify different words in the phrase.
[0062] In one or more embodiments, the 3D modeling system 102 identifies a plurality of tokens in the natural language phrase that correspond to a plurality of strings having the identified meanings. As constructed, the natural language phrase includes a request for the 3D modeling system to place some books on a desk within a three-dimensional modeling environment. The 3D modeling system 102 first identifies each string having the identified meaning such that, in this case, the 3D modeling system 102 identifies each string in the natural language phrase as a token.
[0063] After identifying the tokens in the natural language phrase, the 3D modeling system 102 annotates each token with a label indicating a particular component of speech to which the token belongs. For example, the 3D modeling system 102 can identify and label nouns (including whether the noun is plural), verbs, adjectives, prepositions, and determiners in the natural language phrase. As illustrated in Figure 3A , the 3D modeling system 102 determines that "put" is a verb, "some" is an adjective, "books" is a plural noun, "on" is a preposition, "the" is a determiner, and "desk" is a singular noun. Accordingly, the 3D modeling system 102 has identified each word in the phrase and labeled it with its corresponding token type / category.
[0064] Further, when generating the dependency tree, the 3D modeling system 102 also determines the dependencies of the individual tokens in the context of the natural language phrase. In particular, each dependency indicates a relationship involving the token and one or more other tokens in the natural language phrase. For purposes of illustration, Figure 3AThe 3D modeling system identifies the verb token "put" as acting on the direct object "book" as a noun modifier to place the direct object on the prepositional noun "table" according to the identified preposition "on". Further, the 3D modeling system 102 determines that "some" is an adjective modifying "book" to indicate that a number of books are to be placed on the table. The 3D modeling system 102 also determines that "the" is a determiner for "table". By identifying and labeling the components and their relationships, the 3D modeling system generates a tree in which the verb "put" serves as the structural center (e.g., the root node) of the tree.
[0065] Thus, the 3D modeling system 102 uses a natural language processor to determine the labels for any nouns, verbs, and attributes in the natural language phrase. For more complex sentences, the 3D modeling system 102 can generate a larger dependency tree with more branches based on the corresponding relationships and the number of nouns, verbs, modifiers, and prepositions. As an example, the phrase "In the center of the table is a vase" includes multiple nouns that the 3D modeling system 102 identifies and labels, some of which are base nouns and some of which are compound nouns, etc. Although the structure of this phrase is more complex than Figure 3A the embodiments, the 3D modeling system is able to create an accurate dependency tree for the phrase with "is" as the structural center.
[0066] Once the 3D modeling system 102 has created a dependency tree for the natural language phrase, the 3D modeling system 102 can then generate an entity-command representation for the natural language phrase. Figure 3B Illustrated Figure 3A is the entity-command representation 302 of the shown natural language phrase. In particular, the 3D modeling system 102 generates the entity-command representation 302 based on the components and dependencies identified in the dependency tree 300 for the natural language phrase.
[0067] In one or more embodiments, the 3D modeling system 102 generates an entity-command representation by defining a set of entities and a set of commands for a natural language phrase based on a dependency tree. In particular, the entities include categories of base nouns and are associated with any attributes of the base nouns, the count of the nouns, the relationships that relate the nouns to another entity within the sentence, and any determiners corresponding to the nouns. The entity categories include base nouns that are used to describe objects in a three-dimensional scene (e.g., "table," "plate," "arrangement"). As mentioned, base nouns include those that do not have a compound dependency relationship with another noun, are not abstract concepts, and do not represent a spatial region. The attributes of a base noun include one or more modifiers that are used to modify the base noun (e.g., "modern," "blue (very, dark)"). The count of a noun includes an integer or qualitative descriptor that represents the number of entities in a group (e.g., "2," "three," "many," "some"). The relationships that relate a noun to another entity include a set of (string, entity) pairs that describe the relationship to another specific entity in the sentence (e.g., "on: desk," "to the left of: keyboard"). The determiners corresponding to a noun include words, phrases, or affixes that express a reference to the noun in context (e.g., "a," "the," "another," "each").
[0068] Additionally, the commands include base verbs and are associated with any attributes of the verbs and any targets of the verbs. By way of illustration, the base verbs include verbs that are used to describe commands (e.g., "move," "rearrange"). The attributes of a verb include a set of one or more modifier words that help modify the verb / command (e.g., "closer to," "dimmer (more)"). The targets of a verb are a list of (string, entity) pairs that represent different types of relationships to a specific entity (e.g., "direct object: laptop," "on: table").
[0069] As illustrated in Figure 3B the 3D modeling system 102 converts the dependency tree 300 shown in Figure 3A to an entity-command representation 302 by placing the identified entities and commands in a logical configuration that relates the entities and commands based on the determined dependencies from the dependency tree 300. For example, to convert the phrase "put some books on the desk" to an "entity-command representation, the 3D modeling system 102 identifies two separate entities - a first entity 304a ("books") and a second entity 304b ("desk").
[0070] In one or more embodiments, the 3D modeling system 102 uses pattern matching to determine additional properties of natural language phrases and assigns them to corresponding entities or commands in the entity-command representation 302. For example, if present, the 3D modeling system 102 determines that amod(noun: A, adjective: B) assigns the token B as a property or count of the entity planted at A. When performing pattern matching, the 3D modeling system 102 enhances the standard speech parts used by the natural language processor to create a dependency tree 300 with four classes that are used for scene understanding. The first class includes spatial nouns that are spatial regions relative to an entity (e.g., "right", "center", "side"). The second class includes counting adjectives that are adjectives representing object counts or general determiners (e.g., "all", "many"). The third class includes group nouns that have a special meaning in a set of object classes (e.g., "stack", "arrangement"). The fourth class includes gerunds, which the 3D modeling system 102 can model as property modifiers of a direct object (e.g., "clean", "brighten").
[0071] For illustration, the 3D modeling system 102 identifies the associated components of the phrase corresponding to each of the entities. Specifically, when constructing the entity-command representation 302, the 3D modeling system 102 determines that "some" is a counting adjective modifying the first entity 304a and creates a block 306 corresponding to "some" associated with the first entity 304a. Additionally, the 3D modeling system 102 determines that "the" is a determiner corresponding to the second entity 304b and creates a block 308 corresponding to "the" associated with the second entity 304b.
[0072] In addition to associating appropriate blocks with entities, the 3D modeling system 102 also associates entities with the command(s) within the entity-command representation 302. Specifically, the 3D modeling system 102 first determines that "put" is the command 310 within the phrase. Then, the 3D modeling system 102 determines that the first entity 304a is the target (i.e., direct object) of the command 310 and creates a block 312 to indicate the target relationship between the first entity 304a and the command 310. The 3D modeling system 102 also determines, via a preposition (i.e., "on"), that the second entity 304b has a relationship with the command 310 and the first entity 304a and creates a block 314 that associates the second entity 304b with the command 310. Accordingly, the 3D modeling system 102 associates the first entity 304a and the second entity 304b via the determined relationships and the command 310.
[0073] Although Figure 3B illustrates for Figure 3AThe entity-command representation 302 of the shown phrase, but the 3D modeling system 102 can generate entity-command representations of more complex phrases involving any number of nouns and commands. For example, a natural language phrase can include multiple requests. The 3D modeling system 102 can parse the natural language phrase to determine each separate request and then generate a separate entity-command representation for each request. For illustration, for a natural language phrase stating "Move the chairs around the dining table further away and move some books on the desk to the table", the 3D modeling system 102 can determine that the natural language phrase includes a first request to move the chairs around the dining table further away and a second request to move some books from the desk to the table.
[0074] Additionally, as previously mentioned, the 3D modeling system 102 can use co-reference information to dereference pronouns. This means that: the 3D modeling system 102 does not create a new entity for the pronoun, but determines the corresponding noun and then applies any attributes, relationships, etc. to the corresponding noun rather than a new entity. For example, in the phrase "Add a dining table and place the plate on top of it", the 3D modeling system determines that "it" corresponds to the dining table.
[0075] In one or more embodiments, the 3D modeling system 102 determines that there is ambiguity in a non-pronoun scenario (e.g., two nouns with the same name or meaning), however, when aligning entities with objects in a 3D three-dimensional scene, the 3D modeling system 102 can resolve the ambiguity. For example, in the above-mentioned phrase requesting to move books from the desk to the table, "table" can refer to the table in the first request to move the chairs further away, or "table" can refer to a different table in the same scene. Accordingly, the 3D modeling system 102 can create a new entity for "table" (e.g., in a separate entity-command representation). Then, the 3D modeling system 102 will attempt to resolve the ambiguity when constructing or modifying the requested three-dimensional scene by determining whether the scene includes one table or more than one table.
[0076] As shown in Figure 3C After generating the entity-command representation 302, the 3D modeling system 102 generates a semantic scene graph 316 corresponding to the entity-command representation 302. In one or more embodiments, the 3D modeling system 102 generates the semantic scene graph 316 to create a representation that the 3D modeling system 102 can use to easily identify the previously generated three-dimensional scene that most corresponds to the requested three-dimensional scene. As described below, the semantic scene graph 316 includes a plurality of nodes and edges indicating the spatial relationships between the objects in the three-dimensional scene.
[0077] Specifically, the semantic scene graph is an undirected graph that includes object nodes, relationship nodes, and edge labels. The object nodes represent objects in the 3D scene. The 3D modeling system 102 can annotate the objects in the scene with a list of per-object attributes (e.g., "ancient", "wooden"). The relationship nodes represent specific instances of relationships between two or more objects (e.g., "to the left of", "on each side of", "around", "above"). Additionally, each edge label connects an object node to a relationship node and describes the type of connection between the corresponding object and relationship. Each of these components is described in more detail below.
[0078] In one or more embodiments, the 3D modeling system 102 generates the semantic scene graph 316 by first assigning each base noun to an object category. Specifically, the 3D modeling system 102 can use a model database with a fixed set of object categories to map the base noun for each entity to the corresponding object category and create an object node for the base noun. For example, the 3D modeling system 102 can use the equivalence set obtained from Princeton WordNet. The 3D modeling system 102 discards entities that are not mapped to the categories in the model database. Additionally, the 3D modeling system 102 adds attributes and qualifiers as annotations to the corresponding object nodes.
[0079] For the phrases illustrated in Figures 3A to 3B the 3D modeling system 102 generates an object node for each of the identified entities. Specifically, the 3D modeling system 102 uses the entity count information from the entity-command representation 302 to instantiate a new object node instance for each object instance. For integer counts, the 3D modeling system 102 can simply generate multiple object nodes corresponding to the integer count. For illustration, the integer count "three" causes the 3D modeling system 102 to create three separate object nodes for the entity.
[0080] For imprecise counts such as "some" and "many", the 3D modeling system 102 can determine the count by examining the available 3D scene (e.g., in Figure 1in the content database 106 shown) and count the number of occurrences of two or more instances of the occurrence category to obtain a frequency histogram for each object category. For each count modifier, the 3D modeling system 102 can use this distribution to obtain the lower and upper bounds of the count implied by the modifier-category pairing, and then sample uniformly from this distribution. For example, the 3D modeling system 102 can sample between the 0th percentile and the 2nd percentile for "few", between the 10th percentile and the 50th percentile for "some", and between the 50th percentile and the 100th percentile for "many". The 3D modeling system 102 can infer that a plural noun without a modifier (e.g., "there are chairs around the table") has an implied "some" modifier. The 3D modeling system 102 replicates relationships, attributes, and determiners across each new instance of an object.
[0081] In addition, the 3D modeling system 102 identifies any qualifiers in the entity-command representation to determine the presence of multiple object nodes. For example, qualifiers such as "each" and "all" imply the presence of more than one object node for a particular entity. In one or more embodiments, the 3D modeling system 102 leaves the qualifier on a single object node until the 3D modeling system 102 uses the semantic scene graph 316 to generate a three-dimensional scene.
[0082] The 3D modeling system 102 also transforms the relationship information from the entity-command representation 302 into relationship nodes in the semantic scene graph 316. In particular, the 3D modeling system 102 can generate relationship nodes for each object in the object based on the relationships in the entity-command relationships 302 that are based on prepositions or other relationship information. For illustration, the 3D modeling system 102 can insert relationship nodes indicating "whether an object is above another object in the scene", "whether an object is to the left of another object in the scene", "whether an object is below another object in the scene", "whether an object is in front of another object in the scene", etc. Relationships that support more than one object are grouped together into a single relationship node within the semantic scene graph 316.
[0083] For illustration, Figure 3C is shown by targeting Figure 3BA single relationship node of the illustrated entity-command representation 302 is associated with multiple object nodes. Specifically, the semantic scene graph 316 includes a first object node 318 representing a second entity 304b (corresponding to the "desk" object) from the entity-command representation. Additionally, the semantic scene graph 316 includes multiple object nodes 320a to 320c representing a first entity 304a (corresponding to the "book" object) from the entity-command representation. The semantic scene graph 316 includes a relationship node 322 that indicates that the corresponding objects of the object nodes 320a to 320c are "on the corresponding object of the object node 318". Additionally, as shown, since all the books share the same relationship with the desk, the 3D modeling system 102 includes only one relationship node 322 in the semantic scene graph 316.
[0084] As previously described, the 3D modeling system 102 determines a count for each entity from the entity-command representation 302. Figure 3C Illustrated is that the 3D modeling system 102 infers the count of objects based on the "book" entity (the first entity 304a in the entity-command representation 302). Specifically, the 3D modeling system 102 analyzes the available scenes from the content database to determine the possible count range for the imprecise count "some". In this embodiment, the 3D modeling system 102 determines that "some" includes three book objects by uniformly sampling the 10th percentile to the 50th percentile of the frequency histogram of the available scenes. Accordingly, the 3D modeling system 102 generates three separate object nodes for the entity.
[0085] In addition to object nodes and relationship nodes, the semantic scene graph 316 also includes multiple edges that describe the relationships between each object node via the relationship nodes. Specifically, the edges indicate the directionality of the relationships between two object nodes from the entity-command representation 302. For example, the semantic scene graph 316 includes a first edge 324 representing the information that the desk is a lower object relative to one or more other objects. The semantic scene graph 316 also includes multiple edges 326a to 326c that represent the information that the book is a higher object relative to one or more other objects. Since the relationship node 322 indicates the "on" relationship, the combination of the first edge 324 and the multiple edges 326a to 326c with the relationship node 322 and the object nodes 318, 320a to 320c indicates to the 3D modeling system 102 that the books are on the desk within the three-dimensional scene.
[0086] In one or more embodiments, before generating a semantic scene graph for a natural language phrase, the 3D modeling system 102 first determines a canonical entity-command representation. As briefly described previously, a single conceptual request for generating a particular three-dimensional scene can be parsed in multiple different ways. For illustration, to generate a three-dimensional scene of a book on a desk, a user can say a first phrase "books are stacked on the desk", a second phrase "there is a stack of books on the desk", or a third phrase "stack the books on the desk". As illustrated by the entity-command representations 400 to 404 in Figures 4A to 4C , each phrase includes a unique syntax with different nouns, verbs, prepositions, etc., resulting in different entity-command representations. Specifically, the first entity-command representation 400 corresponds to the first phrase above, the second entity-command representation 402 corresponds to the second phrase above, and the third entity-command representation 404 corresponds to the third phrase above.
[0087] To create a consistent form of entity-command representation, the 3D modeling system 102 can select a particular form of entity-command representation as the canonical representation. In particular, the canonical representation allows the 3D modeling system 102 to determine a single entity command representation to be generated for each form of conceptual request. In one or more embodiments, if possible, the 3D modeling system 102 selects the descriptive form of entity-command representation (the first entity-command representation 400) as the canonical representation. As described herein, the descriptive form of entity-command representation is a representation in which the base noun is an object having a set of attributes that describe the base noun. If the user enters a request using a different form, the 3D modeling system 102 uses a set of pattern matching rules (e.g., similar to the pattern matching rules used to generate entity-command representations) to transform the resulting entity-command representation into the descriptive form.
[0088] In response to generating a semantic scene graph for a natural language phrase, the 3D modeling system 102 uses the semantic scene graph to generate a three-dimensional scene. In one or more embodiments, the 3D modeling system 102 identifies one or more available scenes that are most similar to the scene requested in the natural language phrase. For example, the 3D modeling system 102 compares the semantic scene graph with the semantic scene graphs for multiple available three-dimensional scenes to identify one or more scenes that are similar to the requested scene. For illustration, the 3D modeling system 102 can compare the semantic scene graph of the natural language phrase with each semantic scene graph of the available scenes by comparing the object nodes, relationship nodes, and edges in the semantic scene graph of the natural language phrase with the object nodes, relationship nodes, and edges in the semantic scene graph of the available scenes. A higher number of overlapping nodes and edges indicates a closer match, while a lower number of overlapping nodes and edges indicates a less close match.
[0089] Additionally, the 3D modeling system 102 can rank the available scenes by determining the degree of match between the corresponding semantic scene graph and the semantic scene graph of the natural language phrase. For example, the 3D modeling system 102 can assign a comparison score to each semantic scene graph of the available three-dimensional scenes based on the number and similarity of nodes and edges, as well as the semantic scene graph of the natural language phrase. Then, the 3D modeling system 102 can use the comparison scores to generate a ranked list of scenes, where a higher comparison score is higher in the ranked list of scenes. The 3D modeling system 102 can also use a threshold score to determine whether any of the available scenes in the available scenes are well-aligned with the natural language phrase.
[0090] Figure 5 Illustrated are multiple available three-dimensional scenes identified in the content database based on Figure 3C the semantic scene graph 316 shown. Specifically, the first scene 500 includes a table, and there is a book on the table. The second scene 502 includes a desk, and there are six books on the desk. The third scene 504 includes a bookshelf, and there are three books on the shelves inside the bookshelf. The 3D modeling system 102 compares the semantic scene graph 316 with the semantic scene graphs of each scene, and then selects at least one scene that best matches the expected request in the natural language phrase.
[0091] The 3D modeling system 102 selects the second scene 502 for generating the three-dimensional scene of the natural language phrase. The 3D modeling system selects the second scene 502 because the second scene 502 is the most similar to the semantic scene graph 316 (i.e., the second scene 502 includes a desk, and there are books on the desk). In particular, even though the second scene 502 has some differences based on the semantic scene graph of the second scene 502 (i.e., as in the semantic scene graph 316, there are six books instead of three books), the second scene 502 is still closer to the requested scene than the other available scenes (e.g., the books are on the desk instead of on the table or in the bookshelf). Additionally, since the requested scene includes an imprecise count (\"some\") of the book entity, the 3D modeling system 102 can determine that the second scene 502 is an acceptable deviation from the semantic scene graph 316. Additionally, the 3D modeling system 102 may need to share certain entities, commands, and / or other attributes to make the scene a close match.
[0092] As briefly mentioned previously, the 3D modeling system 102 can generate a 3D scene for a natural language phrase by inserting objects from a selected available 3D scene (or from multiple available 3D scenes) into the user's workspace. Specifically, the 3D modeling system 102 can select objects in the available 3D scene based on object identifiers corresponding to the objects. Then, the 3D modeling system 102 can copy the object and one or more attributes of the object, and paste / duplicate the object in the user's workspace. The 3D modeling system 102 can perform copy-paste operations on multiple objects in the selected available 3D scene until the 3D modeling system 102 determines that the requested 3D scene is complete. For illustration, the 3D modeling system 102 can copy the desk and books of the second scene 502 shown in FIG. 4, and paste the copied objects into the user's workspace at the user's client device.
[0093] In an alternative embodiment, the 3D modeling system 102 presents the selected 3D scene to the user to allow the user to insert objects into the workspace. For example, the 3D modeling system 102 can cause the user's client device to open a new workspace within the client application. Then, the user can select one or more objects in the client application and move the objects to a previously existing workspace or leave the objects in the new workspace to start a new project.
[0094] In one or more embodiments, the 3D modeling system 102 allows the user to enhance an existing 3D scene using natural language requests. Figures 6A to 6C Illustrated are multiple 3D scenes continuously generated using natural language phrases. As shown in Figure 6A , the 3D modeling system 102 generates a 3D scene 600 including a first set of objects 602. Specifically, the 3D modeling system 102 identifies a first natural language phrase including a request to generate the 3D scene 600 to include one or more objects. The 3D modeling system 102 parses the first natural language phrase to generate a semantic scene graph of the first phrase, and then uses the semantic scene graph to select one or more available 3D scenes including the first set of objects 602 to generate the 3D scene 600.
[0095] Additionally, the 3D modeling system 102 identifies a second natural language phrase including a request to enhance the 3D scene 600 with a second set of objects 604. The 3D modeling system 102 parses the second natural language phrase to create a semantic scene graph representing the second phrase. The 3D modeling system 102 uses the semantic scene graph of the second phrase to select one or more available 3D scenes including the second set of objects 604. As illustrated in Figure 6B , the 3D modeling system then enhances the 3D scene 600 by inserting the second set of objects 604 into the 3D scene.
[0096] Figure 6C Illustrated is a third set of objects 606 inserted into the three-dimensional scene 600. Specifically, the 3D modeling system 102 identifies, in a third natural language phrase, a request to further enhance the three-dimensional scene 600. The 3D modeling system 102 parses the third natural language phrase to create a corresponding semantic scene graph and then uses the semantic scene graph of the third phrase to select one or more available three-dimensional scenes including the third set of objects 606. Finally, the 3D modeling system 102 enhances the three-dimensional scene 600 by inserting the third set of objects 606 into the three-dimensional scene 600.
[0097] As shown in the example of Figures 6A to 6C , the 3D modeling system 102 can enhance a three-dimensional scene by inserting additional three-dimensional objects at specific locations within the scene. In one or more embodiments, the 3D modeling system 102 determines the location of the new object based on the information contained within the natural language phrase. For example, in response to natural language phrases stating "insert a couch", "place the TV in front of the couch", and "place the coffee table between the couch and the TV", the 3D modeling system 102 can insert the coffee table between the couch and the television. The 3D modeling system 102 can use additional intelligence to determine the location of the objects within the scene and, if necessary, can adjust the location of the objects already in the scene to accommodate the new object.
[0098] As described with respect to Figures 1 to 6C , the 3D system 102 can thus perform operations for processing natural language phrases to generate three-dimensional scenes. Figure 7 A detailed schematic diagram of an embodiment of the 3D system 102 described above is illustrated. As shown, the 3D modeling system 102 can be implemented within a content creation system 112 on one or more computing devices 700 (e.g., client devices and / or server devices described in Figure 1 and further described below with respect to Figure 9 ). Additionally, the 3D modeling system 102 can include, but is not limited to: a content manager 702, a communication manager 704, a natural language processor 706, a 3D scene generator 708, and a data storage manager 710. The 3D modeling system 102 can be implemented on any number of computing devices. For example, the 3D modeling system 102 can be implemented in a distributed system of server devices for generating three-dimensional scenes based on natural language phrases. Alternatively, the 3D modeling system 102 can be implemented on a single computing device, such as a single client device running a client application that processes natural language for generating three-dimensional scenes.
[0099] In one or more embodiments, each of the components of the 3D modeling system 102 communicates with other components using any suitable communication technology. Additionally, the components of the 3D modeling system 102 can communicate with one or more other devices, including: other computing devices of the user, server devices (e.g., cloud storage devices), license servers, or other devices / systems. It should be recognized that although the components of the 3D modeling system 102 are shown as being separate in Figure 7 it can be combined into fewer components, such as a single component, or divided into more components, as may be used for a particular implementation. Additionally, although Figure 7 the components shown are described as being connected to the 3D modeling system 102, at least some of the components used to perform operations in conjunction with the 3D modeling system 102 described herein can be implemented on other devices within the environment.
[0100] The components of the 3D modeling system 102 can include software, hardware, or both. For example, the components of the 3D modeling system 102 can include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (e.g., (a) computing device(s) 700). When executed by one or more processors, the computer-executable instructions of the 3D modeling system 102 can cause the (a) computing device(s) 700 to perform the three-dimensional generation operations described herein. Alternatively, the components of the 3D modeling system 102 can include hardware, such as a dedicated processing device for performing a particular function or group of functions. Additionally, or alternatively, the components of the 3D modeling system 102 can include a combination of computer-executable instructions and hardware.
[0101] Furthermore, for example, the components of the 3D modeling system 102 that perform the functions described herein for the 3D modeling system 102 can be implemented as part of a stand-alone application, a module of an application, a plug-in for an application including a marketplace application, one or more library functions that can be called by other applications, and / or a cloud computing model. Thus, the components of the 3D modeling system 102 can be implemented as part of a stand-alone application on a personal computing device or a mobile device. Alternatively or additionally, the components of the 3D modeling system 102 can be implemented in any application that allows for three-dimensional content generation, including but not limited to: CREATIVE FUSE, and Software. "ADOBE", "ADOBE FUSE", "CREATIVE CLOUD", "ADOBE FUSE", "ADOBE DIMENSION" and "ILLUSTRATOR" are registered trademarks of Adobe Inc. in the United States and / or other countries.
[0102] As mentioned, the 3D modeling system 102 includes a content manager 702 that facilitates the storage and management of three-dimensional content. Specifically, the content manager 702 can access a content database (e.g., Figure 1 the content database 106 shown), which includes a plurality of three-dimensional scenes previously created by any number of content creators. The content manager 702 allows content creators to store and share three-dimensional content across applications and devices and / or with other users. Additionally, the content manager 702 can store and / or manage assets for generating three-dimensional scenes.
[0103] The 3D modeling system 102 includes a communication manager 704 that facilitates communication between the 3D modeling system 102 and one or more computing devices and / or systems. For example, the communication manager 704 can facilitate communication with one or more client devices of a user to receive natural language requests for generating three-dimensional scenes. Additionally, the communication manager 704 can facilitate communication with client devices to provide three-dimensional scenes generated in response to natural language requests. Further, the communication manager 704 can allow client devices to store three-dimensional scenes on the content manager 702 for later use by one or more users.
[0104] The 3D modeling system 102 also includes a natural language processor 706 that facilitates the processing and analysis of natural language phrases. Specifically, the natural language processor 706 can use natural language processing techniques to generate a dependency tree, entity-command representation, and semantic scene graph of natural language phrases. The natural language processor 706 can also communicate with one or more other components (e.g., the 3D scene generator) to provide the semantic scene graph for comparison with available three-dimensional scenes in the content manager 702.
[0105] The 3D modeling system 102 also includes a 3D scene generator 708 that facilitates the generation of three-dimensional scenes. In particular, the 3D scene generator 708 can use the semantic scene graph from the natural language processor 706 of natural language phrases to identify one or more available three-dimensional scenes that the 3D scene generator can use to generate a three-dimensional scene based on the natural language phrase. The 3D scene generator 708 can also perform operations within the three-dimensional environment to insert, remove, manipulate, or otherwise modify three-dimensional content in the three-dimensional content scene.
[0106] The 3D modeling system 102 also includes a data storage manager 710 (which includes non-transitory computer memory) that stores and maintains data associated with processing natural language phrases and generating three-dimensional scenes. For example, the data storage manager 710 may store semantic scene graphs for natural language phrases and three-dimensional scenes. The data storage manager 710 may also store information for accessing and searching the content database in conjunction with the content manager 702.
[0107] Turning now to Figure 8 , the figure shows a flowchart of a series of actions 800 for synthesizing a three-dimensional scene using natural language. Although Figure 8 the figure illustrates actions according to one embodiment, alternative embodiments may omit any of the actions shown in Figure 8 , add any of the actions shown in Figure 8 , reorder any of the actions shown in Figure 8 , and / or modify any of the actions shown in Figure 8 . Figure 8 The actions in Figure 8 may be performed as part of a method. Alternatively, a non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause a computing device to perform the actions in Figure 8 . In other embodiments, the system may perform the actions in
[0108] As shown, the series of actions 800 includes an action 802 for analyzing a natural language phrase. For example, action 802 involves using natural language processing to analyze a natural language phrase that includes a request for generating a three-dimensional scene to determine dependencies involving one or more entities and one or more commands in the natural language phrase. Action 802 may involve receiving a text input that includes a natural language phrase or a voice input that includes a natural language phrase. Additionally, the natural language phrase may include multiple requests for generating multiple three-dimensional scenes. Action 802 may also involve receiving multiple natural language phrases.
[0109] This series of actions 800 also includes an action 804 of generating an entity-command representation of a natural language phrase. For example, action 804 involves using the determined dependencies between one or more entities and one or more commands to generate an entity-command representation of a natural language phrase. For example, action 804 may first involve generating a dependency tree that includes multiple tokens representing words in the natural language phrase and the dependency relationships corresponding to the multiple tokens. Then, action 804 may involve converting the dependency tree into: a list of entities including one or more entities annotated with one or more attributes and one or more relationships corresponding to the one or more entities; and a list of commands including one or more command verbs operating on the one or more entities.
[0110] Action 804 may also involve identifying one or more additional natural language phrases including a request to generate a three-dimensional scene in one or more different phrase forms. Action 804 may further involve generating a canonical entity-command representation for the one or more additional natural language phrases in one or more different phrase forms.
[0111] Action 804 may also involve determining that an object node includes an imprecise count corresponding to a base noun. Then, action 804 may involve determining a frequency histogram for an object category by analyzing a scene database to count the number of times two or more entity instances appear in the object category. Action 804 may further involve determining multiple objects to be included in a three-dimensional scene for an entity by sampling the distribution of count modifiers for the entity-command representation, where the distribution is determined based on the frequency histogram. For example, action 804 may involve determining multiple object nodes to be included in a semantic scene graph for a base noun by sampling the distribution determined based on the frequency histogram.
[0112] Additionally, this series of actions 800 includes an action 806 of generating a three-dimensional scene. For example, action 806 involves generating a three-dimensional scene by using at least one three-dimensional scene among multiple available three-dimensional scenes identified based on the entity-command representation of a natural language phrase. For example, action 806 may involve accessing a database including previously generated three-dimensional scenes that have corresponding semantic scene graphs representing the layout of objects within the previously generated three-dimensional scenes.
[0113] Action 806 may also involve converting an entity-command representation of a natural language phrase into a semantic scene graph that indicates the contextual relationship of one or more entities and one or more commands. For example, action 806 may involve mapping the base noun of an entity in one or more entities to an object node of an object category. Action 806 may also involve mapping the relationships corresponding to one or more entities to relationship nodes, where the edges indicate the direction of the relationships corresponding to one or more entities. Action 806 may involve generating edges that indicate the direction of the relationships corresponding to one or more entities.
[0114] Action 806 may also involve adding attributes or qualifiers as annotations to the object nodes of the base nouns. Additionally, action 806 may involve using the semantic scene graph to select a 3D scene from multiple available 3D scenes corresponding to the semantic scene graph.
[0115] Additionally, action 806 may involve comparing the semantic scene graph of a natural language phrase with the semantic scene graphs of multiple available 3D scenes. For example, action 806 may involve ranking multiple available 3D scenes based on the similarity between the semantic scene graph of the natural language phrase and the semantic scene graphs of the multiple available 3D scenes. Action 806 may also involve identifying a 3D scene from the multiple available 3D scenes that has a semantic scene graph that matches the semantic scene graph of the natural language phrase. For example, action 806 may involve selecting the 3D scene with the highest ranking from the multiple available 3D scenes.
[0116] As discussed in more detail below, embodiments of the present disclosure may include or utilize a special-purpose or general-purpose computer that includes computer hardware (such as, for example, one or more processors and system memory). Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be at least partially implemented as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (such as any of the media content access devices described herein). Generally, a processor (such as a microprocessor) receives instructions from a non-transitory computer-readable medium (such as memory) and executes those instructions to perform one or more processes, including one or more of the processes described herein.
[0117] A computer-readable medium can be any available medium that can be accessed by a general-purpose or special-purpose computer system. A computer-readable medium storing computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium carrying computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of the present disclosure can include at least two distinctly different computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0118] Non-transitory computer-readable storage media (devices) include: RAM, ROM, EEPROM, CD-ROM, solid state drives (“SSDs”) (e.g., based on RAM), flash memory, phase change memory (“PCM”), other types of memory, other optical disk storage devices, magnetic disk storage devices, or other magnetic storage devices, or any other medium that can be used to store computer-executable instructions or data structures in the form of desired program code components and that can be accessed by a general-purpose or special-purpose computer.
[0119] “Network” is defined as one or more data links that support the transportation of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer via a network or another communication connection (a hardwired connection, a wireless connection, or a combination of a hardwired connection and a wireless connection), the computer properly views the connection as a transmission medium. The transmission medium can include a network and / or a data link that can be used to carry computer-executable instructions or data structures in the form of desired program code components and that can be accessed by a general-purpose or special-purpose computer. The above combinations should also be included within the scope of computer-readable media.
[0120] Further, when program code components in the form of computer-executable instructions or data structures reach various computer system components, these program code components can be automatically transferred from the transmission medium to the non-transitory computer-readable storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., “NIC”) and then ultimately transferred to the computer system RAM and / or a less volatile computer storage medium (device) at the computer system. Thus, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize a transmission medium.
[0121] Computer-executable instructions include, for example, instructions and data that, when executed at a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a particular function or group of functions. In some embodiments, the computer-executable instructions are executed on a general-purpose computer to transform the general-purpose computer into a special-purpose computer implementing the elements in the present disclosure. The computer-executable instructions can be, for example, binary, intermediate format instructions (such as assembly language), or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or the acts described above. Instead, the described features and acts are disclosed as example forms for implementing the claims.
[0122] Those skilled in the art should understand that the present disclosure can be practiced in a network computing environment having many types of computer system configurations, including: personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablets, pagers, routers, switches, etc. The present disclosure can also be practiced in a distributed system environment in which local computer systems and remote computer systems linked through a network (through hardwired data links, wireless data links, or a combination of hardwired data links and wireless data links) both perform tasks. In a distributed system environment, program modules can be located in local memory storage devices and remote memory storage devices.
[0123] Embodiments of the present disclosure can also be implemented in a cloud computing environment. In this specification, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. A shared pool of configurable computing resources can be quickly provided via virtualization, and the shared pool of configurable computing resources can be released with less management effort or service provider interaction and then scaled accordingly.
[0124] A cloud computing model can consist of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, etc. The cloud computing model can also exhibit various service models such as, for example, Software as a Service (“SaaS”), Platform as a Service (“PaaS”), and Infrastructure as a Service (“IaaS”). Different deployment models (such as private cloud, community cloud, public cloud, hybrid cloud, etc.) can also be used to deploy the cloud computing model. In this specification and the claims, a “cloud computing environment” is an environment that employs cloud computing.
[0125] Figure 9 A block diagram of an exemplary computing device 900 that can be configured to perform one or more of the processes described above is illustrated. It should be understood that one or more computing devices (such as computing device 900) can implement the multi-RNN prediction system. As shown by Figure 9 illustrated, computing device 900 can include a processor 902, a memory 904, a storage device 906, an I / O interface 908, and a communication interface 910, which can be communicatively coupled via a communication infrastructure 912. In certain embodiments, computing device 900 can include fewer or more components than those shown in Figure 9 are shown. The components of computing device 900 shown in Figure 9 will now be described in more detail.
[0126] In one or more embodiments, processor 902 includes hardware for executing instructions (such as those that make up a computer program). By way of example and not limitation, to execute instructions for dynamically modifying a workflow, processor 902 can retrieve (or fetch) instructions from internal registers, internal caches, memory 904, or storage device 906 and decode and execute them. Memory 904 can be a volatile or non-volatile memory used to store data, metadata, and programs for execution by one or more processors. Storage device 906 includes storage means (such as a hard disk, flash drive, or other digital storage device) for storing data or instructions for performing the methods described herein.
[0127] The I / O interface 908 allows a user to provide input to the computing device 900, receive output from the computing device 900, and otherwise transfer data to and receive data from the computing device 900. The I / O interface 908 can include a mouse, keypad or keyboard, touch screen, camera, optical scanner, network interface, modem, other known I / O devices, or a combination of such I / O interfaces. The I / O interface 908 can include one or more devices for presenting output to a user, including but not limited to: a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 908 is configured to provide graphical data to a display for presentation to a user. The graphical data can represent one or more graphical user interfaces and / or any other graphical content that can serve a particular implementation.
[0128] The communication interface 910 can include hardware, software, or both. In any case, the communication interface 910 can provide one or more interfaces for communication between the computing device 900 and one or more other computing devices or networks (such as, for example, packet-based communication). By way of example and not limitation, the communication interface 910 can include: a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network (such as, WI-FI).
[0129] Additionally, the communication interface 910 can facilitate communication with various types of wired or wireless networks. The communication interface 910 can also facilitate communication using various communication protocols. The communication infrastructure 912 can also include hardware, software, or both that couples the components of the computing device 900 to each other. For example, the communication interface 910 can use one or more networks and / or protocols to enable multiple computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. By way of illustration, the digital content activity management process can allow multiple devices (such as, client devices and server devices) to exchange information using various communication networks and protocols for sharing information, such as, electronic messages, user interaction information, engagement metrics, or activity management resources.
[0130] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments of the present disclosure. Various embodiments and aspects of the present disclosure have been described with reference to the details discussed herein, and the drawings illustrate various embodiments. The above description and drawings are illustrative of the present disclosure and should not be construed as limiting the present disclosure. Numerous specific details have been described to provide a thorough understanding of the various embodiments of the present disclosure.
[0131] Without departing from the spirit or essential characteristics of the present disclosure, the present disclosure may be embodied in other specific forms. The described embodiments should be considered illustrative rather than restrictive in all respects. For example, the methods described herein may be performed using fewer or more steps / acts, or the steps / acts may be performed in a different order. Additionally, the steps / acts described herein may be repeated or performed in parallel with each other, or repeated or performed in parallel with different instances of the same or similar steps / acts. Accordingly, the scope of the present application is indicated by the appended claims rather than the foregoing description. All changes within the equivalent meaning and scope of the claims will be included within the scope of the claims.
Claims
1. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause a computing device to: Analyze a natural language request for generating a desired three-dimensional scene using natural language processing to determine dependencies of a plurality of entities involved in the natural language request and one or more commands operating on the plurality of entities; Generate a semantic scene graph indicating relative positioning of a plurality of objects in a scene based on a contextual relationship of the plurality of entities and the one or more commands indicated by the determined dependencies; Compare the semantic scene graph with a plurality of semantic scene graphs corresponding to a plurality of existing three-dimensional scenes to identify one or more three-dimensional scenes; And Generate a three-dimensional scene that satisfies the natural language request using the one or more three-dimensional scenes.
2. The non-transitory computer-readable medium according to claim 1, wherein the instructions, when executed by the at least one processor, cause the computing device to determine the dependencies by generating a dependency tree.
3. The non-transitory computer-readable medium according to claim 2, wherein generating the dependency tree includes setting a verb node as a center of the dependency tree and constructing the dependency tree based on noun dependencies associated with the verb node.
4. The non-transitory computer-readable medium according to claim 1, wherein the instructions, when executed by the at least one processor, cause the computing device to generate the semantic scene graph by: Classifying base nouns according to object categories; Determining entity counts, entity qualifiers, and relationships; And Constructing the semantic scene graph to include a plurality of object nodes, one or more relationship nodes, and edges connecting the plurality of object nodes.
5. The non-transitory computer-readable medium according to claim 4, wherein generating the semantic scene graph includes setting the edges to indicate a directionality of a spatial relationship between the plurality of object nodes.
6. The non-transitory computer-readable medium according to claim 1, wherein the instructions, when executed by the at least one processor, cause the computing device to compare the semantic scene graph with the plurality of semantic scene graphs corresponding to the plurality of existing three-dimensional scenes by generating a comparison score that indicates how similar a structure of the semantic scene graph is to structures of the plurality of semantic scene graphs.
7. The non-transitory computer-readable medium according to claim 1, wherein the instructions, when executed by the at least one processor, cause the computing device to generate the three-dimensional scene that satisfies the natural language request using the one or more three-dimensional scenes by loading a first three-dimensional scene among the one or more three-dimensional scenes.
8. The non-transitory computer-readable medium according to claim 7, wherein the instructions, when executed by the at least one processor, cause the computing device to generate the three-dimensional scene that satisfies the natural language request by inserting one or more three-dimensional objects from a second three-dimensional scene among the one or more three-dimensional scenes into the loaded first three-dimensional scene.
9. A system for synthesizing a three-dimensional scene using natural language in a digital media environment for three-dimensional computer modeling, comprising: A database of previously generated three-dimensional scenes; And At least one processor configured to cause the system to: Parse a natural language phrase including a request for generating a desired three-dimensional scene using natural language processing; Generate a semantic scene graph of the natural language phrase based on the parsing of the natural language phrase, the semantic scene graph including a plurality of object nodes, one or more relationship nodes, and edges connecting two or more of the plurality of object nodes via the one or more relationship nodes; Compare the semantic scene graph of the natural language phrase with a plurality of semantic scene graphs corresponding to the previously generated three-dimensional scenes to identify one or more three-dimensional scenes; And Generate a three-dimensional scene that satisfies the natural language phrase using the one or more three-dimensional scenes.
10. The system according to claim 9, wherein the at least one processor is further configured to cause the system to set the edges of the semantic scene graph to indicate the directionality of the spatial relationship between the two or more of the plurality of object nodes.
11. The system according to claim 9, wherein the at least one processor is further configured to cause the system to: Generate a ranking of the previously generated three-dimensional scenes based on the similarity between the semantic scene graph of the natural language phrase and the plurality of semantic scene graphs of the previously generated three-dimensional scenes; and Select the one or more three-dimensional scenes from the previously generated three-dimensional scenes based on the ranking.
12. The system according to claim 11, wherein the at least one processor is further configured to cause the system to generate the three-dimensional scene that satisfies the natural language phrase using the one or more three-dimensional scenes by loading the highest-ranked three-dimensional scene among the one or more three-dimensional scenes.
13. The system according to claim 9, wherein the at least one processor is further configured to cause the system to generate a dependency tree based on the parsing of the natural language phrase by setting a verb node as the center of the dependency tree and constructing the dependency tree based on noun dependencies related to the verb node.
14. The system according to claim 13, wherein the at least one processor is further configured to cause the system to generate an entity-command representation of the natural language phrase from the dependency tree, the entity-command representation including a list of entities annotated with corresponding attributes and relationships and a list of command verbs operating on the entities.
15. The system according to claim 14, wherein the at least one processor is further configured to cause the system to generate the semantic scene graph of the natural language phrase by converting the entity-command representation into the semantic scene graph.
16. A computer-implemented method for synthesizing a three-dimensional scene using natural language, comprising: Performing an analysis of a natural language request for generating a three-dimensional scene using natural language processing; Generate a semantic scene graph for the natural language request based on the analysis, the semantic scene graph including a plurality of object nodes, one or more relationship nodes, and edges connecting the plurality of object nodes; Compare the semantic scene graph of the natural language request with a plurality of semantic scene graphs corresponding to a plurality of existing three-dimensional scenes to identify one or more three-dimensional scenes; And Generate a three-dimensional scene that satisfies the natural language request using the one or more three-dimensional scenes.
17. The computer-implemented method according to claim 16, wherein generating the semantic scene graph of the natural language request includes setting the edges to indicate the directionality of the spatial relationships between the plurality of object nodes.
18. The computer-implemented method according to claim 16, wherein generating the three-dimensional scene that satisfies the natural language request includes loading a first three-dimensional scene that is most similar to the semantic scene graph of the natural language request among the one or more three-dimensional scenes.
19. The computer-implemented method according to claim 18, wherein generating the three-dimensional scene that satisfies the natural language request includes inserting one or more three-dimensional objects from a second three-dimensional scene among the one or more three-dimensional scenes into the loaded first three-dimensional scene.
20. The computer-implemented method according to claim 19, wherein inserting the one or more three-dimensional objects from the second three-dimensional scene among the one or more three-dimensional scenes into the loaded first three-dimensional scene is in response to receiving a second natural language request for enhancing the first three-dimensional scene.
Citation Information
Cited By
Natural language driven cross-modal three-dimensional scene generation method and system
CN121767566A