Space generation device, space generation method, and space generation program
The spatial generation device addresses the cost and effort issues of creating and modifying large-scale layouts by using an agent within the virtual space to autonomously execute actions based on input instructions, efficiently generating spaces that match user inputs.
Patent Information
- Application Number
- JP2023180065
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-19
- Publication Date
- 2025-05-02
AI Technical Summary
Existing technologies require repetitive user operations for characters and object placement/deletion, making it costly to create and modify large-scale layouts in virtual or real spaces.
A spatial generation device equipped with an input device, processor, and output device, utilizing an agent within the virtual space to autonomously execute actions based on current state and input instructions, efficiently generating a space that reflects the input instruction.
Enables efficient generation of spaces that match user input instructions, reducing the cost and effort required for creating and modifying large-scale layouts, and can be applied to both virtual and real spaces.
Smart Images

Figure 2025070037000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a space generating device, a space generating method, and a space generating program. [Background technology]
[0002] Conventionally, there is a technology described in JP 2022-097359 A (Patent Document 1) for placing objects in a virtual space. This publication states that "a program in one embodiment displays a first field of view corresponding to the position of a movable character in a virtual space on a display, and, in response to an input operation, places an object at a first target position related to the position of the character, or deletes an object already present at the first target position, determines a second target position based on the position of the character and the first target position, and sets the second field of view according to the second target position." [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2022-097359 A Summary of the Invention [Problem to be solved by the invention]
[0004] In the technology disclosed in Patent Document 1, the user must repeatedly operate the characters and place or delete objects, which results in huge costs when creating or modifying a large-scale layout. Therefore, the present invention aims to receive input instructions from a user regarding a desired layout and efficiently generate a space that reflects the input instructions. This problem is not limited to virtual spaces, but also occurs when recognizing real spaces and automatically changing the layout, etc. [Means for solving the problem]
[0005] A representative space generation device of the present invention comprises an input device that accepts conditions of a desired virtual space as input instructions, a processor that generates a virtual space that satisfies the input instructions, and an output device that outputs the virtual space generated by the processor, wherein the processor places an agent that exists with a position within the virtual space and serves as a virtual operating subject that changes the state of the virtual space, in the virtual space as an initial state, and generates a virtual space that satisfies the input instructions by repeating a process of determining an action of the agent based on the current state of the virtual space and the input instructions, and a process of having the agent execute the determined action to update the current state of the virtual space, wherein the action of the agent includes movement, which is an action that changes the position of the agent in the virtual space, and operation, which is an action that changes the state of the virtual space using the position of the agent as a base point. Furthermore, one representative space generation method of the present invention includes a space generation device having an input device that accepts conditions of a desired virtual space as input instructions, a processor that generates a virtual space that satisfies the input instructions, and an output device that outputs the virtual space generated by the processor, the space generation device including a step of accepting the input instructions, a step of placing an agent that exists with a position within the virtual space and serves as a virtual operating subject that changes the state of the virtual space, in the virtual space as an initial state, a step of determining an action of the agent based on the current state of the virtual space and the input instructions, and a step of having the agent execute the determined action to update the current state of the virtual space, thereby generating a virtual space that satisfies the input instructions, and a step of outputting a virtual space that satisfies the input instructions, wherein the action of the agent includes a movement that is an action that changes the position of the agent in the virtual space, and an operation that is an action that changes the state of the virtual space using the position of the agent as a base point. Furthermore, one representative space generation program of the present invention causes a computer having an input device that accepts conditions of a desired virtual space as input instructions, a processor that generates a virtual space that satisfies the input instructions, and an output device that outputs the virtual space generated by the processor to execute the following steps: accepting the input instructions; placing an agent that exists with a position within the virtual space and serves as a virtual operating subject that changes the state of the virtual space, in the virtual space as an initial state; determining an action of the agent based on the current state of the virtual space and the input instructions; and having the agent execute the determined action to update the current state of the virtual space, thereby generating a virtual space that satisfies the input instructions; and outputting a virtual space that satisfies the input instructions, wherein the action of the agent includes movement, which is an action that changes the position of the agent in the virtual space, and operation, which is an action that changes the state of the virtual space using the position of the agent as a base point. Effect of the Invention
[0006] According to the present invention, a space reflecting an input instruction can be efficiently generated. Problems, configurations and effects other than those described above will become apparent from the following description of the embodiment. [Brief description of the drawings]
[0007] [Figure 1] Schematic diagram of the space generating device [Diagram 2] Example of a flow chart executed by the space generating device [Diagram 3] Detailed example of the flow executed by the space generation device [Figure 4] Examples of inputs, outputs, and actions of space generating devices [Diagram 5] Example of spatial state given to action generation model [Figure 6] Schematic diagram of the object acquisition unit [Figure 7] Example of a method that projects pre-acquired point cloud information into the initial space [Figure 8]A schematic diagram of a space generator that holds multiple agents [Figure 9] Schematic diagram of a space generating device having a communication unit [Figure 10] Example of a user interface for a space generating device [Figure 11] Terminal diagram DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] Hereinafter, an embodiment will be described with reference to the drawings. In the description of the drawings, elements having the same functions are given the same numbers, and duplicated descriptions will be omitted. In addition, the following embodiment is one form for carrying out the present invention, and the present invention is not limited to this embodiment.
[0009] In this embodiment, the "space" refers to any extent in which objects can be arranged. The space in this embodiment is structured information on the position (coordinates) of each object in space. In this embodiment, the space is described using a "virtual three-dimensional space", but is not limited to this. For example, a virtual two-dimensional space (a planar space with no height) or a real three-dimensional space may also be the subject of the description.
[0010] In this embodiment, an "object" is an object that exists in space. An object in a virtual three-dimensional space is, for example, a 3D model, an object in a virtual two-dimensional space is, for example, a two-dimensional image, and an object in a real three-dimensional space is, for example, a physical object. An object may be associated with not only texture information (visual information consisting of shape and color) such as a 3D model or a two-dimensional image, but also additional information such as an explanatory document about the object or a program code for operating the object in the space. Intangible concepts that are not associated with texture information, such as the properties of the space (physical laws of the space, the area in which an agent can act, the brightness of the space, etc.), may also be treated as objects and set in the space during the course of the agent's actions.
[0011] In this embodiment, an "agent" is something that exists in the space handled by the space generating device and is operated by the space generating device to change the state of the space. When the space is a virtual space, it is, for example, a virtual robot (character, avatar) that acts in response to instructions in the virtual world. When the space is a real space, it is, for example, a physical robot that actually operates in the real space. Agents may also be a type of object and may be subject to generation and deletion. EXAMPLES
[0012] 1 is a diagram showing an overview of a space generating device 101. The space generating device 101 is composed of at least a control / recording unit 102, a behavior generation model 103, an object acquisition unit 104, and a space recording unit 105.
[0013] The space generating device 101 receives a space generation instruction and an initial space as input. The input instruction is an instruction generated by a user or an external system, which indicates what kind of space the space generating device 101 should ultimately create. The input instruction may be in the form of text, audio, image, video, or any other arbitrary form, or a combination of these. For example, by giving a text instruction such as "Please create a space with one house and trees on both sides" as an input instruction, the space generating device 101 can generate a space reflecting the instruction. For example, by using an image as an input instruction, the space generating device 101 can create a space reflecting the scenery of the image. For example, by giving text and an image as an input instruction, the space generating device 101 can create a space reflecting the scenery of the image while following the instructions of the text. In this embodiment, an example will be described using only text as an input instruction.
[0014] The initial space is a space that the space generating device 101 treats as the initial state of the space in which the operation is performed. The initial space may not be received as an input from the outside, but may be a default initial space that is set in advance in the space generating device 101. Any object may be placed in the initial space, or no object may be placed at all. An agent may be placed in the initial space (that is, the initial placement of the agent may be determined in advance), or may not be placed. If an agent is not placed in the initial space, the initial position of the agent may be determined by some method (such as a random position, or a position with as many objects as possible around it) when the space recording unit 105 is initialized. The properties of the initial space (physical laws of the space, the area in which the agent can act, the brightness of the space, etc.) may be specified in the input initial space, or may not be specified. For properties of the space that are not specified, the space generating device 101 may use properties that are set in advance, or the object acquiring unit 104 may generate the properties of the space as intangible objects and set them in the space.
[0015] The space generating device 101 outputs at least the final state of the space, that is, the state of the space recorded in the space recording unit at the time of the end of execution. The agents present in the space recorded by the space recording unit 105 may be deleted when the final space is output, or may be output as part of the final space.
[0016] An action history may be output as an additional output of the space generating device 101. The action history is a record of what action the agent performed for each step. By additionally outputting the action history, it is possible to use the action history as a reference for searching for improvements when the final space state is significantly different from the input instruction, or to use the space generating device 101 for a virtual space and to refer to the action history when a worker or robot acts in the real space.
[0017] The behavior history may include not only the behavior actually performed by the agent at each step, but also the input / output of each function of the space generation device 101 at each step, and an error message when an error occurs. For example, when the behavior generation model 103 outputs a reason for generating the next behavior, this may be included in the behavior history. For example, when the object acquisition unit 104 is used to acquire a new object, if an object corresponding to the object acquisition instruction cannot be acquired, an error message indicating that the object acquisition has failed may be included in the behavior history. For example, when an object is present in the direction in which the agent moved and the agent could not move, an error message indicating that the movement has failed may be included in the behavior history.
[0018] The control and recording unit 102 performs input and output with the behavior generation model 103, object acquisition unit 104, and space recording unit 105, and manages the overall operation of the space generating device 101 and records the output of each function.
[0019] The behavior generation model 103 receives an input from the control and recording unit 102 and outputs the next behavior to be taken by the agent. For example, a multimodal text generation AI model that receives text and image input and generates corresponding text may be used as the behavior generation model. When using an AI model as the behavior generation model, a pre-trained general-purpose AI model may be used, or an AI model trained by machine learning for this space generation device may be used. Whatever model is used, it is necessary to use a model that is suitable for the input / output requirements of the behavior generation model 103.
[0020] The object acquisition unit 104 receives an instruction for the next object to be placed, and selects an object that matches the instruction from a set of already created objects, or generates a new object and outputs it. The object acquisition unit 104 is called by the control and recording unit 102 when the next action generated by the action generation model 103 is the acquisition of a new object.
[0021] The space recording unit 105 records the state of the space, and when it receives the input of the next action from the control / recording unit 102, it operates the agent to change the state of the space and outputs the space state at that time. When the space generating device 101 receives the initial space, it initializes the space state of the space recording unit 105 with the initial space. The space recording unit 105 may additionally output the action result when the agent is operated according to the next action. For example, when the action can be completed as instructed by the input of the next action, it may additionally output a message indicating that the action was successful as the action result. For example, when the agent is unable to move because there is an object in the direction of movement or the agent is unable to move because the destination is an area set as a no-go area for the agent, it may additionally output a message indicating that the action has failed and the reason for the failure as the action result.
[0022] FIG. 2 is an example of a flowchart executed by the space generating device 101. Step 202: The space generating device 101 receives an input instruction and an initial space as an input. Step 203: Input instructions and the current spatial state are input to the action generation model 103 to obtain the next action. Step 204: If the next action output by the action generation model 103 is “end the action”, proceed to step 207. If it is not “end the action”, proceed to step 205. Step 205: If the next action output by the action generation model 103 is an action that requires acquisition of a new object, such as "place an object of xx in front of you," a new object is generated using the object acquisition unit 104. Step 206: The next action output by the action generation model 103 and, if a new object is acquired in step 205, the new object are used to operate the agent present in the space recording unit 105. Step 207: Output the space recorded in the space recording unit 105 at that time. As an additional output, a sequence of actions generated by the action generation model 103 may be output as an action history.
[0023] 3 shows in detail the flow performed by the space generating device 101, including an example of input and output. However, the control and recording unit 102 in the space generating device 101 is omitted.
[0024] First, the space generating device 101 initializes the state of the space recorded by the space recording unit 105 using the initial space received as an input (corresponding to step 202). After that, the following flow (1. to 3.) is treated as one processing group, and this is repeated. 1. The behavior generation model 103 receives a specified input and generates the next behavior to be taken by an agent present in the space recorded by the space recording unit 105 (corresponding to step 203). 2. If the generated action instruction is an action instruction for obtaining a new object, the corresponding object is obtained by the object obtaining unit 104 (corresponding to step 205). 3. The agent in the space acts according to the generated action instructions (corresponding to step 206).
[0025] If the action is successful, the state of the space changes according to the process groups 1 to 3. The process groups (1 to 3) are repeated until a space corresponding to the input instruction is generated, that is, until the action generation model 103 generates an action instruction of "end the action" as the next action.
[0026] In the above flow, if each function fails to perform a specified output and fails to operate, the spatial state recorded in the spatial recording unit 105 at that point in time may be output and execution may be terminated, or execution may be continued by transitioning to the beginning of the step (the input part to the behavior generation model, step 203). In the latter case, an error message for the part where the operation failed may be included in the behavior history (described later), which is one of the inputs to the behavior generation model 103, to change the generated behavior and prevent the same error from always occurring. Alternatively, the same error may be prevented from always occurring by introducing randomness into the operation of each function.
[0027] Instead of ending a series of actions when the next action output by the action generation model 103 is "end action", the action end may be determined by a separately prepared action end determination model. In other words, the current space state may be input to the action end determination model, and it may be determined whether or not the current space state reflects the input instruction.
[0028] Instead of terminating a series of actions by determining whether the current space state matches the input instruction using the behavior generation model 103, the action termination may be determined by combining other additional termination determination rules. For example, the action may be terminated and the space generation device 101 may output when the next action output by the behavior generation model 103 is "terminate the action" or the number of steps reaches 1000. For example, the action may be terminated and the space generation device 101 may output when the next action output by the behavior generation model 103 is "terminate the action" or the placement, movement, or deletion of an object has not been performed even once for 10 consecutive steps.
[0029] The behavior generation model 103 receives "input instructions" and "current spatial state" as inputs. It may also receive "options for next action", "options for objects that can be placed", and "action history" as additional inputs. The behavior generation model 103 outputs "next action". It may also output "reason for generating the next action".
[0030] An "input instruction", which is one of the inputs to the behavior generation model 103, is an instruction that serves as a guideline for generating the next behavior. It may be the input instruction received by the space generation device 101 itself, or it may be a processed version of the input instruction.
[0031] The "current space state", which is one of the inputs to the behavior generation model 103, is the space state recorded by the space recording unit 105 that has been processed in order to be provided to the behavior generation model 103. For example, a first-person perspective image of an agent existing in the space may be provided as the current space state.
[0032] "Next action options", which are one of the inputs to the behavior generation model 103, are options for the next action generated by the behavior generation model. Possible action options are, for example, "move forward / left / right / backward / up / down xx meters", "rotate xx degrees clockwise / counterclockwise", "place xx object in front", "delete object in front", "hold object in front", "place held object in front", "end action", etc. Here, "xx" is a variable, and the behavior generation model may freely set a value depending on the action to be taken. The behavior generation model 103 is expected to refer to this input option and generate an appropriate action.
[0033] Various restrictions can be imposed on the movement of agents. One of the restrictions is based on a virtual foothold within the virtual space. For example, restrictions can be set such as horizontal movement on the ground, not crossing steps of a certain height or higher, not entering areas where entry is prohibited, etc. One of the restrictions is a restriction that prohibits the overlap of the position of the agent with the position of a virtual object in the virtual space. This restriction is intended to avoid interference with other objects. It is also possible to set a restriction for the agent so that the size and shape in the virtual space are specified and the agent moves while avoiding contact with other objects. In addition, a restriction may be imposed such that when the agent moves while holding an object, the object held by the agent does not come into contact with other objects.
[0034] The options for the next action may be input from outside, or may use external information. For example, for the movement of the agent, the options may be selected from "move forward" and "rotate the agent." For the operation of an object by the agent, the options may be selected from "hold the object in front of the agent," "place the held object in front of the agent," and "rotate the object."
[0035] "Choice of objects that can be placed", which is one of the inputs to the behavior generation model 103, is a list of objects that the agent can place. For example, object candidates such as "house, tree, fence, post, flower, ..." may be specified. When the object acquisition unit 104 selects and outputs an object suitable for an instruction from a set of objects prepared in advance, by specifying choices of objects that can be placed in advance, it is possible to prevent the generation of a behavior to place an object that is not prepared in advance.
[0036] The "behavior history", which is one of the inputs to the behavior generation model 103, is a series of actions that the behavior generation model 103 has generated up to the present time. By providing the behavior history, it is expected that the agent can refer to what actions it has performed up to now and improve the accuracy of the next action it takes. The form of the behavior history output by the space generation device 101 and the form of the behavior history input to the behavior generation model 103 may be different. In other words, the information granularity of the behavior history may be different. For example, instead of including all actions from the start point to the present time, the number of actions given as input may be limited, such as using only history from a certain number of steps in the past from the present, or using only important actions. The behavior history may include not only the actions actually performed by the agent at each step, but also input / output of each function of the space generation device 101 at each step, and error messages in the event of an error.
[0037] The "next action", which is one of the outputs of the behavior generation model 103, is an instruction for the next action to be taken by the agent. The control and recording unit 102 receives this and has the agent on the space recording unit 105 perform this action, thereby changing the state of the space. The next action may be output in a natural language format, or a json format or the like may be used to make it easier to handle in later processing. For example, when using a json format, it may be output in a format such as {"action type":"movement", "direction":"forward", "distance":3}. The "next action option" to be input to the behavior generation model 103 may be input in the same format as the "next action". When generating an object as the next action, the object to be generated may be specified in a different form such as an image instead of text. The next action may be output as a combination of multiple action options input in the "next action option", and the agent on the space recording unit 105 may perform the action of the output. For example, in the case of moving forward and then placing an object, [{“action type”: “move”, “direction”: “forward”, “distance”: 3}, {“action type”: “obtain object”, “object”: “tree”}] may be output as the next action.
[0038] "Reason for generating next action", which is one of the outputs of the behavior generation model 103, is the reason why the behavior generation model 103 generated the next action to be performed by the agent. By providing the reason for generating the next action as part of the past action history to the input of the behavior generation model 103 at the next step, the behavior generation model can grasp the past action intention, and as a result, improvement in accuracy can be expected. Furthermore, by the space generation device 101 additionally outputting the reason for generating each action, the user can grasp the reason for performing the action of each step.
[0039] Figure 4 shows an example of an initial space that serves as input, a final space that serves as output, and the corresponding input instructions and the actions of an agent. For example, a text instruction (401) such as "Create a space with one house and trees on either side" is given as an input instruction, and a plane that resembles the ground and a space (402) with clouds floating in the air are given as the initial space. In response to this, an agent (404) placed in the space acts sequentially according to the actions generated by the action generation model, and generates the final space (403). An example of the order of actions at this time is shown in floor plan 405, and the order of actions is as follows: 1. "Place a house object in front of you", 2. "Move 3m to the left", 3. "Place a wooden object in front of you", 4. "Move 6m to the right", 5. "Place a wooden object in front of you", 6. “End the action.” In this case, when the behavior generation model as shown in FIG. 4 generates the behavior of "end the behavior", the agent ends the behavior, and the state of the space recorded in the space recording unit 105 at that time is output as the output of the space generation device 101. The above order of actions is one example, and even if the initial space, input instructions, and final space are the same, a different order of actions may be taken. For example, the agent may place a tree first without placing a house. Also, the agent may take an unnecessary action such as first moving backwards 1m and then moving forward 1m immediately afterwards.
[0040] In addition to the embodiment shown in FIG. 4, the following input instructions and initial space may be used. For example, to create a layout of a room, the initial space may be a room (a space covered with a floor, walls, and ceiling), and a text instruction such as "Create a layout of a 6-tatami room for a person living alone" may be given as an input instruction. For example, when creating a layout of a town, the initial space may be a space consisting of a certain area of ground and several roads, and the input instruction may be a text instruction such as "Create a layout of a town. The town must have more than 100 detached houses, one apartment building, and public facilities that are useful to the residents." For example, to correct the layout inside a factory, the initial space may be the existing layout of the factory, and the input instruction may be a text instruction such as "Find and correct unnecessary parts of the current layout. Move devices that can be improved by moving them a little, and delete unnecessary devices."
[0041] Fig. 5 shows an example of a method of expressing the current spatial state to be given to the behavior generation model 103. As shown in Fig. 5, the spatial state 501 recorded by the space recording unit 105 may be converted into a first-person viewpoint image 502 of the agent 404 and given to the behavior generation model 103.
[0042] As another method of expressing the spatial state, for example, an image of a third-person viewpoint of the agent (an image shot from a position slightly behind the agent so that the agent is included) may be used. For example, a plan view centered on the agent (an image shot from directly above the agent looking down on the agent and its surroundings) may be used. For example, the image representation of the spatial state obtained by the above method may be converted into a text form corresponding to the image using an image captioning model or an object detection model, and then provided to the behavior generation model 103. For example, if additional information such as document information is linked to each object included in the agent's field of view, the additional information may be provided to the behavior generation model 103. For example, the space within a certain distance from the agent may be converted into a three-dimensional point cloud form and provided to the behavior generation model 103. For example, a combination of multiple types of the above-mentioned methods of expressing the spatial state may be provided to the behavior generation model.
[0043] 6 is a schematic diagram of the object acquisition unit 104. The object acquisition unit 104 receives an object acquisition instruction and outputs an object corresponding to the instruction. The object acquisition unit 104 is composed of an object acquisition control unit 601, and either or both of an object recording unit 602 and an object generating unit 603.
[0044] The object acquisition control unit 601 performs input / output with an object recording unit 602 and an object generating unit 603 , and manages the overall operation of the object acquisition unit 104 .
[0045] The object recording unit 602 records a set of objects that can be placed in the space. When the space generating device 101 is initialized, an object created separately in advance may be registered. Also, an object created by the object generating unit 603 may be registered. The object recording unit 602 receives an object acquisition instruction, searches for an object corresponding to the instruction, and outputs an object that is most similar to the input instruction. If there is no object corresponding to the instruction, that is, if the similarity between the object acquisition instruction and all objects is equal to or less than a threshold value, an error message indicating that no corresponding object exists may be output instead of returning the object.
[0046] The object generation unit 603 generates an object corresponding to an object acquisition instruction. For example, the object generation unit 603 receives the text "house with red roof" as the object acquisition instruction and generates an object imitating a house with a red roof. For example, when a text-type instruction is received as the object acquisition instruction and a corresponding three-dimensional object is generated, a 3D model generation model may be used as the object generation unit. For example, when a text-type instruction is received as the object acquisition instruction and a corresponding two-dimensional object (i.e., a still image) is generated, an image generation model may be used as the object generation unit. For example, when a text-type instruction is received as the object acquisition instruction and a program code that specifies the texture of a 3D model and the behavior of the object is generated as an object, a combination of a 3D model generation model and a program code generation model may be used as the object generation unit.
[0047] When object acquisition unit 104 has only object recording unit 602, it acquires an object using object recording unit 602. When object acquisition unit 104 has only object generation unit 603, it acquires an object using object generation unit 603. When object acquisition unit 104 has both object recording unit 602 and object generation unit 603, it may acquire an object by combining them. For example, first, an object corresponding to the object acquisition instruction is searched for in object recording unit 602, and if it is registered, it is acquired and output. If it is not registered, it is acquired and output using object generation unit 603. At that time, the object generated by object generation unit 603 may be registered in object recording unit 602, and the recorded object may be reused when a similar object acquisition instruction is input in the next step or later.
[0048] FIG. 7 shows an example of a case where a method of projecting point cloud information acquired in advance onto an initial space is used. When a virtual three-dimensional space is treated as the space and the purpose is to reproduce the state of the real space on the virtual space, the state of the real space may be acquired in advance as a point cloud or the like and projected onto the initial space to be used as a reference when generating behavior. In other words, when acquiring the current space state from the space recording unit 105 in the form of an image of the first person viewpoint of the agent, the point cloud information projected onto the space may also be acquired as part of the image of the first person viewpoint. Since this point cloud information can be regarded as an example of the form of an input instruction, only the point cloud information may be given to the behavior generation model as an input instruction, or it may be given to the behavior generation model together with an input instruction in the form of text as shown in FIG. 7. In order to distinguish between an object actually placed in the space and the point cloud information that is merely projected, a method such as increasing the transparency of the point cloud compared to the object or changing the brightness or color of the point cloud may be used. Instead of the point cloud, another means for reading three-dimensional information of the real space and projecting it into the space in a three-dimensional form, such as NeRF or depth sensing camera data, may be used. Instead of using a means to read the three-dimensional information of the real space to be reproduced, for example, a two-dimensional plan view of the real space to be reproduced viewed from directly above may be obtained in advance and projected onto the floor of the virtual space as a reference for generating behavior. The above method may also be used to reproduce a different virtual space that already exists, rather than reproducing a real space.
[0049] FIG. 8 shows a schematic diagram of the space generation device 101 holding multiple agents. The space generation device 101 may hold multiple agents (agent A 803 and agent B 804 in FIG. 8), and each agent may act individually in the space recorded by the space recording unit 105. In other words, multiple agents may exist in one space recorded in the space recording unit 105, and the behavior generation model 103 may generate a different next action for each agent based on the behavior generation instruction corresponding to each agent, the current space state, and other inputs, and each agent may act according to the generated next action. In this case, the control / recording unit 102 may hold multiple recording units (a recording unit 801 for agent A and a recording unit 802 for agent B in FIG. 8) to separately record the behavior history of each agent. By giving the same behavior generation instruction to the behavior generation model 103, the efficiency of the work may be improved by performing labor-intensive work in parallel. Alternatively, by giving different behavior generation instructions to the behavior generation model 103, each agent may be assigned a different role and act. Agents may communicate with each other by sharing a part of their action history with other agents, or by generating additional messages for other agents by the action generation model and sharing them with other agents. The timing of the actions of the agents may be synchronous (performing actions at the same timing) or asynchronous (performing actions at different timings when the next action of each agent and, if necessary, the object to be placed are acquired). When the next action corresponding to at least one agent or all agents is "end the action", the space state recorded in the space recording unit 105 at that time may be output and the space generating device 101 may be terminated. In the latter case, only the agent that generated the action "end the action" may not perform the action after the next step, and may wait until all other agents have ended their actions. The number of agents held by the space recording unit 105 may be three or more, not two as shown in FIG. 8.When the behavior generation model 103 generates a behavior such as "placing an agent in front of the user as an object" in order to make work more efficient, a new agent may be generated, and the control and recording unit 102 may control and record the new agent as an entity independent of itself in the subsequent steps.
[0050] FIG. 9 shows a schematic diagram of the space generating device 101 having a communication unit 901. The space generating device 101 may search for necessary information from external knowledge such as the Internet and use it as additional input for each function. For example, when the behavior generation model receives an input instruction in the form of text, an image related to the input instruction may be searched for from external knowledge and provided as additional input to the behavior generation model. For example, when the object acquisition control unit 601 receives an object acquisition instruction and the object generating unit 603 generates a new object, an image of the object to be generated may be searched for from external knowledge and provided as additional input to the object generating unit 603. By acquiring information related to the input from external knowledge and using it as additional input for each function as described above, the accuracy of each function can be expected to be improved.
[0051] FIG. 10 is a diagram showing an example of a user interface of the space generating device 101. As in the above embodiment, the space generating device 101 does not only receive input and output after completion of execution, but the operating status of the space generating device 101 may be externally confirmed by the user interface during execution. For example, the behavior history including the output results of each function and the space state at that time recorded by the space recording unit 105 may be confirmed. If the output results of each function confirmed from the outside or the space state at that time are not intended by the user, the execution may be interrupted midway. In that case, the space state or the input instruction may be manually corrected, the behavior history may be inherited or initialized, and then the execution may be resumed.
[0052] Fig. 11 is a diagram showing an example of the hardware configuration of the space generating device 101. The space generating device 101 in Fig. 11 is a computer equipped with a processor 1101, a memory 1102, an auxiliary storage device 1103, an input device 1104, an output device 1105, and a communication device 1106. This is an example of the hardware configuration, and other configurations may be used.
[0053] For example, the processor 1101 deploys a space generation program in the memory 1102 and executes it to realize functions such as the control / recording unit 102 and the object acquisition unit 104. The auxiliary storage device 1103 stores the behavior generation model 103. The auxiliary storage device 1103 also operates as the space recording unit 105. The input device 1104 accepts inputs such as input instructions and initial spaces. The output device 1105 outputs the final space and the like. The communication device 1106 operates as the communication unit 901, and is also capable of receiving inputs and transmitting outputs.
[0054] In the above embodiment, the space is described as a "virtual three-dimensional space," the object is described as a "virtual three-dimensional object arranged in the virtual three-dimensional space," and the agent is described as a "virtual robot operating in the virtual three-dimensional space." As described above, the space may be a "virtual two-dimensional space" or a "real three-dimensional space."
[0055] For example, when dealing with a "virtual two-dimensional space" as the space, the object may be a "virtual two-dimensional object (i.e., a still image)" and the agent may be a "virtual robot that operates in the virtual two-dimensional space." In this case, it is possible to implement the same configuration and flow as the above embodiment.
[0056] For example, when dealing with a "real three-dimensional space" as the space, the object may be a "real three-dimensional object (i.e., an object that exists in reality)" and the agent may be a "physical robot that operates in the real three-dimensional space". In this case, a physical robot is placed in the real space, and the next action is generated by communicating with a computer in the robot or an external computer, and the robot acts accordingly. The initial space and final space input and output by the space generation device 101 are real spaces, not virtual spaces. When the action generation model 103 selects the generation of an object as the next action, for example, a 3D model of the object may be generated by the object acquisition unit 104, and then a pseudo 3D object may be created using a 3D printer mounted on the robot (agent) and placed in the place. The placed pseudo 3D object may later be replaced manually or automatically with a physical object to be actually placed. Alternatively, candidates for the physical object to be placed may be placed together in an object placement area at an appropriate location in the space, and when the behavior generation model 103 selects the generation of an object as the next action, the robot may move to the object placement area, pick up the object to be placed, and then return to the correct position to place the object.
[0057] If there are dynamic objects (for example, objects that move according to a certain rule, brightness of the space that changes at a certain cycle, a robot that behaves regardless of behavior instructions generated by the behavior generation model 103, etc.) in the space handled by the space recording unit 105, they may be made to move regardless of the behavior instructions generated by the behavior generation model 103. In that case, the movement of each object may be performed at the same timing as the behavior timing of the agent, or may be made to move regardless of the behavior timing of the agent. When dynamic objects are present, those objects may be fixed without moving while the space generation device 101 is running.
[0058] When the agent places a new object on the space recording unit 105, an object that can be obtained at low cost, such as a roughly shaped pseudo object, may be obtained and temporarily placed, and the object acquisition process may be performed by the object acquisition unit 104 in parallel with the subsequent steps, and upon completion of the acquisition, the temporarily placed pseudo object may be replaced with the finally acquired object. By using such a means, if it takes a long time for the object acquisition unit 104 to acquire an object, it is possible to improve the execution efficiency of the space generating device 101 by executing functions in parallel.
[0059] As described above, the space generating device 101 comprises an input device 1104 that accepts the conditions of the desired virtual space as input instructions, a processor 1101 that generates a virtual space that satisfies the input instructions, and an output device 1105 that outputs the virtual space generated by the processor. The processor 1101 places an agent, which has a position within the virtual space and acts as a virtual operating subject that changes the state of the virtual space, in the virtual space as an initial state, and generates a virtual space that satisfies the input instructions by repeating a process of determining an action of the agent based on the current state of the virtual space and the input instructions, and a process of having the agent execute the determined action to update the current state of the virtual space. The behavior of the agent includes movement, which is a behavior that changes the position of the agent in the virtual space, and manipulation, which is a behavior that changes the state of the virtual space with the position of the agent as a base point. With this configuration and operation, a space that reflects an input instruction can be efficiently generated.
[0060] In addition, the operation by the agent can be selected from one or more of placing a new object, which is a virtual object in the virtual space, deleting the object, the agent holding the object, and placing the object held by the agent. In this way, having an agent perform operations that mimic human work has the following advantages: It is possible to intuitively understand the changes in space caused by the agent. It is easy to evaluate whether the agent's actions are correct or not. The agent's behavior can be used as a reference for approaching the real space.
[0061] Furthermore, the processor 1101 sets a viewpoint based on the position of the agent, and compares the state of the virtual space seen from that viewpoint with the input instructions to determine the behavior of the agent. This makes it possible to efficiently generate a space that matches the input instruction without taking into account the state of the entire space. This advantage is particularly noticeable when the space is large or when there are a large number of objects.
[0062] Moreover, the output device 1105 outputs the history of the agent's actions. Therefore, for example, if there is a problem with the generated space, the behavior that caused it can be easily identified. In addition, when laying out the real space with reference to the agent's actions, the actions to be performed in the real space can be identified.
[0063] In addition, the processor 1101 imposes restrictions on the movement of the agent based on a virtual foothold within the virtual space, and a restriction that prohibits the position of the agent from overlapping with the position of an object, which is a virtual object in the virtual space. This makes it easier to intuitively understand the agent's behavior. Furthermore, when applied to layout in the real space, it is possible to obtain actions that can be realized in the real space.
[0064] The space in the initial state is, for example, a space generated corresponding to a real space. With this configuration, it is possible to simulate the real space and verify the actions taken in the real space. Also, the actions taken in the virtual space can be used to control the movement of a robot placed in the real space. In other words, an agent is made to act virtually in a virtual space generated corresponding to the real space, and if the desired result is obtained as a result, the robot is made to act in the real space in the same way as the action taken by the agent in the virtual space.
[0065] Furthermore, the processor 1101 can deploy a plurality of the agents and cause each agent to act individually. By operating multiple agents independently and reflecting them in the same virtual space, the desired virtual space can be generated more efficiently.
[0066] Moreover, the space generating device 101 further comprises a communication device 1106 for communicating with the outside, and the processor 1101 further uses information acquired from the outside via the communication device 1106 to determine the behavior of the agent. This makes it possible to generate a desired space with higher accuracy.
[0067] The present invention is not limited to the above-mentioned embodiment, and various modifications are included. For example, the above-mentioned embodiment is described in detail to easily explain the present invention, and is not necessarily limited to the embodiment having all the described configurations. Moreover, the present invention is not limited to the deletion of the configurations, and it is also possible to replace or add the configurations. [Explanation of symbols]
[0068] 101: space generation device, 102: control / recording unit, 103: behavior generation model, 104: object acquisition unit, 105: space recording unit, 106: communication unit, 1101: processor, 1102: memory, 1103: auxiliary storage device, 1104: input device, 1105: output device, 1106: communication device
Claims
1. an input device that receives input instructions indicating desired conditions for the virtual space; A processor for generating a virtual space that satisfies the input instructions; an output device that outputs the virtual space generated by the processor; The processor, An agent that exists in a virtual space and acts as a virtual operator that changes the state of the virtual space is placed in the virtual space as an initial state. a process of determining an action of the agent based on a current state of the virtual space and the input instruction, and a process of updating the current state of the virtual space by having the agent execute the determined action, thereby generating a virtual space that satisfies the input instruction; The action of the agent includes a movement which is an action for changing the position of the agent in the virtual space, and an operation which is an action for changing the state of the virtual space with the position of the agent as a base point. A space generating device characterized by:
2. The space generating device according to claim 1, A space generating device characterized in that the operation by the agent can be selected from one or more of placing a new object which is a virtual object in the virtual space, deleting the object, the agent holding the object, and placing an object held by the agent.
3. The space generating device according to claim 1, The processor, A space generating device, comprising: a viewpoint set based on the position of the agent; and a state of the virtual space as seen from the viewpoint compared with the input instructions to determine the behavior of the agent.
4. The space generating device according to claim 1, The output device is A space generating device that outputs a history of the actions of the agent.
5. The space generating device according to claim 1, The processor, A space generation device characterized by imposing restrictions on the movement of the agent based on a virtual foothold within the virtual space and a restriction that prohibits the position of the agent from overlapping with the position of an object that is a virtual object in the virtual space.
6. The space generating device according to claim 1, A space generating device, characterized in that the space in the initial state is a space generated corresponding to a real space.
7. The space generating device according to claim 1, The processor, A space generating device characterized in that a plurality of the agents are arranged and each agent is allowed to act individually.
8. The space generating device according to claim 1, Further comprising a communication device for communicating with the outside, The space generating device is characterized in that the processor determines the behavior of the agent by further using information acquired from the outside via the communication device.
9. a space generating device including an input device for receiving a condition of a desired virtual space as an input instruction, a processor for generating a virtual space that satisfies the input instruction, and an output device for outputting the virtual space generated by the processor; receiving the input instruction; A step of placing an agent, which exists in the virtual space and has a position therein and serves as a virtual operating subject that changes the state of the virtual space, in the virtual space as an initial state; a step of repeating a process of determining an action of the agent based on a current state of the virtual space and the input instruction, and a process of causing the agent to execute the determined action and updating the current state of the virtual space, thereby generating a virtual space that satisfies the input instruction; outputting a virtual space that satisfies the input instruction; Including, The action of the agent includes a movement which is an action for changing the position of the agent in the virtual space, and an operation which is an action for changing the state of the virtual space with the position of the agent as a base point. A space generating method comprising:
10. A computer including an input device that receives a condition of a desired virtual space as an input instruction, a processor that generates a virtual space that satisfies the input instruction, and an output device that outputs the virtual space generated by the processor, receiving the input instruction; A step of placing an agent, which exists in the virtual space and has a position therein and serves as a virtual operating subject that changes the state of the virtual space, in the virtual space as an initial state; a step of repeating a process of determining an action of the agent based on a current state of the virtual space and the input instruction, and a process of causing the agent to execute the determined action and updating the current state of the virtual space, thereby generating a virtual space that satisfies the input instruction; outputting a virtual space that satisfies the input instruction; Run the command, The action of the agent includes a movement which is an action for changing the position of the agent in the virtual space, and an operation which is an action for changing the state of the virtual space with the position of the agent as a base point. A space generation program comprising:
Citation Information
Patent Citations
program
JP2022097359A