Method for controlling robot device
By using large language models to predict human behavior and generate task plans, the method addresses the challenge of controlling robots in dynamic environments with moving agents, achieving reliable and cooperative interactions.
Patent Information
- Application Number
- JP2024192889
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-11-01
- Publication Date
- 2025-05-19
AI Technical Summary
Existing robot control methods struggle to reliably navigate and interact in dynamic environments with moving agents, such as humans, due to the inability to accurately predict and respond to their behavior.
A method involving the collection of state information about agents in the environment, conversion into text-based state descriptions, and utilization of large language models (LLMs) to predict agent behavior and generate task plans for the robot, enabling more reliable control and cooperation with humans.
This approach allows for more reliable and effective control of robots in dynamic environments by accurately predicting human behavior and generating appropriate task plans, enhancing cooperation between humans and robots.
Smart Images

Figure 2025078058000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for controlling a robot device.
Background Art
[0002] In many applications, it is desirable for a robot to be able to act autonomously in an environment where humans are present or in an environment where (other) obstacles exist. For this purpose, the robot must be able to take into account, in particular, the possible behavior of dynamic objects or "agents" such as humans in its environment during control. Thus, for example, an approach that enables more reliable automatic control of a robot device in an environment where agents such as humans move through the environment like objects and / or affect the state of each environment is desired.
Summary of the Invention
Means for Solving the Problems
[0003] According to various different embodiments, there is provided a method for controlling a robot device, the method comprising collecting state information regarding an agent located in the environment of the robot device, converting the state information regarding the agent into a state description in text form, supplying the state description in text form to a large language model for generating the behavior of the agent, the future state of the agent, and / or a prediction of the environment, generating a task plan for the robot device in consideration of the prediction, and controlling the robot device according to the task plan.
[0004] By the above method, in a dynamic environment, that is, in an environment where one or more agents are acting or interacting with one or more agents, that is, for example, in an environment where there are humans, it becomes possible to more reliably and better control a robot device. Thereby, in particular, cooperation between one or more humans and a robot becomes possible. According to various different embodiments, in this case, the control of the robot device (in particular, prediction of the state of the environment and task planning) (for example, in a dynamic environment with poor visibility) is implemented using one or more large language models (LLMs). Regarding LLMs, it has been found that LLMs can react like humans (for example, as a chatbot). Therefore, LLMs are also suitable for predicting human behavior.
[0005] Various different examples are presented below.
[0006] Example 1 is a method for controlling a robot device as described above.
[0007] Example 2 is the method described in Example 1, including training a large language model to predict the behavior of an agent from a state description in text form.
[0008] An LLM can be trained, for example, using deep learning (for example, by minimizing the difference between the target output from a training data set and the actual output from the model using an optimizer). Thereby, the LLM is adapted to be used for control itself. The LLM may be pre-trained.
[0009] Example 3 is the method described in Example 1 or 2, including creating a task plan by supplying a text-form target description, a text-form environment description, and a text-form description of the state of the robot device to an LLM or a further LLM ("LLM planning unit").
[0010] This may be a further LLM or may further instruct the same LLM appropriately (such that the same LLM takes into account the prediction). This enables efficient generation of a task plan.
[0011] Example 4 is the method according to Example 1 or 2, including creating a task plan by supplying a prediction (in text form as the output of the LLM), a target description in text form, an environment description in text form, and a description of the state of the robot device in text form to a further LLM (the "LLM planning unit").
[0012] By dividing into two LLMs, it becomes possible to fine-tune or train both LLMs for each task (e.g., by reinforcement learning, e.g., by reinforcement learning with human feedback).
[0013] Example 5 is the method according to Example 4, including training a further large language model to generate a task plan.
[0014] A further LLM can be trained, for example, using reinforcement learning (e.g., using the reward obtained for controlling the robot device using the task plan generated for this further LLM). Thereby, the further LLM is adapted to be used for control.
[0015] Example 6 is the method according to any one of Examples 1 to 5, including generating a scene graph from the collected state information and further information about the environment of the robot device, and generating a state description in text form from the scene graph.
[0016] This enables efficient generation of a state description in text form because, for example, the relationship between the agent and other objects (or also a room, etc.) represented by the scene graph (by the edges) can be directly converted into text.
[0017] Embodiment 7 is a control device (for example, a robot) configured to implement the method described in any one of Embodiments 1 to 6.
[0018] Embodiment 8 is a computer program comprising instructions for causing a processor to implement the method described in any one of Embodiments 1 to 6 when executed by the processor.
[0019] Embodiment 9 is a computer-readable medium storing instructions for causing a processor to implement the method described in any one of Embodiments 1 to 6 when executed by the processor.
[0020] In the drawings, like reference numerals generally relate to the same parts throughout all the different illustrations. The drawings are not necessarily to scale; instead, emphasis is generally placed on illustrating the principles of the invention. In the following specification, various aspects are described with reference to the following drawings.
Brief Description of the Drawings
[0021]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Mode for Carrying Out the Invention
[0022] The following detailed description refers to the accompanying drawings, which are shown for the purpose of illustrating specific details and aspects of the disclosure in which the invention can be practiced. Other aspects may be used and structural, logical, and electrical changes may be made without departing from the scope of the invention. The various different aspects of the disclosure are not necessarily mutually exclusive, because some aspects of the disclosure can be combined with one or more other aspects of the disclosure to form new aspects.
[0023] In the following, various different examples will be described in more detail.
[0024] Figure 1 shows a control scenario.
[0025] The mobile robot 100 is located within an environment 101 (e.g., a factory hall or a construction site). The robot 100 has a starting position 102 and is required to reach a target position 103. Within the environment 101, there are obstacles 104 to be bypassed by the robot 100. The obstacles 104 should not be passed through by the robot 100 (e.g., machines, walls, or trees), or should be avoided, because the robot may damage or injure the obstacles 104 (e.g., within a factory hall where the robot 100 is used to transport goods, such as a human).
[0026] The robot 100 includes a control device 105 (the control device 105 may be spatially separated from the robot 100, i.e., the robot 100 may be remotely controlled). In the exemplary scenario of Figure 1, the goal is for the control device 105 to control the robot 100 to navigate from the starting position 102 through the environment 101 to the target position 103. The robot 100 may be, for example, an autonomous vehicle, but may also be a robot equipped with legs or caterpillars, or other types of drive systems (e.g., deep-sea exploration vehicles or Mars exploration vehicles).
[0027] Furthermore, the embodiment is not limited to the scenario where the robot should be moved (as a whole) between positions 102 and 103, and the embodiment can also be used for controlling a robot arm having an end effector that should be moved between positions 102 and 103 (without colliding with the obstacle 104), etc.
[0028] Therefore, in the following, terms such as robot, vehicle, machine, etc. are used as examples of the "object" or agent to be controlled, i.e., a computer-controlled technical system (e.g., a machine). The approach described herein may be applied in various different types of computer-controlled machines such as robots or vehicles and others (and in some cases their environments). The general term "robot device" is used hereinafter for all types of technical systems (mobile technical systems and / or technical systems having one or more movable components) that can be controlled by the approach described hereinafter.
[0029] In an ideal case, the control device 105 learns a control strategy that enables the control device 105 to successfully control the robot 100 (from the starting position 102 to the target position 103 without colliding with the obstacle 104) for any scenario (i.e., environment, starting position, and target position) in a particular scenario that the control device 105 has not yet encountered.
[0030] The control device 105 is appropriately trained for this purpose. For such training, scenarios (especially the environment 101) can be simulated, but the scenarios are usually real during use. However, the robot 100 and the environment can still be simulated during use (e.g., when the simulation of a vehicle (by a control strategy) is used to test other control strategies of an autonomous vehicle).
[0031] According to various embodiments, a control device (e.g., control device 105) for a robotic device is provided, and this control device uses a large language model (LLM) to predict the behavior (or the result of the behavior) of an agent within the environment of the robotic device. In that case, this prediction can be made by a task planning module of the control device for task planning. For example, in order to avoid a collision between a human and a robot, the behavior of a human within an indoor environment, e.g., where the human will move, is predicted.
[0032] In recent years, in the research of natural language processing (NLP) and the research of generative artificial intelligence (AI), various different high-performance LLMs that can create text like that written by a human have been developed with great progress.
[0033] According to various embodiments, one or more LLMs are utilized to predict the future activities or behavior patterns of an agent (e.g., a human). Thereby, for example, the robot 100 can be navigated through an environment with poor visibility with (optionally a number of) humans, i.e., human "obstacles" 104. Because the LLM can learn "common sense" and thus has been found to be also suitable for predicting human behavior. In the following, one or more humans are taken as examples, but the behavior of other "agents" can be predicted in the same way, e.g., like an animal (e.g., when navigating a robot within a cowshed).
[0034] According to various embodiments, for example, a prediction and planning module for a control device of a mobile robot within an indoor environment crowded with many humans is provided, and this prediction and planning module takes the following inputs (information), i.e., · A description of the type of environment · A description of the task · Sensor data regarding the current position of the robot · For example, a scene graph with all objects in the environment and their relationships, including one or more agents (e.g., humans). Receive it.
[0035] The prediction and task planning module uses one or more LLMs to predict human behavior and create (high-level) task plans in order to achieve the goals (i.e., tasks) defined by the user and, at the same time, minimize certain costs (e.g., driving routes, costs of human-robot interactions, etc.).
[0036] Figure 2 shows the data flow for controlling a robot device (e.g., an autonomous system such as a self-driving robot) according to an embodiment in each control device.
[0037] Cognition (or perception) 201 provides information about the environment of the robot device (in the form of a scene graph according to various embodiments). From this information, a prediction 202 about the future state of the environment is made. This prediction 202 is used for task planning 203. When the execution of a task is planned (e.g., first move to point A, wait for 5 seconds, and then move to point B), a movement plan 204 is made accordingly (i.e., the trajectories for each movement according to the task plan are specified). Then, the control 205 of the robot device is carried out according to this movement plan.
[0038] The functions of task planning 203 and movement planning 204 can be integrated into the function of "planning", which can be made the function of planning module 206. That is, task planning 203 (which can be regarded as a "high-level" plan) can be regarded as part of the planning level of each autonomous system (at least mainly). In this specification, the functions of prediction 202 and task planning 203 are integrated into the function of "prediction and task planning module" 207.
[0039] According to various embodiments, the perception 201 provides information in the form of a scene graph, which includes all objects (static and dynamic objects, i.e., including robots and humans) in the environment as nodes, and includes the relationships between the objects as edges.
[0040] FIG. 3 shows an example of a scene graph 300 for a control scenario in which a mobile robot is to navigate through one floor of a building.
[0041] The scene graph 300 has a tree form in this example. The root 301 of the scene graph is associated with the floor in this example. For each room 302 of the floor, the scene graph 300 includes an internal (i.e., neither the root nor the leaf) node 303 associated with that respective room within the scene graph 300. Each leaf (i.e., each leaf node) 304 of the scene graph 300 is associated with a respective object 305 (or also with a respective location within the room) and is connected to a node 303 associated with the room in which the respective object 305 is located.
[0042] Each node (root, internal node, and leaf) may have one or more properties (also referred to as attributes). This property may be a binary property. For example, in the case of an internal node, it may be "Objekt_im_Raum=wahr (Object in the room=true)", "Roboter_im_Raum=wahr (Robot in the room=true)". Or in the case of a leaf, it may be "ist_zu_oeffnen=falsch (~should be opened=false)", "ist_aufnehmbar=wahr (~can be picked up=true)". Also, this property may be a non-binary attribute. For example, in the case of a leaf, it may be "Kategorie=Bett (Category=Bed)", "Gewicht=10kg (Weight=10kg)". Or in the case of an internal node, it may be "Raumtyp=Vorratsraum (Room type=Storage room)", the position and / or speed of an object, etc. Or in the case of the root, it may be "Stockwerkstyp=Keller (Floor type=Basement)". One or more of the objects are, for example, agents, i.e., in this example, humans (these objects are represented, for example, by the binary property "ist_Mensch=wahr (~is a human=true)" or generally "ist_dynamisches_Objekt=wahr (~is a dynamic object=true)"). The robot itself can also be represented by a leaf. The node associated with an object (object node) may be a human node, a robot node, or an environment node according to what those objects represent (humans, robots, or other objects in the environment).
[0043] The scene graph 300 can also have yet another level. For example, a room can first be subdivided into a plurality of locations (each of these locations is associated with one node each that is linked to a respective room node), and then objects can be specified for these locations. Relationships can also be established between objects (for example, a box is placed inside a shelf). In that case, an edge can be created to connect both of these object nodes to describe the relationship. If an agent is located in the vicinity (in terms of Euclidean distance) of an object, or if the agent is directly related to the object (for example, a human is sitting on a chair), then both of these corresponding nodes can also be linked to the edge (however, for example, there is no edge between two agent nodes). In that case, in some cases, the scene graph may no longer have a tree shape.
[0044] According to various different embodiments, in prediction 202, an LLM is used to predict human behavior. This prediction is provided as an input for task planning 203, and task planning 203 then outputs a task plan using another LLM for motion planning 204 to calculate a motion trajectory for the robot. Control 205 calls the corresponding actuator to follow the motion trajectory (for example, to move an arm, a wheel, etc.).
[0045] For example, this is used in a factory scenario for assembling an e-bike, in which case the robot 100 is required to collect various different parts of the e-bike, such as a steering wheel, a rear wheel, a frame, a battery, and a saddle, from various different locations and transport them one after another to the assembly area without passing through the area where humans are present (that is, without colliding with humans or disturbing humans).
[0046] The task planning problem is given by the tuple (O, P, A, T, C, I, G). O is the set of all basic objects of the problem. P is the set of properties (or attributes) defined by one or more objects respectively (binary properties are a subclass of attributes with boolean values). A is a finite set of actions taken on object tuples. T is a state transition model, and C represents the cost of state transition. I is the initial state, and G represents one or more target states. A state is an assignment of values to all possible properties of an object. A state is described, for example, by a scene graph as described above (at least partially, for example, information about the robot state (such as joint positions) can also be provided as additional information).
[0047] A language model (LM: language model in English) can be regarded as a distribution over sequences of word tokens. One approach is to model the conditional distribution p(w 1:k at the k-th position in the sequence of previous input tokens w k over the next token w k | w 1:k ). Relatively recent LMs are based on the Transformer architecture, which has significantly extended the LM to bring about large language models (LLMs: large language models in English), that is, large-scale language models having basically tens or hundreds of billions of parameters and capable of generating text or programming code like that written by humans.
[0048] Figure 4 shows the functions of the prediction and task planning module 400 according to one embodiment.
[0049] The prediction and task planning module 400 includes an "SG2NL" module 401 that converts the scene graph (SG) 402 into a description in natural language (English: natural language) of the scene (i.e., the state) described by this scene graph (for example, a description of the properties or states of the objects associated with the leaves in words), an LLM prediction unit "LLM Praed " 403 that predicts human behavior, and an LLM planning unit "LLM Plan " 404 that creates a task plan for the robot.
[0050] The inputs to the prediction and task planning module 400 are the scene graph 402 and the goal description G in natural language defined by the user NL and the output of the prediction and task planning module 400 is the generated task plan π. As described above, the scene graph 402 consists of a plurality of nodes and edges as shown in FIG. 3. That is, some of the nodes (for example, the leaves) represent the objects and agents (for example, humans and the robot itself) in the environment, and the nodes (including the leaves) can have a plurality of attributes, such as the position, orientation, actions, etc. of the human or robot.
[0051] Algorithm 1 shows an example in pseudo-code for creating a task plan from the scene graph for the goal G NL (in this case, common English keywords such as while, do, and end are used).
[0052]
Table 1
[0053] In line 2 of Algorithm 1, the human state S H of one or more humans, the environmental state S E and the robot state S R are identified from the scene graph SG, where S H and S Eare two sets of attributes of human nodes or environmental nodes (and also for other objects and rooms), S R is a set of attributes of robot nodes.
[0054] In line 3, the scene graph 402 is analyzed by the SG2NL module 401. The SG2NL module 401 converts the scene graph 402 into NL descriptions of the human state S H NL , the environmental state S E NL , and the robot state S R NL , where S H NL , S E NL , and S R NL are texts in the form of sentences in natural language.
[0055] Then, S H NL is sent to the LLM prediction unit 403, and the LLM prediction unit 403 generates a prediction of future human behavior in natural language that associates the predicted human behavior with the actual human state
Number
[0056] The description S of the environmental state E NL is sent to the LLM planning unit 404 together with the goal description G NL , the description S of the robot state in natural language R NL , and the prediction of the human state in natural language
Number
[0057] The task plan for putting an apple in the refrigerator may be written, for example, as "Go to the refrigerator, wait until the person leaves the area of the refrigerator, open the refrigerator, put the apple in the refrigerator, and close the refrigerator."
[0058] The LLM prediction unit 403 and the LLM planning unit 404 may be the same pre-trained LLM, but different instructions are given to this same LLM according to each case.
[0059] Next, the task plan π is executed by the robot, and the robot interacts with the environment 405, and the next human state S’ H and the next environmental state S’ E and the next robot state S’ R are observed (line 6). Then, the scene graph is correspondingly updated by the state change (line 7). This process is repeated until the goal defined by the user is reached.
[0060] Algorithm 2 shows an example of converting the scene graph SG into natural language by the SG2NL module (here too, general English keywords such as for, do, end, if, else, then are used).
[0061]
Table 2
[0062] First, the human state S H NL , the environmental state S E NL and the robot state S R NLThe NL description is initialized as blank (line 1). From line 2 to line 10, all nodes in the scene graph are looped through. For each node, access is made to all edges (i.e., relationships) and adjacent nodes (i.e., adjacent objects and humans) (lines 3 - 4). If the node is a human node, the current state of the human (position, orientation, and actions), and the relationships between this human and all adjacent objects and humans are converted to text and added to S H NL (lines 5 - 7). If the node is an environment node representing other objects, the position of the object and the relationships between this object and all adjacent objects (including humans) are recorded in S E NL (lines 8 - 10). Otherwise, the relationships between all adjacent objects and humans of the robot node are added to S R NL (line 12).
[0063] In summary, according to various different embodiments, a method as shown in FIG. 5 is provided.
[0064] FIG. 5 shows a flowchart 500 illustrating a method for controlling a robot device according to an embodiment.
[0065] At 501, state information regarding agents (e.g., humans or animals, i.e., "dynamic objects" that move within the environment and in some cases act on the environment, e.g., a human can remove other objects from the environment or carry them to other locations) located within the environment of the robot device is collected.
[0066] At 502, the state information regarding the agents is converted into a state description in text format.
[0067] In 503, a state description in text form is supplied to an LLM (a "LLM prediction unit") for generating an agent's behavior, the agent's future state, and / or an environment prediction (the LLM can output, for example, an agent's behavior or the result of the behavior in view of the state of the environment (including the agent itself)). In other words, a state description in text form is supplied to the (trained) LLM to generate a prediction.
[0068] In 504, a task plan for the robot device is generated in consideration of the prediction.
[0069] In 505, the robot device is controlled according to the task plan.
[0070] The method of FIG. 5 may be implemented by one or more computers having one or more data processing units. The term "data processing unit" may be understood as any kind of entity that enables the processing of data or signals. The data or signals may be processed, for example, according to at least one (i.e., one or more than one) special function performed by the data processing unit. The data processing unit may include an analog circuit, a digital circuit, a logic circuit, a microprocessor, a microcontroller, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an integrated circuit of a programmable gate array (FPGA), or any combination thereof, or may be composed of these. Any other method for implementing each function described in detail herein may also be understood as a data processing unit or a logic circuit device. One or more of the method steps described in detail herein can be implemented (e.g., implemented) by the data processing unit by one or more special functions performed by the data processing unit.
[0071] That is, according to various different embodiments, the method is particularly computer-implemented.
[0072] Controlling a robotic device in accordance with a task plan includes generating one or more control signals for the robotic device. The term "robotic device" may be understood to relate to any technical system (comprising a mechanical part whose operation is controlled), such as a computer-controlled machine, a vehicle, a household appliance, a power tool, a manufacturing machine, a personal assistant, or an access control system.
[0073] Various different embodiments can receive and use sensor signals from various different sensors, such as video, radar, LiDAR, ultrasonic, motion, heat generation, etc., for example, to obtain sensor data regarding state information. These sensor data can be processed to obtain state information (i.e., for perception). This can include, for example, classifying the sensor data or performing semantic segmentation on the sensor data to detect the presence of an object (in the environment in which the sensor data was acquired).
Claims
1. A method for controlling a robotic device (100), the method comprising: Collecting (501) state information about agents located within an environment (101) of the robotic device (100); converting (502) state information about the agent into a textual state description; feeding (503) the textual state description to a large-scale language model (403) for generating predictions of the agent's behavior, future states of the agent, and / or the environment (101); generating (504) a task plan for the robotic device (100) taking into account the prediction; and controlling (505) the robotic device (100) according to the task plan; A method comprising:
2. training said large-scale language model (403) to predict the behavior of said agent from said textual state description; The method of claim 1 , comprising:
3. creating said task plan by feeding said large-scale language model or a further large-scale language model (404) with a textual goal description, a textual environment description, and a textual description of a state of said robotic device (100); The method of claim 1 or 2, comprising:
4. creating said task plan by feeding said prediction, a textual goal description, a textual environment description, and a textual description of a state of said robotic device (100) to a further large-scale language model (404); The method of claim 1 or 2, comprising:
5. training the further large language model (404) to generate a task plan; The method of claim 4 , comprising:
6. generating a scene graph (402) from the collected state information and further information about the environment (101) of the robotic device (100); generating said textual state description from said scene graph; The method of any one of claims 1 to 5, comprising:
7. A control device configured to carry out the method according to any one of claims 1 to 6.
8. A computer program comprising instructions which, when executed by a processor, cause the processor to carry out a method according to any one of claims 1 to 6.
9. A computer readable medium storing instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 6.