Method for controlling robotic device

By using large language models to predict the behavior of dynamic agents and the future state of the environment, and generate task planning for robot equipment, the problem of safely controlling robot equipment in a dynamic environment is solved, and collaboration between robots and humans is realized.

CN119937361APending Publication Date: 2025-05-06ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411536918.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-10-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to safely and successfully control robotic devices in the presence of dynamic agents, such as humans, environments, especially in dynamic environments, where robots need to consider possible behaviors of humans or other agents.

Method used

By collecting agent status information in the environment, converting it into text state description, using the Large Language Model (LLM) to predict the behavior of the agent and the future state of the environment, and on this basis generate task planning for the robot equipment, ultimately realizing safe and automatic control of the robot equipment.

Benefits of technology

It is possible to achieve safe and successful control of robotic devices in a dynamic environment, especially in the presence of humans or other agents, to enable collaboration between robots and humans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937361A_ABST
    Figure CN119937361A_ABST
Patent Text Reader

Abstract

According to various embodiments, a method for controlling a robotic device is provided, including collecting state information about an agent located in an environment of the robotic device; converting the state information about the agent into a text state description; delivering the textual state description to a large language model to produce a prediction of the behavior of the agent, the future state of the agent, and / or the future state of the environment; generating a task plan for the robotic device taking into account the prediction; and controlling the robotic device according to the task plan.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to methods for controlling a robotic device. Background Art

[0002] In many applications it is desirable that a robot be able to act autonomously in an environment where people are present or (other) obstacles are present. To this end, the robot is in particular able to take into account possible behavior of dynamic objects or "agents" in its environment, such as people, when controlling it. Therefore, it is desirable to have a method that enables safe automatic control of a robotic device in an environment where agents, such as people, are present, which move in the environment and / or influence the state of the corresponding environment, such as causing objects to move. Summary of the invention

[0003] According to various embodiments, a method for controlling a robotic device is provided, comprising collecting state information about an agent located in an environment of the robotic device; converting the state information about the agent into a textual state description; feeding the textual state description into a large language model to generate predictions of the agent's behavior, the agent's future state, and / or the future state of the environment; generating a task plan for the robotic device taking into account the predictions; and controlling the robotic device according to the task plan.

[0004] The method described above enables safe and successful control of a robotic device in a dynamic environment, i.e. in an environment in which one or more agents act or interact, i.e. in which, for example, a person is present. This makes it possible, in particular, for a collaboration between one or more people and a robot to be achieved. According to various embodiments, the robotic device (e.g. in a chaotic and dynamic environment) is controlled (in particular, the prediction of environmental states and task planning) with the aid of one or more large language models (LLMs). It has been shown that the LLM is able to react in a human manner (e.g. as a chatbot). The LLM is therefore also suitable for predicting human behavior.

[0005] Various embodiments are described below.

[0006] Embodiment 1 is a method for controlling a robot apparatus as described above.

[0007] Embodiment 2 is a method according to embodiment 1, comprising training the large language model for predicting agent behavior from textual state descriptions.

[0008] For example, the LLM can be trained using deep learning (for example by minimizing the difference between the nominal output from the training data set and the actual output from the model by means of an optimizer). The LLM is thus adapted to its use for control. The LLM can also be pre-trained.

[0009] Embodiment 3 is a method according to embodiment 1 or 2, comprising generating a task plan by feeding a textual target description, a textual environment description and a textual description of the state of the robotic device to the large language model or another large language model ("LLM planner").

[0010] This can be another LLM, or the same LLM can be further prompted accordingly (so that it takes the prediction into account). Thus, the mission plan can be generated efficiently.

[0011] Embodiment 4 is a method according to embodiment 1 or 2, comprising generating a task plan by feeding a prediction (in textual form as an output of the LLM) and a textual target description, a textual environment description, and a textual description of the state of the robotic device to another large language model ("LLM planner").

[0012] The separation into two LLMs enables fine-tuning or training of the two LLMs for corresponding tasks (eg, through reinforcement learning, such as reinforcement learning from human feedback).

[0013] Embodiment 5 is a method according to embodiment 4, comprising training the other large language model to generate task planning.

[0014] For example, another LLM can be trained using reinforcement learning (eg, with the aid of rewards obtained for controlling the robotic device using the task plan generated for the robotic device), thereby adapting the other LLM to its use for control.

[0015] Embodiment 6 is a method according to one of embodiments 1 to 5, comprising generating a scene graph from the collected state information and from other information about the environment of the robotic device and generating the textual state description from the scene graph.

[0016] This enables efficient generation of textual state descriptions, since for example the relationships between agents and other objects (or spaces, etc.), represented by a scene graph (via edges), can be converted directly into text.

[0017] Embodiment 7 is a (eg robot) control device which is configured to carry out the method according to one of embodiments 1 to 6.

[0018] Embodiment 8 is a computer program having instructions, which, when executed by a processor, cause the processor to execute the method according to one of embodiments 1 to 6.

[0019] Embodiment 9 is a computer-readable medium storing instructions, which, when executed by a processor, cause the processor to perform a method according to one of embodiments 1 to 6. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In the accompanying drawings, similar reference numerals generally refer to the same parts throughout the different views. The accompanying drawings are not necessarily drawn to correct scale, with emphasis instead generally being placed on depicting the principles of the invention. In the following description, various aspects are described with reference to the following drawings.

[0021] Figure 1 A control scenario is shown.

[0022] Figure 2 The data flow in a corresponding control device for controlling a robotic device according to one specific embodiment is explained.

[0023] Figure 3 An example of a scene graph for a control scenario in which a mobile robot should be guided through the floors of a building is shown.

[0024] Figure 4 The functionality of the prediction and mission planning module according to one embodiment is illustrated.

[0025] Figure 5 A flow chart is shown which illustrates a method for controlling a robotic device according to one specific embodiment.

[0026] The following detailed description relates to the attached drawings, which illustrate specific details and aspects of the present disclosure in which the present invention may be implemented for the purpose of explanation. Other aspects may be used and structural, logical and electrical changes may be made without departing from the scope of protection of the present invention. Different aspects of the present disclosure are not necessarily mutually exclusive, as some aspects of the present disclosure may be combined with one or more other aspects of the present disclosure to form new aspects. DETAILED DESCRIPTION

[0027] Different examples are described in more detail below.

[0028] Figure 1 A control scenario is shown.

[0029] A mobile robot 100 is located in an environment 101 (e.g. in a factory hall or on the grounds of a construction site). The robot 100 has a starting position 102 and should reach a target position 103. Obstacles 104 that should be avoided by the robot 100 are located in the environment 101. For example, the obstacles cannot be traversed by the robot 100 (e.g. machines, walls or trees) or should be avoided because the robot would damage or injure them (e.g. people, for example in a factory hall where the robot 100 is used to transport objects).

[0030] The robot 100 has a control device 105 (the control device may also be separated from the robot 100 in space, that is, the robot 100 may be remotely controlled). Figure 1 In the example scenario of FIG. 1 , the goal is to control a robot 100 by a control device 105 to navigate through an environment 101 from a starting position 102 to a target position 103. The robot 100 is, for example, an autonomous vehicle, but may also be a robot with legs or tracks or other types of drive systems (e.g. a deep sea or Mars rover).

[0031] Furthermore, the embodiments are not limited to scenarios where the robot (as a whole) should move between positions 102, 103, but can also be used to control a robot arm, the end effector of which can be moved between positions 102, 103 (without hitting obstacle 104), etc.

[0032] Accordingly, in the following, terms such as robot, vehicle, machine, etc. are used as examples of "objects" or agents to be controlled, i.e. computer-controlled technical systems (e.g. machines). The solutions described here can be applied in the case of various types of computer-controlled machines (and possibly their environment), such as robots or vehicles. In the following, the general term "robotic device" can also be used for all types of technical systems that can be controlled with the solutions described below, which are mobile and / or have one or more movable parts.

[0033] Ideally, the control device 105 has learned a control strategy that enables it to successfully control the robot 100 (from the starting position 102 to the target position 103 without hitting the obstacle 104) in a specific scenario that the controller 105 has not yet encountered for any scenario (i.e., environment, starting position, and target position).

[0034] For this purpose, the control device 105 is suitably trained. For such a training, the scenario (especially the environment 101) can also be simulated, but is usually real in use. However, the robot 100 and the environment can also still be simulated in use (for example if the simulation of the vehicle (according to the control strategy) is used to test the (other) control strategy of the autonomous vehicle).

[0035] According to various embodiments, a control device (e.g., control device 105) for a robotic device is provided that uses a large language model (LLM) to predict the behavior (or the results of the behavior) of an agent in an environment of the robotic device. These predictions can then be created by a task planning module of the control device for task planning. For example, predicting human behavior in an indoor environment, such as where a person moves to avoid collisions between a person and a robot.

[0036] In recent years, with the tremendous progress in the research of Natural Language Processing (NLP) and Generative Artificial Intelligence (KI), various powerful LLMs that can produce human-like texts have been developed.

[0037] According to various embodiments, one or more LLMs are used to predict the future activities or behavior patterns of agents (e.g., people). Thus, for example, the robot 100 can navigate through a cluttered environment with (if necessary, many) people, i.e., human "obstacles" 104, because it has been shown that LLMs can learn "common sense" and are therefore also suitable for predicting the behavior of people. In the following, one or more people are used as an example, but the behavior of other "agents" such as animals (e.g., for navigating a robot in a cowshed) can also be predicted in a similar manner.

[0038] According to various embodiments, a prediction and planning module is provided for a control device of a mobile robot, for example in a crowded indoor environment, wherein the control device receives the following inputs (information):

[0039] Description of the type of environment

[0040] Description of the task

[0041] Sensor data about the robot’s current position

[0042] A scene graph with all objects in the environment and their relationships, including, for example, one or more agents (e.g., humans).

[0043] The prediction and task planning module uses one or more LLMs to predict human behavior and create a (high-level) task plan to complete user-defined goals (i.e., tasks) while minimizing specific costs (e.g., travel segments, human-machine interaction costs, etc.).

[0044] Figure 2 A data flow in a corresponding control device for controlling a robotic device, for example an autonomous system such as a self-driving robot, according to one specific embodiment is explained.

[0045] Perception (or sensing) 201 provides information about the environment of the robotic device (in the form of a scene graph, according to various embodiments). From this information, a prediction 202 is made about the future state of the environment. This prediction 202 is used to perform a task planning 203. If the execution of a task is planned (e.g. first move to point A, wait 5 seconds, then move to point B), a movement planning 204 is performed for this purpose (i.e. determine (one or more) trajectories for performing (one or more) corresponding movements according to the task plan). Control of the robotic device is then performed 205 according to the movement plan.

[0046] The functionality of mission planning 203 and movement planning 204 may be combined into the functionality of "planning", which may be the functionality of planning module 206. Thus, mission planning 203 (which may be considered "high-level" planning) may be considered (at least primarily) as part of the planning layer of the corresponding autonomous system. In the present description, the functionality of prediction 202 and mission planning 203 is combined into the functionality of "prediction and mission planning module" 207.

[0047] According to various embodiments, perception 201 provides information in the form of a scene graph containing all objects in the environment (including static and dynamic objects, ie robots and humans) as nodes and relations between objects as edges.

[0048] Figure 3 An example of a scene graph 300 for a control scenario in which a mobile robot should be navigated through floors of a building is shown.

[0049] In this example, the scene graph 300 has a tree form. Here, the root 301 of the scene graph is assigned to a floor. For each room 302 of the floor, the scene graph 300 contains an internal (i.e., neither a root nor a leaf) node 303 of the scene graph 300 assigned to the corresponding room. Each leaf (i.e., each leaf node) 304 of the scene graph 300 is assigned to a corresponding object 305 (or a corresponding position in the room) and is connected to the node 300, which is assigned to the room where the corresponding object 305 is located.

[0050] Each node (root, internal nodes and leaves) can have one or more properties (also called attributes). This can be a binary property, such as "Ojekt_im_Raum (object is in the room) = true", "Robot_im_Raum (robot is in the room) = true" for an internal node or " (open) = false", "ist_aufnehmbar (can be picked up) = true" and also other non-binary attributes, such as for leaves "category = bed", "weight = 10 kg" or for internal nodes "room type = storage room", the position and / or speed of the object, etc., or for the root "floor type = basement". One or more of the objects are, for example, agents, i.e. in this example people (the objects are characterized, for example, by the binary property "ist_Mensch (is a person) = true" or in general "ist_dynamisches_Objekt (is a dynamic object) = true"). The robot itself can also be represented by a leaf. Depending on what the node assigned to the object (object node) represents (a person, a robot or other object in the environment), the node can be a person node, a robot node or an environment node.

[0051] The scene graph 300 can also have other levels, for example, a room can first be divided into locations (the locations are each assigned to a node connected to the corresponding room node) and then objects are described for these locations. There may also be relationships between objects (for example, a box is located in a shelf). Then an edge can be created to connect two object nodes and describe this relationship. If an agent is located near an object (in Euclidean distance) or the agent is directly related to the object (for example, a person sits on a chair), two corresponding nodes can also be connected using an edge (but for example, there is no edge between two agent nodes). The scene graph can then no longer have the form of a tree.

[0052] According to various embodiments, LLMs are used to predict human behavior for prediction 202. These predictions are provided as input to task planning 303, which then outputs a task plan to calculate the robot's movement trajectory using another LLM for motion planning 204. The control device 205 calls the corresponding actuators (e.g., for moving arms, wheels, etc.) to follow the movement trajectory.

[0053] For example, this is used in a factory scenario for assembling electric bicycles, where the robot 100 should collect different parts of the electric bicycle, such as the steering wheel, rear wheel, frame, battery and seat, from different locations and bring the parts to the assembly area in sequence without driving through areas occupied by people (i.e., without colliding with or obstructing people).

[0054] The task planning problem is given by a tuple (O, P, A, T, C, I, G). O is the set of all elementary objects of the problem. P is a set of properties (or attributes), each of which is defined by one or more objects (binary properties are a subclass of attributes with Boolean values). A is a finite set of actions that operate on tuples of objects. T is a state transition model, C represents the cost of state transitions. I is the initial state, and G represents one or more target states. A state is an assignment of values ​​to all possible properties of an object. As described above, the state is described, for example, by a scene graph (at least in part, for example information about the robot state (such as joint positions, etc.) may also exist as additional information).

[0055] A language model (LM stands for English: language model) can be viewed as a distribution over sequences of word tokens. One approach is to model the conditional distribution p(wk|w1:k) over wk (the next token at the kth position) given a sequence of previous input tokens w1:k. Newer LMs are based on transformer architectures that significantly amplify LMs and lead to LLMs (English: large language models), i.e. large language models, which typically have tens or hundreds of billions of parameters and are able to produce human-like text or programming code.

[0056] Figure 4 The functionality of the forecasting and mission planning module 400 is illustrated according to one embodiment.

[0057] The prediction and mission planning module 400 includes the "SG2NL" module 401, the LLM predictor for predicting human behavior 403 and the LLM planner "LLMPlan" 404 that creates a task plan for the robot, the "SG2NL" module 401 converts the scene graph (SG) 402 into a natural language (English: natural language) description of the scene (i.e., state) described by the scene graph (e.g., a description of the characteristics or state of the objects assigned to the leaves in words).

[0058] The inputs of the prediction and task planning module 400 are the scene graph 402 and the user-defined natural language goal description G NL , and its output is the generated task plan π. As explained above, the scene graph 402 consists of multiple nodes and edges, such as Figure 3 As shown in: a part of the node (such as the leaf) represents the objects and agents in the environment (such as people and the robot itself), and the node (including the leaf) can have multiple attributes, such as position, orientation, action of the person or robot, etc.

[0059] Algorithm 1 gives the pseudocode for obtaining the target G from the scene graph.NL Example of creating a task plan (common English keywords such as while, do, and end are used here).

[0060]

[0061] Algorithm 1: Task Planning with Human Behavior Prediction

[0062] In line 2 of Algorithm 1, the person state S of (one or more persons) is determined from the scene graph SG H , Environmental status S E and the robot state S R , where S H and S E are two attribute sets of a person node or (for other objects and rooms) an environment node, and S R Is a collection of attributes for a robot node.

[0063] In row 3, the scene graph 402 is parsed by the SG2NL module 401. The SG2NL module 401 converts the scene graph 402 into a NL description SHNL of the human state, a NL description SENL of the environment state, and a NL description SRNL of the robot state, where SHNL, SENL, and SRNL are text sentences in natural language.

[0064] The SHNL is then passed to the LLM predictor 403 which generates predictions of future human behavior in natural language. (Line 4), which appends the predicted human behavior to the current human state. An example of an output prediction could be: “A person is standing in front of a refrigerator. The person will open the refrigerator soon.”

[0065] Combine the description of the environment state SENL with the target description GNL, the robot state SRNL in natural language and the prediction of the human state in natural language It is forwarded to the LLM planner 404, which generates a task plan π (e.g., a sequence of steps) therefrom, which guides the robot to a given task goal (line 5).

[0066] For example, a task plan for putting an apple in a refrigerator might look like: "Walk to the refrigerator, wait until the person leaves the refrigerator area, open the refrigerator, put the apple in the refrigerator, close the refrigerator."

[0067] The LLM predictor 404 and the LLM planner 403 may be the same pre-trained LLM, but different hints are given to them accordingly.

[0068] Then, the task plan π is executed by the robot, and the robot interacts with the environment 405 and observes the next human state S'H , environmental state S' E and the next robot state S' R (Line 6). Then the scene graph is updated accordingly through state changes (Line 7). This process is repeated until the user-defined goal is reached.

[0069] Algorithm 2 gives an example of converting the scene graph SG into natural language through the SG2NL module (common English keywords such as for, do, end, if, else, then are also used here).

[0070]

[0071] Algorithm 2: Convert scene graph to natural language

[0072] First, the NL description SHNL of the human state, the NL description SENL of the environment state, and the NL description SRNL of the robot state are initialized to empty (line 1). From line 2 to line 10, all nodes in the scene graph are traversed in a loop. For each node, all edges (i.e., relationships) and neighbor nodes (i.e., adjacent objects and people) are visited (lines 3-4). If the node is a human node, the current state (position, orientation, and action) of the person and its relationship with all adjacent objects and people are converted into text and added to SHNL (lines 5-7). If the node is an environment node representing other objects, the position of the object and its relationship with all adjacent objects (including people) are recorded in SENL (lines 8-10). Otherwise, the relationship with all adjacent objects and people of the robot node is added to SRNL (line 12).

[0073] In summary, according to various embodiments, there is provided Figure 5 The method shown in .

[0074] Figure 5 A flow chart 500 is shown which illustrates a method for controlling a robotic device according to one specific embodiment.

[0075] In 501, state information is collected about an agent (e.g., a person or an animal, i.e., a "dynamic object" that moves in the environment and affects the environment when necessary, e.g., a person can remove another object from the environment or carry another object to another place) located in the environment of the robotic device.

[0076] At 502, state information about the agent is converted into a textual state description.

[0077] At 503, the textual state description is fed to an LLM ("LLM predictor") to generate predictions of the agent's behavior, the agent's future state, and / or the environment's future state (e.g., the LLM may output the agent's behavior or the results of the behavior given the state of the environment (including the agent itself)). In other words: the textual state description is fed to a (trained) LLM to generate predictions.

[0078] At 504 , a task plan is generated for the robotic device taking into account the predictions.

[0079] At 505, the robotic device is controlled according to the mission plan.

[0080] Figure 5 The method can be performed by one or more computers having one or more data processing units. The term "data processing unit" can be understood as any type of entity that enables processing of data or signals. For example, data or signals can be processed according to at least one (i.e., one or more) specific functions performed by the data processing unit. The data processing unit may include an analog circuit, a digital circuit, a logic circuit, a microprocessor, a microcontroller, a central unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable gate array integrated circuit (FPGA), or any combination of these, or may be constructed by an analog circuit, a digital circuit, a logic circuit, a microprocessor, a microcontroller, a central unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable gate array integrated circuit (FPGA), or any combination of these. Any other means for implementing the corresponding functions described in more detail herein may also be understood as a data processing unit or a logic circuit device. One or more of the method steps described in detail here may be performed (e.g., implemented) by a data processing unit through one or more specific functions performed by the data processing unit.

[0081] According to various embodiments, the method is therefore in particular computer-implemented.

[0082] Controlling the robotic device according to the task plan comprises generating one or more control signals for the robotic device. The term "robotic device" may be understood to relate to any technical system having mechanical parts whose movement is controlled, such as a computer-controlled machine, a vehicle, a household appliance, a power tool, a manufacturing machine, a personal assistant or an access control system.

[0083] Various embodiments may receive sensor signals from various sensors, such as video, radar, lidar, ultrasound, motion, thermal imaging, etc., and use these sensor signals to, for example, provide sensor data for state information. The sensor data may be processed to obtain state information (i.e., for perception). This may include classifying the sensor data or performing semantic segmentation on the sensor data, for example, to detect the presence of an object (in the environment in which the sensor data was obtained).

Claims

1. A method for controlling a robotic device (100), comprising: collecting (501) state information about agents located in an environment (101) of the robotic device (100); converting (502) state information about the agent into a textual state description; feeding (503) the textual state description to a large language model (403) to generate predictions of the agent's behavior, the agent's future state, and / or the environment's future state; generating (504) a task plan for the robotic device (100) taking into account the prediction; as well as The robotic device (100) is controlled (505) according to the task plan.

2. The method according to claim 1, comprising training the large language model (403) for predicting agent behavior from textual state descriptions.

3. The method according to claim 1 or 2 includes generating the task plan by feeding a textual target description, a textual environment description and a textual description of the state of the robotic device (100) to the large language model or another large language model (404).

4. The method according to claim 1 or 2 includes generating the task plan by feeding the prediction and textual target description, textual environment description and textual description of the state of the robotic device (100) to another large language model (404).

5. The method of claim 4, comprising training the other large language model (404) to generate a task plan.

6. The method according to any one of claims 1 to 5, comprising: A scene graph (402) is generated from the collected state information and from other information about the environment (101) of the robotic device (100), and the text state description is generated from the scene graph. 7 . A control device configured to carry out the method according to claim 1 .

8. A computer program comprising instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable medium storing instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 6.