Apparatus and method for controlling a robotic device

A control model for robotic devices uses demonstrations to quickly determine skill sequences based on state transitions and geometric conditions, addressing the inefficiencies of existing methods and enabling efficient robotic task execution.

JP7728139B2Active Publication Date: 2025-08-22ROBERT BOSCH GMBH
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2021164686
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-07
Filing Date
2021-10-06
Publication Date
2025-08-22
Estimated Expiration
2041-10-06

AI Technical Summary

Technical Problem

Existing robotic devices require long analysis times to determine skills for execution, especially when multiple parameters are involved, which can hinder efficient task performance.

Method used

A control model is developed that considers the relationship between state transitions and geometric conditions to quickly determine the sequence of skills required for a robotic device to perform a task, using a trained model based on demonstrations.

Benefits of technology

This approach allows for rapid skill determination with reduced computational costs and no manual selection process, enabling efficient and scalable control of robotic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007728139000298
    Figure 0007728139000298
  • Figure 0007728139000299
    Figure 0007728139000299
  • Figure 0007728139000300
    Figure 0007728139000300
Patent Text Reader

Abstract

To provide a device and a method for controlling a robotic device.SOLUTION: A method for controlling a robotic device, includes: a step of assigning a set of task parameters to each state transition; a step of obtaining a set consisting of a state transition-state-state transition triple included in a supplied control state sequence; a step of matching a parameter model so that, for each state transition-state-state transition triple, the parameter model finds a probability distribution relating to each task parameter belonging to the task parameter set assigned to the state transition following the state; a step of matching an object model so that the object model finds a probability distribution relating to the state of an object for each object in one or more objects; and a step of controlling the robotic device with a control model using a trained parameter model and a trained object model.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Various embodiments relate generally to apparatus and methods for controlling robotic devices. [Background technology]

[0002] A robotic device may have multiple skills and may execute multiple skills to perform a task. In this case, the behavior of the robotic device when executing a skill may depend on multiple parameters. A skill to be executed to perform a task may be determined, for example, using a robot control model. However, a large number of skills associated with a large number of parameters may require a long analysis time to determine a skill to be executed. When the robotic device is in operation, for example, a short analysis time may be required to determine a skill to be executed.

[0003] A further potential advantage is to transfer robot capabilities based on the demonstration to a robot control model. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Publication: “Optimizing sequences of probabilistic manipulation skills learned from demonstration,” by L. Schwenkel, M. Guo, and M. Buerger, Conference on Robot Learning, 2019 (hereinafter referred to as Reference [1]) Summary of the Invention [Problem to be solved by the invention]

[0005] The publication "Optimizing sequences of probabilistic manipulation skills learned from demonstration" by L. Schwenkel, M. Guo and M. Buerger, Conference on Robot Learning, 2019 (hereinafter referred to as Reference [1]) describes a skill-centric approach, whereby skills are learned independently under different scenarios and are not tied to one specific task. [Means for solving the problem]

[0006] According to a method and an apparatus with the features of independent claim 1 (first embodiment) and independent claim 9 (ninth embodiment), skills for performing a task can be determined by a robotic device with a small computational structure and / or a small time expenditure. For example, the method and apparatus can determine skills to be performed by a robotic device with respect to a current state and a target state of the robotic device with a small computational structure and / or a small time expenditure. In particular, the method and apparatus can provide skills to be performed using a trained control model. Furthermore, a method and an apparatus are provided that can transfer robot capabilities to the control model based on demonstrations.

[0007] The robotic device can be any type of computer-controlled device, such as a robot (e.g., a manufacturing robot, a maintenance robot, a household robot, a medical robot, etc.), a vehicle (e.g., an autonomous vehicle), a home appliance, a manufacturing machine, a personal assistant, an access control system, etc.

[0008] The control model formed according to the first embodiment takes into account, for example, on the one hand, the relationship between the last executed state transition (e.g., the last executed skill) and the subsequently executable state transition (e.g., the subsequently executable skill), and on the other hand, the geometric conditions underlying the sequence of these two state transitions.

[0009] This control model is advantageous in that, for example, complex calculations are not required to determine a sequence of state transitions to be performed (e.g., a sequence of skills to be performed that can be induced by the state transitions to be performed). For example, skills to be performed can be determined in a short amount of time, regardless of the number of skills, parameters, and / or objects involved in the surroundings of the robotic device. Furthermore, this control model has the advantage that no manual selection process is required to select skills to be executed.

[0010] This control model is advantageous in that it can be linearly scaled, for example, for additional skills of the robotic device.

[0011] Each control state sequence may include an alternating series of states and state transitions.The features described in this paragraph may be combined with the first embodiment to form a second embodiment.

[0012] The step of controlling the robotic device using the control model may include, for each state, the following steps: determining an individual task parameter set for each executable state transition in the state in response to inputting a goal state and the state transition that achieved the state into the trained parameter model; determining an individual probability distribution for each object in the one or more objects in response to inputting the goal state, individual states of other objects in the one or more objects, and the task parameter set determined using the trained parameter model into the trained object model; determining a probability for each executable state transition in the state using the probability distributions determined for the one or more objects; and determining the executable state transition with the highest determined probability as the state transition to be executed. The features described in this paragraph may be combined with the first or second embodiment to form a third embodiment.

[0013] Determining a respective set of task parameters for each executable state transition in the state may include the steps of: determining a respective probability distribution for each task parameter in the task parameter set; and determining an expectation of the respective probability distribution as a task parameter in the task parameter set. The features described in this paragraph may be combined with the third embodiment to form a fourth embodiment.

[0014] The step of providing a control state sequence for each initial state-goal state pair in the multiple initial state-goal state pairs may include the steps of: selecting state transitions in each state of the robotic apparatus from the initial state to the goal state; and determining, by simulation, task parameter sets assigned to the selected state transitions and states of the robotic apparatus resulting from the state transitions. The features described in this paragraph may be combined with one or more of the first to fourth embodiments to form a fifth embodiment.

[0015] Providing a control state sequence for the initial state-goal state pair may include determining a number of possible control state sequences for the initial state-goal state pair, each control state sequence having an alternating sequence of states and state transitions, and determining as the control state sequence for the initial state-goal state pair the possible control state sequence having the shortest sequence of states and state transitions. The features described in this paragraph may be combined with one or more of the first through fifth embodiments to form a sixth embodiment.

[0016] The control model may further include a robot trajectory model, a precondition model, and an exit condition model for each state transition. The step of training the control model may further include the following steps: providing a demonstration to perform each state transition in the plurality of state transitions; training a robot trajectory model for each state transition using the demonstration, where each robot trajectory model is a hidden semi-Markov model having one or more initial states and one or more end states; and training a precondition model and an exit condition model for each state transition using the demonstration, where the precondition model has, for each initial state of the robot trajectory model assigned to the state transition, a probability distribution of a robot configuration before execution of the state transition, and the exit condition model has, for each end state of the robot trajectory model assigned to the state transition, a probability distribution of a robot configuration after execution of the state transition. The step of controlling the robot device using the control model may include the following steps for each state. That is, it may include the steps of determining a task parameter set using the trained parameter model, determining a state transition to be executed using the trained object model, determining a robot trajectory according to a robot trajectory model using the state transition to be executed and the task parameter set, and controlling the robot device to execute the determined robot trajectory. The features described in this paragraph may be combined with one or more of the first to sixth embodiments to form a seventh embodiment.

[0017] The effect of training a control model based on a demonstration is that the learning efficiency is increased (e.g., the required computational costs are reduced, e.g., the required time consumption is reduced, e.g., the cost of generating learning or training data is reduced).

[0018] Furthermore, the effect obtained by training a robot trajectory model based on a demonstration is that the performance (e.g., efficiency) of determining the next skill to be executed using the trained robot trajectory model is significantly improved, which has the advantage that complex simulation is not required.

[0019] Providing a control state sequence for an initial state-goal state pair may include the steps of: selecting a state transition at each state of the robotic device from the initial state to the goal state; determining a task parameter set assigned to the selected state transition using a precondition model assigned to the selected state transition; and determining a state of the robotic device resulting from the state transition using an exit condition model assigned to the selected state transition. The features described in this paragraph may be combined with the seventh embodiment to form an eighth embodiment.

[0020] Specifically, by using a control state sequence determined based on a prior problem statement regarding the current state and the target state, a control model can be trained, taking into account geometric conditions, and thus improving the performance of the robot device during operation.

[0021] A computer program product can store program instructions that, when executed, perform a method according to one or more of the first through eighth embodiments. A computer program product having the features set forth in this paragraph constitutes a tenth embodiment.

[0022] The non-volatile storage medium can store program instructions that, when executed, perform a method according to one or more of the first through eighth embodiments. A non-volatile storage medium having the features described in this paragraph constitutes an eleventh embodiment.

[0023] The non-transitory storage medium can store program instructions that, when executed, perform a method according to one or more of the first through eighth embodiments. A non-transitory storage medium having the features described in this paragraph constitutes a twelfth embodiment.

[0024] The drawings show exemplary embodiments of the invention and are explained in more detail below. [Brief explanation of the drawings]

[0025] [Figure 1] FIG. 1 illustrates an exemplary robotic device, according to various embodiments. [Figure 2] FIG. 1 illustrates a flowchart for generating a control model according to various embodiments. [Figure 3] FIG. 10 illustrates steps for determining an exemplary control state sequence according to various embodiments. [Figure 4] FIG. 1 illustrates a flowchart of demonstration-based learning in accordance with various embodiments. [Figure 5] FIG. 1 illustrates an apparatus for recording a user's demonstration according to various embodiments. [Figure 6A] FIG. 1 illustrates a flowchart for controlling a robotic device, according to various embodiments. [Figure 6B] FIG. 1 illustrates a flowchart for controlling a robotic device, according to various embodiments. [Figure 7] FIG. 1 illustrates operational aspects of a control model for an exemplary task in accordance with various embodiments. [Figure 8] FIG. 1 illustrates a method for controlling a robotic device, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0026] In one embodiment, a "computer" may be understood to refer to any type of entity that implements logic, and may be hardware, software, firmware, or a combination thereof. Thus, in one embodiment, a "computer" may be a hardwired logic circuit or a programmable logic circuit, such as a programmable processor, for example, a microprocessor (e.g., CISC (large instruction set) or RISC (reduced instruction set)). A "computer" may include one or more processors. A "computer" may also be software implemented or executed by a processor, such as any type of computer program, for example, a computer program using virtual machine code such as Java. Any other type of implementation of the individual functions described in more detail below may also be understood to be a "computer" according to one alternative embodiment.

[0027] When controlling a robotic device to perform a task, a robot control model can be used to determine skills that are or should be performed by the robotic device to accomplish the task. When the robotic device is in operation, it may be necessary to determine the skills to be performed with little latency (e.g., short analysis time). Various embodiments relate to apparatus and methods for controlling a robotic device that can determine the skills to be performed by the robotic device in consideration of a task with little latency. Furthermore, various embodiments relate to methods and apparatus that can use demonstrations to generate a model for determining the skills to be performed.

[0028] 1 illustrates a robotic device assembly 100. The robotic device assembly 100 may include a robotic device 101. For illustrative purposes, the robotic device 101 illustrated in FIG. 1 and illustratively described below represents one exemplary robotic device and may include, for example, an industrial robot in the form of a robotic arm for moving, installing, or handling an unfinished product. It is noted that the robotic device may be any type of computer-controlled device, such as a robot (e.g., a manufacturing robot, a maintenance robot, a household robot, a medical robot, etc.), a vehicle (e.g., an autonomous vehicle), a home appliance, a manufacturing machine, a personal assistant, an access control system, etc.

[0029] The robotic device 101 has robotic components 102, 103, 104 and a base (or generally a holding device) 105 that supports the robotic components 102, 103, 104. The term "robotic components" refers to the movable members of the robotic device 101 that can be manipulated to physically interact with their surroundings, e.g., to perform a task, e.g., to execute or perform one or more skills.

[0030] For control purposes, the robotic assembly 100 has a controller 106, which is configured to implement interaction with its surroundings according to a control program. The last of the robot components 102, 103, 104 (as viewed from the base 105) 104 is also referred to as the end effector 104 and may include one or more tools such as a welding burner, a gripping tool, a lacquer device, or the like.

[0031] Other robot components 102, 103 (closer to base 105) can form a positioning device and are therefore paired with an end effector 104 to provide a robot arm (or articulated arm) with an end effector 104 at its end. This robot arm is a mechanical arm that can perform functions similar to a human arm (possibly with a tool at its end).

[0032] The robotic device 101 may include link members 107, 108, and 109 that connect the robotic components 102, 103, and 104 to each other and to the base 105. The link members 107, 108, and 109 may include one or more joints, each of which may provide rotational and / or translational (i.e., shifting) movement of the associated robotic components relative to each other. Movement of the robotic components 102, 103, and 104 may be effected using actuators controlled by the controller 106.

[0033] The term "actuator" can be understood to mean a component suitable for exerting an effect on a mechanism in response to being actuated. An actuator can translate commands (so-called activates) sent from the controller 106 into mechanical movements. An actuator can be, for example, an electromechanical converter configured to convert electrical energy into mechanical energy in response to being controlled.

[0034] The term "controller" (also referred to as "control mechanism") may be understood to refer to any type of logic-implemented unit, which may include, for example, circuitry and / or a processor capable of executing software, firmware, or a combination thereof stored on a storage medium, and in this example, providing instructions to, for example, actuators. The controller may be configured, for example, by program code (e.g., software) to control the operation of the system, in this example, the operation of the robot.

[0035] In this example, the controller 106 includes a computer 110 and a storage device 111, where the storage device 111 stores code and data based on which the computer 110 controls the robotic device 101. According to various embodiments, the controller 106 controls the robotic device 101 based on a robot control model 112 stored in the storage device 111.

[0036] According to various embodiments, the robotic device assembly 100 may include one or more sensors 113. The one or more sensors 113 may be configured to provide sensor data characterizing the state of the robotic device. For example, the one or more sensors 113 may include, for example, an imaging sensor such as a camera (e.g., a standard camera, a digital camera, an infrared camera, a stereo camera, etc.), a radar sensor, a LIDAR sensor, a position sensor, a velocity sensor, an ultrasonic sensor, an acceleration sensor, a pressure sensor, etc.

[0037] The robotic device 101 can be in one of many states. According to various embodiments, the robotic device 101 can have one current state of many states at any given time. Each state of the many states can be determined using sensor data provided by one or more sensors 113 and / or the configuration of the robotic device 101.

[0038] A state transition may occur between each state and the state that follows this state. As used herein, the term "state transition" may correspond to an action and / or skill of the robotic device 101. Specifically, the robotic device 101 may perform one action and / or skill in one state, thereby leading to a new state of the robotic device 101.

[0039] The robotic device 101 can be configured to perform multiple skills. A plurality of skills within the multiple skills can be predefined, for example, in the program code of the controller 106. One or more skills within the multiple skills can include, for example, mechanical movement of one or more robotic components 102, 103, 104. One or more skills within the multiple skills can include, for example, an end effector action (e.g., grasping, releasing, etc.). According to various embodiments, a skill performed in a current state of the robotic device 101 can lead to one of multiple resulting states of the robotic device 101.

[0040] The robot control model 112 can be configured to determine a state transition to be performed, and the controller 106 can be configured to control the robotic device 101 to perform the state transition. The robot control model 112 can be configured to determine a skill to be performed, and the controller 106 can be configured to control the robotic device 101 to perform the skill.

[0041] According to various embodiments, at least a portion of the control model 112 can be configured to provide a state transition to be performed (e.g., a skill to be performed) for a certain state of the robotic device 101 and a goal state of the robotic device 101. The control model 112 can be configured to provide a skill to be performed and task parameters assigned to the skill for a certain state of the robotic device 101 and a goal state of the robotic device 101. The goal state can be, for example, a state in which a task to be performed is or has been completed.

[0042] According to various embodiments, the control model 112 can be generated (e.g., trained) while the robotic device 101 is not in motion. According to various embodiments, the generated control model 112 can be used while the robotic device 101 is in motion to determine skills to be performed by the robotic device 101.

[0043] 2 illustrates a flowchart 200 for generating a control model 206 according to various embodiments. The control model 206 thus generated can be used, for example, as control model 112 and / or as part of control model 112. A computer can be configured to generate the control model 206. This computer can be, for example, computer 110 of controller 106. As described herein, the control model 206 can be generated (trained) even when the robotic device 101 is not being operated, and thus, for example, this computer can be a different computer than computer 110. For example, the control model 206 can be trained spatially separate from the robotic device assembly 100.

[0044] According to various embodiments, a number of initial state-goal state pairs 202 {(S0, S N )} for the same task (e.g., multiple initial state-goal state pairs 202 {(S0, S NAccording to various embodiments, a computer is configured to generate a number of initial state-goal state pairs 202 {(S0, S N )}. For example, according to an embodiment, multiple initial state-goal state pairs 202 {(S0, S N )} for each initial state-goal state pair 202 (S0, S N ) is the initial state S0 and the target state S N Each initial state S0 may be one of many states of the robotic device 101. Each initial state S0 may represent one state of the robotic device 101 and one or more objects. Each goal state S N can be one of many states of the robotic device 101. Each goal state S N can represent one state of the robotic device 101 and one or more objects. For example, a computer can be configured to select each initial state-goal state pair (S0, S1) from a state space having many states. N ) (e.g., substantially randomly, e.g., using a predefined algorithm).

[0045] According to various embodiments, a computer is configured to generate a number of initial state-goal state pairs 202 {(S0, S N )}, each initial state-goal state pair (S0, S N ), the control state sequence ξ can be configured to be determined for each of the initial state-goal state pairs 202 {(S0, S N )}, multiple control state sequences 204

number

number

[0046] For concreteness, the state transition will be explained below based on the skills of the robotic device 101.

[0047] Specifically, in the initial state S0, the skill

number

number

number

number

[0048] According to various embodiments, the assigned initial state-goal state pair (S0, S N ) control state sequence ξ as

number

[0049] Specifically, the state transition is

number

number

[0050] According to various embodiments, a task parameter set

number

number

[0051] The initial state-goal state pair (S0, S N ) can be determined using, for example, a Task and Motion Planning (TAMP) solution. The TAMP solution can combine, for example, discrete logical reasoning with geometric constraints on the robotic device 101. The TAMP solution can be implemented, for example, as a neural network. According to various embodiments, the control state sequence for an initial state-goal state pair (S0, S N The control state sequence for (i.e., ∑ i = 1 ∑ i = 1 ∑ i = 2 ... *It may include an algorithm, where each state transition (e.g., each skill) can be executed (e.g., theoretically executed) in each current state to move the system to a resulting state.

[0052] skill

number

number

[0053] According to various embodiments, the skill

number

number

number

number

number

number

number

number

number

[0054] According to various embodiments, a number of initial state-goal state pairs 202 {(S0, S N )}, each initial state-goal state pair (S0, S N ), many possible control state sequences can be found. For a given initial state-goal state pair (S0, S N ) is a sequence of control states from an initial state S0 to a target state S N The control state sequence may include an alternating sequence of states and state transitions (e.g., skills) from the initial state to the target state pair (S0, S1). According to various embodiments, the control state sequence with the smallest determined cost among the many possible control state sequences is called the initial state-goal state pair (S0, S1). N) can be determined as the control state sequence ξ for the initial state-goal state pair (S0, S1). According to various embodiments, among the many possible control state sequences, the possible control state sequence that has the shortest sequence of states and state transitions (e.g., including at least a plurality of states or a plurality of state transitions) is determined as the control state sequence ξ for the initial state-goal state pair (S0, S1). N ) can be obtained as the control state sequence ξ for

[0055] 3 illustrates steps for determining an exemplary control state sequence 204A. For example, if an initial state-goal state pair (S0, S1) is 12 An example control state sequence 204A for the control state sequence 202A can be obtained, where the initial state is S0=S0 and the target state is S N =S N12 The exemplary control state sequence 204A can be determined, for example, by a graph search algorithm, using the specifically illustrated state-state transition diagram 302. As described herein, each state transition can be assigned an individual task parameter (represented as p in the state-state transition diagram 302). The state-state transition diagram 302 shows an exemplary determined sequence 304 of states and state transitions from an initial state to a goal state. Based on this example, the exemplary control state sequence 204A can be described as follows:

number

[0056] Referring to FIG. 2, the determined sequence of control states 204

number

number

number

number

number

[0057] The following describes an exemplary generation of the control model 206 according to various embodiments.

[0058] According to various embodiments, for each control state sequence ξ, a virtual initial skill is assigned at the beginning of the control state sequence (i.e., a sequence of states and skills).

number

number

number

[0059] According to various embodiments, a computer can be configured to determine a set of state transition-state-state transition triples contained within a provided sequence of control states. According to various embodiments, the computer can determine a set of state transition-state-state transition triples contained within a provided sequence of control states.

number

number

number

[0060] Each first triple

number

number

number

number

number

number

number

number

[0061] According to various embodiments, each skill transition

number

number

number

number

number

number

number

number

[0062] According to various embodiments, the control model 206 is a parameter model

number

number

number

number

number

number

number

number

number

number

number

number

[0063] According to various embodiments, each task parameter

number

number

number

number

number

number

number

number

number

number

number

[0064] According to various embodiments, a parametric model

number

number

number

[0065] According to various embodiments, each object

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0066] According to various embodiments, an object model

number

number

number

number

[0067] The TP-GMM can be trained, for example, using the Expectation Maximization (EM) algorithm, which is described in more detail with reference to Figures 4 and 5 and reference [1].

[0068] According to various embodiments, a parametric model

number

number

number

number

number

number

[0069] Specifically, the trained parameter model

number

number

number

number

number

number

number

number

number

number

[0070] According to various embodiments, the control model 206 may be configured to handle a number of skill transitions (e.g., a set of skill transitions).

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0071] Control Model 206

number

number

number

number

number

number

number

number

[0072] The control model 206 thus generated takes into account, on the one hand, the possible transitions of skills and, on the other hand, the geometric conditions underlying these transitions, whereby the control model 206 is parameterized to the target state. Specifically, the control model 206 thus generated is a compact representation for the TAMP strategy.

[0073] demonstration

number

number

number

[0074] According to various embodiments, the control model 206 can be trained for a predefined task. According to various embodiments, multiple control models can be trained, where each control model in the multiple control models is assigned to a separate task.

[0075] FIG. 4 illustrates a flowchart 400 of demonstration-based learning according to various embodiments.

[0076] For example, to teach a robot a skill, such as moving along a desired trajectory, kinesthetic demonstrations can be performed in which the robot is moved directly, for example, by physically pushing it or using teleoperation. Besides the experience required, safety risks, and demands (e.g., for tasks requiring precise movement), robotic movement is also significantly less intuitive for humans to perform tasks compared to using their own hands.

[0077] In light of the above, various embodiments provide an approach that allows a human user to teach a robot a job (skill) simply by having the human user perform the job themselves. In this case, a demonstration is recorded, for example, by tracking the user's hand (and optionally any objects involved) instead of recording the trajectory of the end effector. The demonstration is then used to learn a compact mathematical representation of the skill, which can then be used (e.g., by the controller 106) to reproduce the skill by the robot in new scenarios (e.g., new relative positions between the robot and the object to be manipulated).

[0078] Various embodiments are based on technological advances in two areas: first, camera-based hand tracking is commonly available in areas where robots are used, e.g., in factories; and second, robot training methods based on human demonstrations enable efficient learning (i.e., efficient training of robots) and flexible replication. One example in this regard is the Task-Parameterized Hidden Semi-Markov Model (TP-HSMM), which allows learned motor skills to be expressed in terms of task parameters.

[0079] Tracking objects and human hands is an active area of ​​research (especially in machine vision) and is extremely important for industrial applications. Apart from applying the corresponding techniques to human-machine interaction (e.g., for video games), according to various embodiments, it is used for robot training (lessons) and learning.

[0080] In the demonstration phase, a user (or generally a demonstrator) demonstrates a desired skill. This demonstration is recorded; for example, a video recording is made using a camera, and a sequence of positions of the user's hand (generally a part of the demonstrator) is determined from the video images and represented in the form of a trajectory. This is repeated over multiple demonstrations 402. Note that this can be done in a decoupled manner, i.e., using a large amount of video that has been pre-recorded without the intention of teaching the skill to a robot.

[0081] In the learning or training phase, a mathematical model is trained based on the collected demonstrations. For example, a TP-HSMM is trained, which includes hand pose as one of the task parameters. A "pose" may include, for example, information about position and / or orientation, or may also include information about the state (e.g., "hand closed" vs. "hand open").

[0082] FIG. 5 illustrates an apparatus 500 for recording a user's demonstration according to various embodiments.

[0083] A user 501 demonstrates a skill by movements of his / her hand 502. For example, the user picks up an object 504 from a first position 505 and moves it to a second position 506. A camera 507 records the user's movements. Multiple cameras can be provided, each recording the demonstration from a different viewing angle, in particular from the perspective of the start position 505 and the perspective of the end position 506 of the object 504.

[0084] Each demonstration is thus represented as a series of images, which are provided to a controller 508, e.g., corresponding to controller 106. Controller 508 may, for example, include a computer for performing calculations. Controller 508 trains a statistical model 509 based on the demonstration, e.g., robot trajectory model 404 and / or TP-GMM 406 (e.g., precondition model and termination condition model, as described herein). It is further assumed that different coordinate systems, referred to as task parameters, are used to calculate the robot trajectory model 404 and / or the TP-GMM 406 (e.g., precondition model and termination condition model, as described herein).

number

[0085] For example, at the end of the demonstration phase, the demonstration can be abstracted (e.g., represented as a progression of coordinates of the hand 502 or object 504) and further stored as a trajectory (e.g., of the hand 502 or object 504, or even of multiple hands and / or multiple objects), for example, in a memory device of the control device 508.

[0086] Referring to Fig. 4, TP-HSMM enables efficient learning and flexible replication to learn robot capabilities based on human demonstration. More precisely, the recorded trajectory of a user's hand 502 is treated as the desired motion to be learned, while the trajectory of an object 504 is used to generate various task parameters for the skill, which represent various configurations of the workspace. These task parameters can be determined, for example, based on the current state. The task parameters can be selected, for example, arbitrarily.

[0087] According to various embodiments, a robot trajectory model 404 can be determined using the demonstration 402. The robot trajectory model 404 can be a TP-HSMM.

[0088] The HSMM (Hidden Semi-Markov Model) extends the simple HMM (Hidden Markov Model) so that temporal information is embedded in the underlying stochastic process. In the case of the HMM, it is assumed that the underlying stochastic process has the Markov property, i.e., the probability of transitioning to the next state depends only on the current state, whereas in the case of the HSMM, the process, i.e., the probability of transitioning to the next state, depends on the current state and the dwell time in the current state. The HSMM is typically applied in speech synthesis, among other things.

[0089] For example, a task-parameterized HSMM (TP-HSMM) such as the robot trajectory model 404, according to one embodiment,

number

number

number

number

number

number

number

number

number

number

number

number

[0090] The TP-GMM is an output probability (or emission probability, i.e., probability for an observation) for each state k=1,...K. Such a mixture model (unlike a simple GMM) cannot be trained independently for each coordinate system. The reason is that the mixture coefficients

number

number

[0091] Once the TP-GMM has been trained, the model can be used to reproduce the trajectory for the learned ability or skill during execution by the controller 508 and / or controller 106.

[0092] However, the a priori probability

number

[0093] In this considered TP-HSMM, each state corresponds to one Gaussian component in the imputed TP-GMM.

[0094] The robotic device 101 can operate within a static, known work field. Within the reach of the robotic device 101 (in many embodiments, referred to as a robot):

number

number

[0095] A further assumption is that there exists a set of core manipulation skills that allow a robot to manipulate (e.g., move) objects. We define these core manipulation skills as

number

[0096] For each job (corresponding to a skill), the user 501 performs multiple demonstrations, which define how the robotic device 101 should perform each job.

number

number

number

number

number

number

number

number

[0097] TP-GMM

number

number

number

number

number

number

number

number

number

number

number

[0098] The TP-HSMM (e.g., by the controller 508) uses the demonstration of the user 501 in the learning phase.

number

[0099] The training results are the values ​​for the parameter set that characterizes the TP-HSMM.

number

[0100] According to various embodiments, the controller 106 can control the robotic device 101 using the TP-HSMM robot trajectory model 404 to perform a job, for example, for a new scenario. For example, the controller 106 can use the robot trajectory model 404 to determine a reference trajectory for the new scenario and control the robotic device 101 so that the robotic device 101 follows the reference constraints. The term "scenario" here refers to a particular selection of modeled task parameters (e.g., a start position 505 or current position and a target position 506, e.g., a current state and a target state).

[0101] According to various embodiments, one or more TP-GMMs 406 may be determined (e.g., by the controller 508). For example, during the training phase, the precondition model

number

number

[0102] Prerequisite Model

number

number

number

number

number

number

number

number

number

[0103] Termination Condition Model

number

number

number

number

number

number

number

number

number

[0104] According to various embodiments, skill-specific manifolds

number

number

number

number

[0105] Specifically, the TP-HSMM robot trajectory model 404 is based on a single skill.

number

number

number

number

number

number

[0106] For example, training the robot trajectory model 404 as a TP-HSMM and further training the precondition model

number

number

[0107] 6A illustrates a flowchart 600A for controlling a robotic device according to various embodiments. The flowchart 600A may be a flowchart for controlling the robotic device 101 during operation.

[0108] According to various embodiments, the robotic device 101 is in an initial state S0 and the target state S F can be supplied and can be input using, for example, a user interface. For example, the control device 106 can determine the current state of the robot device 101 as the initial state S0. For example, the target state S F to the controller 106. For example, the controller 106 may determine a goal state S F Thus, in the initial state of the robot device 101, the initial state-goal state pair (S0, S F ) 602 can be provided. At any point in time, the robot device 101 can be moved from an initial state S0 to a goal state S F Until the current state S k 604.

[0109] According to various embodiments, the generated control model 206 is k 604 and Target S F In response to input from

number

number

number

number

number

number

number

[0110] According to various embodiments, the control model 206 includes a skill transition set

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0111] According to various embodiments, the control model 206 (e.g., using the trained object model) can determine, for each possible state transition in a state, an individual probability distribution for each object in the one or more objects in response to inputs of the target state, the individual states of the other objects in the one or more objects, and the determined set of task parameters. According to various embodiments, the control model 206 can determine, for each possible skill transition, an individual probability distribution for each object in the one or more objects in the state.

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0112] According to various embodiments, possible skill transitions

number

number

number

[0113] Specifically, the skills to be performed

number

number

number

number

number

[0114] According to various embodiments, the robotic device 101 can be controlled according to the control model 206. According to various embodiments, the robotic device 101 can be configured to control the skill to be performed.

number

number

number

number

number

number

[0115] A flowchart 600A for controlling the robotic device 101 can be described by Algorithm 2.

number

[0116] 6B illustrates a flowchart 600B for controlling the robotic device 101 according to various embodiments. Flowchart 600B may substantially correspond to flowchart 600A, where the controller 106 (e.g., the computer 110 of the controller 106) determines the skill to be performed.

number

number

[0117] In this case, the robot trajectory model 404 learned through the demonstration can be used to control the robotic device 101 to perform the skill in the demonstrated scenario or also to perform the skill in a scenario that was not demonstrated.

[0118] According to various embodiments, the controller 106 may select a skill to be performed.

number

number

number

number

number

number

number

number

number

[0119] 7 illustrates the operational aspects of the control model 206 for an example task 700 according to various embodiments. The robotic device 101 can be in an initial state (“STRT”) and can provide a target state (“STP”) (see, e.g., 602 in FIGS. 6A and 6B). The control model 206 can be configured with a skill transition set

number

number

number

number

number

number

number

number

number

number

number

number

[0120] FIG. 8 illustrates a method 800 for controlling a robotic device according to various embodiments.

[0121] The method 800 may include training a control model (blocks 802 through 806). The control model may include a parameter model and an object model.

[0122] Training the control model may include providing, for each initial state-goal state pair in the number of initial state-goal state pairs, a control state sequence having states and state transitions belonging to a set of possible states and state transitions (block 802). The initial states may represent states of the robotic device and one or more objects. The goal states may represent states of the robotic device and one or more objects. A task parameter set may be assigned to each state transition.

[0123] Training the control model may include determining a set of state transition-state-state transition triples contained within the supplied sequence of control states (block 804).

[0124] Training the control model may include matching a parameter model and matching an object model for each state transition-state-state transition triple in the set of state transition-state-state transition triples (block 806). The parameter model may be matched such that the parameter model determines a probability distribution for each task parameter belonging to a task parameter set assigned to a state transition following a state in response to input of the state transition-state-state transition triple and an assigned goal state of the control state sequence in which the state transition-state-state transition triple is included. The object model may be matched such that the object model determines, for each object in the one or more objects, a probability distribution for the state of the object in response to input of two state transitions in the state transition-state-state transition triple, individual states of other objects in the one or more objects, the task parameter set, and the assigned goal state.

[0125] The method 800 may include controlling the robotic device using the control model starting from a predetermined initial state through a series of states to a desired goal state (block 808). At each state, a task parameter set may be determined using the trained parameter model. At each state, a state transition to be performed may be determined using the trained object model. At each state, the robotic device may be controlled to execute the desired state transition using the determined task parameter set.

Claims

1. 1. A method for controlling a robotic device, comprising: training a control model having a parameter model and an object model, the training step comprising: for each initial state-goal state pair in a plurality of initial state-goal state pairs, providing a control state sequence having states and state transitions belonging to a set of possible states and state transitions, the initial states representing states of the robotic device and one or more objects, and the goal states representing states of the robotic device and one or more objects, and further assigning a task parameter set to each state transition; determining a set of state transition-state-state transition triples contained in the provided sequence of control states; For each state transition-state-state transition triple in the set of state transition-state-state transition triples, - adapting the parameter model for the triples such that, in response to input of the state transition-state-state transition triple and the assigned target state of the control state sequence in which the state transition-state-state transition triple is included, the parameter model determines a probability distribution for each task parameter belonging to the task parameter set assigned to the state transition following the state; - adapting the object model for the triples, such that for each object in the one or more objects, the object model determines a probability distribution over the states of the object in response to inputs of two state transitions in the state transition-state-state transition triple, individual states of other objects in the one or more objects, a task parameter set, and the assigned goal state; a step of controlling the robotic device using the control model from a predetermined initial state through a series of states until a target state is reached, in which, in each state, a task parameter set is determined using the trained parameter model, a state transition to be executed is determined using a trained object model, and the robotic device is controlled so as to execute the state transition to be executed using the determined task parameter set; Including, A method wherein each control state sequence comprises an alternating series of states and state transitions.

2. The step of controlling the robot device using the control model includes, in each state, determining, for each state transition executable in the state, a respective task parameter set in response to inputting the goal state and the state transition that achieved the state into the trained parameter model; for each state transition possible in the state, determining an individual probability distribution for each object in the one or more objects in response to inputting into the trained object model the goal state, individual states of other objects in the one or more objects, and the task parameter set determined using the trained parameter model; determining a probability for each possible state transition in said state using said probability distribution determined for one or more of said objects; determining the feasible state transition having the highest determined probability as the state transition to be executed; Including, The method of claim 1.

3. determining an individual task parameter set for each executable state transition in the state, determining an individual probability distribution for each task parameter of the task parameter set; determining the expectation of each probability distribution for the task parameters of said task parameter set; Including, The method of claim 2.

4. The step of providing a control state sequence for each initial state-goal state pair in the plurality of initial state-goal state pairs comprises: selecting a state transition for each state of the robot device from the initial state to the goal state; determining, using simulation, the task parameter sets assigned to the selected state transitions and the states of the robotic device resulting from the state transitions; Including, 4. The method according to any one of claims 1 to 3.

5. The step of providing a control state sequence for an initial state-goal state pair comprises: determining a plurality of possible control state sequences for the initial state-goal state pair, each control state sequence having an alternating sequence of states and state transitions; ● determining a possible control state sequence for the initial state-goal state pair that is the shortest possible sequence of states and state transitions; Including, 5. The method according to any one of claims 1 to 4.

6. the control model further includes, for each state transition, a robot trajectory model, a precondition model, and a termination condition model; The step of training the control model further comprises: providing a demonstration for performing each state transition in said plurality of state transitions; training the robot trajectory model for each state transition using the demonstrations, each robot trajectory model being a hidden semi-Markov model having one or more initial states and one or more end states; training the precondition model and the termination condition model for each state transition using the demonstration, wherein the precondition model has, for each initial state of the robot trajectory model assigned to the state transition, a probability distribution of the robot configuration before the state transition is performed, and the termination condition model has, for each final state of the robot trajectory model assigned to the state transition, a probability distribution of the robot configuration after the state transition is performed; Including, The step of controlling the robot device using the control model includes, in each state, determining a task parameter set using the trained parameter model; determining a state transition to be performed using the trained object model; determining a robot trajectory using the state transition to be executed and the task parameter set according to the robot trajectory model; controlling the robotic device to execute the determined robot trajectory; Including, 6. The method according to any one of claims 1 to 5.

7. The step of providing a control state sequence for an initial state-goal state pair comprises: selecting a state transition for each state of the robot device from the initial state to the goal state; determining the task parameter set assigned to the selected state transition using the precondition model assigned to the selected state transition; determining a state of the robotic device resulting from the selected state transition using the exit condition model assigned to the selected state transition; Including, The method of claim 6.

8. A computer program comprising program instructions for causing a computer to carry out the method according to any one of claims 1 to 7 when the computer runs the computer program.

9. A non-volatile storage medium storing the computer program according to claim 8.

10. An apparatus comprising the non-volatile storage medium of claim 9.

Citation Information

Patent Citations

  • Industrial robot obstacle avoidance path planning method based on machine learning

    CN109960880A

  • METHOD FOR PLANNING A MOVEMENT PATH OF A MANIPULATOR

    DE102016211470A1

  • Route planning device

    JP2006155559A

  • Trajectory searching device

    JP2009066692A

  • Operation control device of robot

    JP2019155561A