Task Planning Method and System for a Robot to Pick and Place Objects under Limited Observation
By improving the belief tree search algorithm and considering the feasibility of action, the unreliable problem of robots in picking and placing objects under observation is solved, and the task planning reliability and optimization performance are improved.
Patent Information
- Application Number
- CN202211729991.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In the case of limited robot observation, it is difficult for the prior art to effectively plan the task of picking and placing objects, especially when the object information is incomplete, resulting in unreliable planning results and poor optimization performance.
Improve the belief tree search algorithm, establish the action feasibility probability, set the action branch expansion of the belief tree based on the action feasibility probability, and design a correction return function to ensure that the action feasibility is fully considered in the planning.
It improves the planning reliability and optimization performance of robot picking and placing objects, ensuring feasibility of task execution in situations of limited observations.
Smart Images

Figure CN115958603B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot task planning. Specifically, it relates to a task planning method and system for a robot to pick up and place objects under limited observation, which is used to arrange the action sequence required for the robot to complete a task, especially for the planning of the robot's task of picking up and placing objects when the object information is incomplete. Background Art
[0002] In many cases, a robot needs to pick up and place different objects multiple times to achieve a certain task goal, such as arranging messy objects neatly or finding a target object among messy objects. Generally, the task planning algorithm for a robot to pick up and place objects assumes that all object information is completely known, including the name and pose of each object. When a robot can only obtain object information through its own observation, due to the mutual occlusion of objects, the robot is very likely unable to observe all the object information, which brings uncertainty to the planning.
[0003] Partially observable Markov decision process (POMDP) is a general model for describing decision-making processes under uncertainty. POMDP planning calculates the action with the maximum expected return under uncertainty. It is very difficult to calculate a complete policy offline in one go for POMDP planning. Generally, an online planning method is adopted, that is, only the action that the robot needs to execute in this step is planned at each step. After executing the action and re-observing, the action for the next step is planned. The widely used online POMDP planning method is belief tree search. Currently, relatively advanced belief tree search algorithms include DESPOT and POMCP, etc.
[0004] Although POMDP planning has been successfully applied in various industries, there are still deficiencies when applied to the task of a robot picking up and placing objects. Not all actions are feasible at each step in the task of a robot picking up and placing objects. However, the past belief tree search algorithms did not consider this point. When constructing a belief tree, the action branches expanded under each node usually come from the entire action space. Sometimes it is assumed that which actions are feasible at each step is known. When constructing a belief tree, the action branches expanded under each node come from the feasible actions. But which actions are feasible at each step in the task of a robot picking up and placing objects is not known in advance. Whether a pick-up or place object action is feasible depends on whether motion planning can find a feasible motion trajectory for this action.
[0005] Patent document CN113190012A (application number: 202110506117.7) discloses a robot task autonomous planning method and system. Among them, the method includes obtaining the semantic positions of static objects and the positional relationships between static objects and dynamic objects based on a home environment semantic knowledge model; based on the semantic positions of static objects and the positional relationships between static objects and dynamic objects, performing action planning according to a hybrid task planner until the task sequence executed by the robot completes the task; among them, during the process of performing action planning by the hybrid task planner, offline task planning is first performed, and it is determined whether the action impact of the offline task planning is deterministic to judge whether to continue executing the offline task sequence. When the action impact of the offline task planning is non-deterministic, then online action planning is performed. Patent document CN112131754A (application number: 202011060344.3) discloses an extended POMDP planning method and system based on a robot accompanying behavior model, including during the standard POMDP planning process, when the invariant of the task action aT being executed matches a certain observation action aO, forming an accompanying relationship between the task action aT and the observation action aO based on the matching predicate statement to form an accompanying behavior model; during the execution of the task action aT, obtaining the observation value obs of the observation action aO; updating the system knowledge base kb of the robot based on the invariant of the task action aT and the observation value obs; judging whether the truth value of the invariant in the knowledge base kb is false. If it holds, task replanning is triggered. The above two patents provide robot task planning algorithms based on POMDP, but do not consider the feasibility of actions
[0006] Patent document CN112356031A (application number: 202011220903.2) discloses an online planning method based on Kernel sampling strategy in an uncertain environment for the planning when a robot executes a task. In this uncertain environment, the uncertainty represented by the POMDP model is the main factor restricting the reliable operation of the robot; in the POMDP model, the robot can observe some of its own states, and the robot obtains the strategy with the maximum reward by continuously interacting with the environment; in the online planning method, when dealing with the observable part, the state of the robot is represented as a belief, denoted as belief, which belongs to a set of states, and the forward search is performed by constructing a belief tree through the POMDP algorithm to obtain the optimal strategy under the current belief; each node of the belief tree represents a belief, and the parent node and the child node are connected by an action-observation branch; the POMDP algorithm is the online POMDP planning algorithm Kernel-DESPOT. Patent document CN114118441A (application number: 202111401793.4) discloses an online planning method based on an efficient search strategy in an uncertain environment. The state of the robot is regarded as a belief. After initializing the upper and lower bounds of the current belief by the POMDP algorithm, the full information of the current belief is represented by discounting the upper and lower bounds, and then the forward search is performed to construct a belief tree to obtain the optimal strategy under the current belief; each node of the belief tree represents a belief, and the parent node and the child node are connected by an action-observation branch. The above two patents improve the DESPOT algorithm, but do not consider the action feasibility. For the task that the robot needs to pick up and place objects, if the objects are placed in a messy and crowded manner, not considering the action feasibility will lead to unreliable planning results and poor optimization performance. Summary of the Invention
[0007] Aiming at the defects in the prior art, the purpose of the present invention is to provide a task planning method and system for a robot to pick up and place objects under limited observation.
[0008] According to a task planning method for a robot to pick up and place objects under limited observation provided by the present invention, it includes:
[0009] Step S1: Improve the belief tree search algorithm, including: establishing the action feasibility probability, setting the belief tree action branch expansion based on the action feasibility probability, and setting the corrected reward function;
[0010] Step S2: Use the improved belief tree search algorithm for task planning and execution.
[0011] Preferably, the establishment of the action feasibility probability adopts:
[0012] Step S1.1: Under the observation z, put all possible feasible actions into the available action set A a (z);
[0013] Step S1.2: The characteristic quantity σ(z, a) of the action a ∈ A a (z) under the observation z should be able to reflect the action difficulty. Let the probability of the action a being feasible under this characteristic quantity be P(a|σ(z, a));
[0014] Step S1.3: Randomly arrange the objects that may appear in the task, obtain the observation z, record whether each action a ∈ A a (z) is feasible and the corresponding characteristic quantity σ(z, a). Repeat triggering Step S1.2 to Step S1.3, and calculate the frequency of the occurrence of feasible actions under each characteristic quantity σ as the estimated value of the action feasibility probability P(a|σ).
[0015] Preferably, the belief tree action branch expansion based on the action feasibility probability adopts:
[0016] For the root node b0, first screen out the available action set according to the current observation of the robot Then for each Call the motion planning to check whether it is feasible, and put the feasible actions into the set of feasible actions The action branches expanded under the root node b0 come from For other nodes b, first screen out the available action set A according to the parent observation branch z b Then for each a ∈ A a (z b ), extract the characteristic quantity σ(z a )(z b ), a), use a random number generator to generate a random number uniformly distributed on [0, 1]. If this random number is not greater than P(a|σ(z b , a)), then put a into the set of feasible actions A b )(z f ), and the action branches expanded under the node b come from A b )(z f ). b )
[0017] Preferably, the modified reward function adopts: the modified reward function Γ(h(s), a) ≤ 0, where h(s) is the action and observation history experienced from the start of the task to the state s, and it evaluates the executed action according to the historical information.
[0018] Preferably, Step S2 adopts:
[0019] Step S2.1: Update the current belief;
[0020] Step S2.2: Use the improved belief tree search algorithm to plan the actions that the robot needs to execute currently;
[0021] Step S2.3: The robot executes the planned actions and then obtains new observations;
[0022] Step S2.4: Update the action feasibility probability under the new observations; if the task is completed, end the task, otherwise repeat and trigger Steps S2.1 to S2.4.
[0023] Preferably, in Step S2.2, when constructing the belief tree, the action branches expanded under each node are determined according to the improved belief tree search algorithm; when searching the belief tree, the return value is the sum of the return function and the corrected return function.
[0024] Preferably, Step S2.4 adopts:
[0025] Let the set of available actions of the robot under the current observation z be A a (z), and the set of feasible actions obtained through motion planning be A f (z). For each a ∈ A a (z), extract the feature quantity σ(z, a), and update the action feasibility probability P(a|σ(z, a)) to
[0026]
[0027] where N > 0 is the period of the moving average, approximately regarded as estimating the probability using the recent N samples.
[0028] According to a task planning system for a robot to pick up and place an object under limited observations provided by the present invention, it includes:
[0029] Module M1: Improve the belief tree search algorithm, including: establishing the action feasibility probability, setting the expansion of the belief tree action branches based on the action feasibility probability, and setting the corrected return function;
[0030] Module M2: Use the improved belief tree search algorithm for task planning and execution.
[0031] Preferably, the establishment of the action feasibility probability adopts:
[0032] Module M1.1: Under the observation z, put all possible feasible actions into the set of available actions A a (z);
[0033] Module M1.2: Under the observation z, for the action a ∈ A aThe characteristic quantity σ(z, a) of (z) should be able to reflect the action difficulty. Let the probability that action a is feasible under this characteristic quantity be P(a|σ(z, a));
[0034] Module M1.3: Randomly arrange the objects that may appear in the task, obtain the observation z, and record each action a ∈ A a Whether (z) is feasible and the corresponding characteristic quantity σ(z, a). Repeat triggering Module M1.2 to Module M1.3, calculate the frequency of the occurrence of feasible actions under each characteristic quantity σ, and use it as an estimated value of the action feasibility probability P(a|σ);
[0035] The belief tree action branch expansion based on the action feasibility probability adopts:
[0036] For the root node b0, first according to the current observation of the robot Screen out the available action set Then for each Call the motion planning to check whether it is feasible, and put the feasible actions into the available action set Among them, the action branches expanded under the root node b0 come from For other nodes b, first according to the parent observation branch z b Screen out the available action set A a (z b ), then for each a ∈ A a (z b ), extract the characteristic quantity σ(z b , a), use the random number generator to generate a random number uniformly distributed on [0, 1]. If this random number is not greater than P(a|σ(z b , a)), then put a into the available action set A f (z b ), and the action branches expanded under the node b come from A f (z b );
[0037] The modified reward function adopts: The modified reward function Γ(h(s), a) ≤ 0, where h(s) is the action and observation history experienced from the start of the task to state s, and it evaluates the executed action according to the historical information.
[0038] Preferably, the module M2 adopts:
[0039] Module M2.1: Update the current belief;
[0040] Module M2.2: Use the improved belief tree search algorithm to plan the action that the robot needs to execute currently;
[0041] Module M2.3: The robot executes the planned actions and then obtains new observations;
[0042] Module M2.4: Update the action feasibility probability under the new observation; if the task is completed, end the task, otherwise repeat to trigger Modules M2.1 to M2.4;
[0043] The said Module M2.2 adopts: when constructing the belief tree, the action branches expanded under each node are determined according to the improved belief tree search algorithm; when searching the belief tree, the return value is the sum of the return function and the corrected return function;
[0044] The said Module M2.4 adopts:
[0045] Let the set of available actions of the robot under the current observation z be A a (z), and the set of executable actions obtained through motion planning be A f (z). For each a ∈ A a (z), extract the feature quantity σ(z, a), and update the action feasibility probability P(a|σ(z, a)) to
[0046]
[0047] where N > 0 is the period of the moving average, which is approximately regarded as estimating the probability using the recent N samples.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] 1. When using the POMDP online planning algorithm based on belief tree search to plan the task of the robot picking up and placing objects, the action branches expanded under each node of the constructed belief tree come from the set of executable actions under the parent observation branch, fully considering the action feasibility and ensuring the reliability of the planning.
[0050] 2. Establish the action feasibility probability. The feasibility of the action branches expanded under the nodes except the root node of the belief tree is determined by sampling the action feasibility probability, saving the number of calls of the motion planning and improving the planning speed.
[0051] 3. Design the corrected return function to adjust the belief tree search direction in real time according to the historical information of the task execution and suppress the undesired search direction. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objects and advantages of the present invention will become more obvious:
[0053] Figure 1 It is an example of a belief tree.
[0054] Figure 2 For the planning and execution process.
[0055] Figure 3 For the task scenario of the implementation example.
[0056] Figure 4 For the task objective of the implementation example. Detailed implementation manners
[0057] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all belong to the protection scope of the present invention.
[0058] Example 1
[0059] The present invention is used in the task where a robot needs to pick up and place an object, planning what action the robot should perform at each step, that is, planning the operation sequence of picking up and placing the object. Assume that the robot can only rely on its own observation of the object information, and the objects are randomly placed, and all object information may not be directly obtained due to mutual occlusion. The present invention makes some improvements on the basis of the POMDP online planning algorithm using belief tree search, so that the action feasibility is fully considered in the planning, the reliability of the planning is improved, and at the same time, the computational amount of the planning is not significantly increased.
[0060] The algorithm of the present invention requires a POMDP model of the known task process. The POMDP model and the planning principle are briefly introduced below.
[0061] The discrete POMDP model is defined as a tuple (S, A, Z, T, O, R), where S, A, and Z represent the state space, action space, and observation space respectively, and T, O, and R represent the transition function, observation function, and reward function respectively. At each step, the robot executes an action a ∈ A in the state s ∈ S and obtains a reward r = R(s, a), then the state transfers to s' ∈ S with probability T(s'|s, a), and the robot obtains a new observation z' ∈ Z with probability O(z'|s', a).
[0062] In the task of a robot picking up and placing objects, the general state is defined as the names and poses of all objects. The observation is defined as the names and poses of the objects directly observed by the robot. The action is defined as picking up a certain object or placing an object in a certain pose. The transition function and the observation function are determined according to the specific task model, and there are also some methods to learn the transition function and the observation function. The reward function is determined according to the optimization goal. Generally, it is hoped that the task can be completed as quickly as possible. Each action execution incurs a certain time cost, corresponding to a negative reward.
[0063] Since the true state is unknown, the robot needs to continuously maintain a belief b ∈ B, which represents the probability distribution in the state space. B is the set of all possible beliefs, called the belief space. b(s) represents the probability that the state in belief b is s. At each step, the robot updates the original belief b to b′ according to the executed action a and the obtained observation z′, and the update process is represented by b′ = τ(b, a, z′).
[0064] The policy is a mapping π from the belief space to the action space: Executing according to the policy means that the robot executes the action corresponding to the current belief according to the policy at each step. The value function V π (b) represents the expected total discounted reward obtained by executing according to the policy π starting from the belief b:
[0065]
[0066] where γ ∈ [0, 1] is the discount factor, representing the importance of the long-term reward relative to the short-term reward. POMDP planning is to find a policy that maximizes the value function as much as possible.
[0067] It is difficult to calculate the complete policy at once. Generally, only the action to be executed currently is calculated through online planning at each step. The widely used online POMDP planning algorithm is belief tree search. At each step, the current belief is used as the root node b0 of the belief tree, and many child nodes are expanded through action branches and observation branches. Figure 1 Shows an example of a belief tree. The hollow circles on the tree are nodes, representing a belief; the line segments connecting the nodes are branches, representing an action or an observation; the solid dots are the connection points of the action branches and the observation branches. The total depth of the belief tree is the planning step length, representing the maximum number of steps considered in the planning.
[0068] Use Δ(b) to represent the depth of node b on the belief tree. A b represents the set of action branches expanded from node b; Z b,a represents the set of observation branches expanded from node b through action branch a; τ(b, a, z) represents the child node expanded from node b through action branch a and observation branch z; B bDenotes the set of all child nodes expanded by node b. If node b has no child nodes, b is called a leaf node. z b Denotes the parent observation branch of node b; a b Denotes the parent action branch of node b; τ -1 (b) Denotes the parent node of node b. The root node b0 has no parent branches and parent nodes. Denote as the virtual parent observation branch of b0, representing the current observation; Denote as the virtual parent action branch of b0, representing the action executed in the previous step; Denote τ -1 (b0) is the virtual parent node of b0, representing the belief in the previous step, Δ(τ -1 (b0)) = -1.
[0069] The general method for using the belief tree to search for the optimal policy π * is dynamic programming. First, estimate the values of all leaf nodes through a heuristic method, and the optimal value V of other nodes * can be calculated through the Bellman equation:
[0070]
[0071] By recursively calculating layer by layer from the leaf nodes upwards, the optimal value V * (b0) and the corresponding action π * (b0) can be obtained. Here, the policy π * only involves the nodes on the belief tree, so it is not a complete policy.
[0072] According to a task planning method for a robot to pick up and place an object under limited observation provided by the present invention, it includes:
[0073] Step S1: Improve the belief tree search algorithm, including: establishing the action feasibility probability, setting the expansion of the belief tree action branch based on the action feasibility probability, and setting the modified reward function;
[0074] Step S2: Use the improved belief tree search algorithm for task and planning and execution.
[0075] The present invention can be combined with any POMDP online planning algorithm that uses belief tree search. When implementing the present invention, it is first necessary to determine which belief tree search algorithm to specifically select. At the same time, the present invention fully considers the feasibility of actions in the planning. Whether an action is feasible depends on whether the motion planning can find a feasible motion trajectory for this action. Therefore, it is also necessary to select a motion planning algorithm suitable for the task actions.
[0076] The establishment of the action feasibility probability adopts:
[0077] Not all actions are feasible at every step during the task. The present invention defines the available action set and the executable action set. The available action set A a (z) represents the set of actions that may be feasible under the observation z. The executable action set A f (z) represents the set of actions that the robot can execute after inspection under the observation z. After obtaining an observation z, first initially screen out the available action set A a (z), and then check the feasibility of each action in A a (z) to obtain the executable action set A f (z) under this observation. At each step, the robot can only execute the actions in the executable action set under the current observation.
[0078] Step S1.1: Design the screening rule for the available action set.
[0079] According to the observation z, some actions are obviously infeasible, while some actions may be feasible. The rule for screening the available action set is to put those actions that may be feasible into the available action set A a (z) based on simple logical judgments. For example, when the robot does not hold an object, picking up an unobstructed object may be feasible; when the robot holds an object, placing the object at an unobstructed and unoccupied pose may be feasible.
[0080] Step S1.2: Design the calculation method for the action feature quantity.
[0081] Let σ(z, a) be a certain feature quantity of the available action a ∈ A a (z) calculated according to the observation z. Let the probability that the available action a is feasible under this feature quantity be P(a|σ(z, a)). The calculation method of the feature quantity is not unique, but it should be able to reflect the difficulty of the action, thus being related to the probability of the action being feasible. Generally, it can be calculated by the relative position relationship between the pose where the robot picks up the object or the pose where the robot places the object and the surrounding objects. The actions of the robot picking up and placing the object are reciprocal. If the robot can pick up the object o at the pose p, then it can also place the object o back at the pose p. Therefore, the feature quantities of the picking-up and placing actions with reciprocal relationships should be the same. For easy maintenance, the feature quantity can be discretized, so that only individual independent probabilities P(a|σ) need to be saved.
[0082] Step S1.3: Estimate the action feasibility probability value.
[0083] Next, estimate the specific value of the action feasibility probability. Randomly arrange the objects that may appear in the task and obtain the observation z. Screen out the available action set A a (z) under the observation z. For each action a ∈ Aa (z), call motion planning to check its feasibility and record the corresponding feature quantity σ(z, a). Repeat the above steps a certain number of times, calculate the frequency of the occurrence of feasible actions under each feature quantity σ, and use it as an estimated value of the action feasibility probability P(a|σ). The probability value estimation can be carried out in simulation to save time. With the estimated value of the action feasibility probability, POMDP online planning can be carried out. During the subsequent task execution, the action feasibility probability value can also be continuously updated.
[0084] The belief tree action branch expansion based on the action feasibility probability is adopted as follows:
[0085] All belief tree search algorithms need to construct a belief tree. The action branches expanded under each node b on the belief tree must come from the set of feasible actions A under the parent observation branch z b under it. f (z b ). The motion planning of a multi-joint manipulator needs to search for a feasible motion trajectory in a high-dimensional space. If it is checked by motion planning for which action branches can be expanded under each node, it will take a lot of time. And online POMDP planning requires planning the next action immediately after each step is executed, and has strict requirements for the planning time.
[0086] When constructing a belief tree, the action branches expanded under the root node represent the actions that the robot may currently execute, and the action feasibility needs to be strictly checked. However, the action branches expanded under other nodes have nothing to do with the actions actually executed by the robot, and the accuracy requirements for action feasibility are not strict. At this time, the action feasibility can be sampled using the action feasibility probability.
[0087] For the root node b0, first filter out the available action set according to the current observation of the robot Then for each Call motion planning to check its feasibility, and put the feasible actions into the set of feasible actions The action branches expanded under the root node b0 come from For other nodes b, first filter out the available action set A according to the parent observation branch z For other nodes b, first filter out the available action set A(z b ) according to the parent observation branch z. Then for each a ∈ A(z a ), extract the feature quantity σ(z b ), generate a random number uniformly distributed on [0, 1] using a random number generator. If this random number is not greater than P(a|σ(z a ), then put a into the set of feasible actions A(z b ). Then for each a ∈ A(z b ), extract the feature quantity σ(z b , a), generate a random number uniformly distributed on [0, 1] using a random number generator. If this random number is not greater than P(a|σ(z f ), then put a into the set of feasible actions A(z b) Among them. The action branches expanded under node b come from A f (z b ). In this way, it not only reasonably simulates the feasibility of actions in future steps but also saves the number of calls for motion planning.
[0088] The designed modified reward function adopts:
[0089] Since the feasibility of the action branches on the belief tree cannot be guaranteed to be completely accurate, it may sometimes mislead the belief tree search direction. The present invention designs a modified reward function Γ(h(s), a) ≤ 0, where h(s) is the action and observation history experienced from the start of the task to state s. The modified reward function is only used for belief tree search and does not represent the actual reward. It evaluates the executed actions based on historical information and can artificially suppress the undesired search directions.
[0090] Such as Figure 2 As shown, the step S2 adopts:
[0091] Step S2.1: Update the current belief. If it is the first step of the task, establish the initial belief according to the belief representation method provided by the belief tree search algorithm. If it is not the first step of the task, update the current belief according to the belief update method provided by the belief tree search algorithm.
[0092] Step S2.2: Use the belief tree search algorithm to plan the actions that the robot needs to execute currently. When constructing the belief tree, the action branches expanded under each node are determined according to step 3. When searching the belief tree, the reward value r = R(s, a) + Γ(h(s), a), which is the sum of the reward function and the modified reward function designed in step 4.
[0093] Step S2.3: The robot executes the planned actions and then obtains new observations.
[0094] Step S2.4: Update the action feasibility probability. Let the set of available actions under the current observation z of the robot be A a (z), and the set of executable actions obtained through motion planning be A f (z). For each a ∈ A a (z), extract the feature quantity σ(z, a), and update the action feasibility probability P(a|σ(z, a)) to
[0095]
[0096] where N > 0 is the period of the moving average, which can be approximately regarded as estimating the probability using the recent N samples.
[0097] Step S2.5: If the task is completed, end the task; otherwise, return to step 5.1.
[0098] In summary, the present invention can be combined with any online POMDP planning algorithm that uses belief tree search. Different planning algorithms may have different special treatments for belief representation, belief update, belief tree construction, and belief tree search methods, but all are applicable to the content of the present invention.
[0099] The task planning system for a robot to pick up and place objects under limited observation provided by the present invention can be implemented through the step process in the task planning method for a robot to pick up and place objects under limited observation provided by the present invention. Those skilled in the art can understand the task planning method for a robot to pick up and place objects under limited observation as a preferred example of the task planning system for a robot to pick up and place objects under limited observation.
[0100] Example 2
[0101] Embodiment 2 is a preferred example of Embodiment 1
[0102] The task instance scenario is as Figure 3 shown. There is a shelf in front of the robot. The support surface of the shelf is 40 cm long and 30 cm wide. Six objects are placed on the shelf. The camera is fixed in front of the shelf and is at the same horizontal plane as the objects on the shelf. The task reference coordinate system has the x direction pointing in front of the robot and the y direction pointing to the left of the robot. All objects are cylinders with a diameter of 6 cm and a height of 12 cm, and each object is distinguished by a different name and color. The robot can pick up objects with known names and positions, and can also place the held objects on the shelf or outside the shelf. At the beginning, the robot does not know the positions of the objects, and the task goal is to arrange all the objects into Figure 4 the formation shown.
[0103] First, the POMDP model of the task is given. The state is defined as s = {(o i , p i ) | i = 1, 2,..., 6}, where o i is the name of the i-th object in the state, p i is the position of the object o i , and if o i is held by the robot, then p i = robot. Assume that an object can be observed if more than half of its projected area is visible in the camera image, and an object held by the robot can also be observed. The observation is defined as z = {(o i , p i ) | i = 1, 2,..., m}, where m ≤ 6 is the number of observed objects, o i is the name of the i-th object in the observation, p i is the position of the object o iThe position. The action types are divided into two types: operation actions and termination actions. The operation action is defined as a = (o, p). When p = robot, it represents the action of picking up the object o. Otherwise, it represents the action of placing the object o at the position p. At this time, the value of p can be outside the shelf or Figure 4 a position in the formation. The termination action is represented by a = stop, which does not perform any operation and is only used to terminate the task. In order to minimize the number of operation actions required to complete the task, it is assumed that each operation action incurs a reward of -10. In addition, if the task is completed, performing the termination action incurs a reward of 300, and if the task is not completed, performing the termination action incurs a reward of -∞. In the planning, it is desired that relatively long-term rewards are as important as relatively recent rewards, and the discount factor is set to γ = 0.99.
[0104] Step 1: Select the belief tree search algorithm and the motion planning algorithm.
[0105] In this example, the DESPOT is selected as the belief tree search algorithm. Using the particle belief representation method, a scenario-based sparse belief tree is constructed, and the search can be set to end at any time. The RRT-connect is selected as the motion planning algorithm, which is a variant of the Rapidly-Exploring Random Tree (RRT) and belongs to the sampling-based motion planning method.
[0106] Step 2: Establish the action feasibility probability.
[0107] Step 2.1: Design the screening rules for the available action set.
[0108] Under the observation z = {(o i , p i )|i = 1, 2,..., m}, if the robot does not hold an object, picking up an unobstructed object is an available action; if the robot holds an object, placing the object at an unobstructed empty position is an available action.
[0109] Step 2.2: Design the calculation method for the action feature quantity.
[0110] In this example, whether the picking action is feasible depends to a large extent on the relative position between other objects and the object to be picked up. If other objects are behind the object to be picked up, then there is almost no impact. Otherwise, the distance between other objects and the object to be picked up in the y direction mainly determines the degree to which the robot is obstructed when picking up the object. Placing is the reverse process of picking up, and the degree to which the placing action is obstructed by other objects is the same as that in the picking action.
[0111] Under the observation z = {(o i , p i )|i = 1, 2,..., m}, the feature quantity of the action of picking up the object a = (o i , robot) is where int is the floor function, thus discretizing the feature quantity at a resolution of 0.3. σ(o i , o j ) is the obstruction caused by picking up o j . Let the coordinates of o i in the plane be (x i , y i ). When o i , o i , o j are all on the shelf and i ≠ j, x i ≥ x j , In other cases, σ(o i , o j ) = 0.
[0112] Under the observation z = {(o i , p i )|i = 1, 2,..., m}, the feature quantity of the action a = (o i , p) for placing an object is σ(p, o j ) is the obstruction caused by placing the object at p. Let the coordinates of p in the plane be (x j , y p ). When p, o p are all on the shelf and i ≠ j, x j ≥ x p , j When In other cases, σ(p, o j ) = 0.
[0113] Step 2.3: Estimate the action feasibility probability value.
[0114] In the simulation, 5000 object layouts are randomly generated. For each layout, the unoccluded objects are checked for pickability using motion planning, and the feature quantities of the corresponding actions are recorded. The frequency of action feasibility for each feature quantity is calculated, and the action feasibility probability estimate values are shown in Table 1. In this example, the value range of the action feature quantity is 0 to 15. The larger the feature quantity, the lower the probability of action feasibility.
[0115] Table 1
[0116] σ 0 1 2 3 4 5 6 7 P(a|σ) 1.0 1.0 0.9967 0.9794 0.7534 0.5526 0.3480 0.2356 σ 8 9 10 11 12 13 14 15 P(a|σ) 0.1304 0.0095 0.0 0.0 0.0 0.0 0.0 0.0
[0117] Step 3: Design the action branch expansion rule for the belief tree.
[0118] For the root node b0, first filter out the available action set according to the current observation of the robot Then for each Then for each Call motion planning to check its feasibility, and put the feasible actions into the set of feasible actions Among them. The action branches extended under the root node b0 come from For other nodes b, first filter out the set of available actions A according to the parent observation branch z b Screen out the set of available actions A a (z b ). Then for each a ∈ A a (z b ), extract the feature quantity σ(z b , a), use the random number generator to generate a random number uniformly distributed on [0, 1]. If this random number is not greater than P(a|σ(z b , a)), then put a into the set of feasible actions A f (z b ). The action branches extended under the node b come from A f (z b ).
[0119] Step 4: Design a modified reward function.
[0120] Since the feasibility of most action branches on the belief tree is determined by sampling according to the action feasibility probability, it cannot be guaranteed to be completely accurate, and it may cause the robot to enter a dead end of repeatedly executing certain actions. Design a modified reward function where h(s) is the history of actions and observations experienced from the start of the task to the state s, N exe (a) is the number of times the action a has been executed in h(s), which can suppress the actions that have been repeatedly executed many times in the past. L(h(s)) is the number of steps experienced in h(s), and L exe is the number of steps actually executed by the robot. The modified reward function takes effect only when L(h(s)) = L exe , which is equivalent to taking effect only on the action branches under the root node. Because only the repeatability of the action results of each step of the planning needs to be suppressed, it is not appropriate to limit the exploration of deeper action branches on the belief tree.
[0121] Step 5: Task planning and execution.
[0122] Step 5.1: Update the current belief. If it is the first step of the task, establish the initial belief according to the particle belief representation method adopted by DESPOT. If it is not the first step of the task, update the current belief according to the particle filtering method adopted by DESPOT.
[0123] Step 5.2: Use the DESPOT algorithm to plan the actions that the robot needs to execute currently. When constructing the belief tree, the action branches expanded under each node are determined according to Step 3. When searching the belief tree, the reward value r = R(s, a) + Γ(h(s), a), which is the sum of the reward function and the corrected reward function designed in Step 4.
[0124] Step 5.3: The robot executes the planned actions and then obtains new observations.
[0125] Step 5.4: Update the action feasibility probability. Let the set of available actions of the robot under the current observation z be A a (z), and the set of executable actions obtained through motion planning be A f (z). For each a ∈ A a (z), extract the feature quantity σ(z, a), and update the action feasibility probability P(a|σ(z, a)) to
[0126]
[0127] where N = 100 is the period of the moving average.
[0128] Step 5.5: If the task is completed, end the task; otherwise, return to Step 5.1.
[0129] It should be understood that the present invention is not limited to the above specific embodiments. Those skilled in the art can make various changes and modifications within the scope of the claims, which do not affect the essence of the present invention. The present invention can be combined with any online POMDP planning algorithm that adopts belief tree search. Different methods may have different special treatments for belief representation, belief update, belief tree construction, and belief tree search methods. The present invention algorithm can be used with different belief tree search algorithms.
[0130] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as both software programs for implementing the methods and the structures within the hardware component.
[0131] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be arbitrarily combined with each other.
Claims
1. A task planning method for a robot to pick up and place objects under limited observation, characterized in that Including: Step S1: Improve the belief tree search algorithm, including: establishing the action feasibility probability, setting the belief tree action branch expansion based on the action feasibility probability, and setting the modified reward function; Step S2: Use the improved belief tree search algorithm for task planning and execution; The establishment of the action feasibility probability adopts: Step S1.1: Under the observation z, put all possible feasible actions into the available action set A a (z); Step S1.2: Observe action a ∈ A under observation z a (z)'s characteristic quantity σ(z,a) can reflect the action difficulty. Let the probability that action a is feasible under this characteristic quantity be P(a|σ(z,a)); Step S1.3: Randomly arrange the objects that may appear in the task, obtain the observation z, and record each action a ∈ A a (z) Whether it is feasible and the corresponding feature quantity σ(z, a), repeat triggering Step S1.2 to Step S1.3, calculate the frequency of the occurrence of feasible actions under each feature quantity σ, and use it as an estimate of the action feasibility probability P(a|σ(z, a)); The setting of the belief tree action branch expansion based on the action feasibility probability adopts: For the root node b0, first, based on the current observation of the robot filter out the set of available actions Then, for each call the motion planning to check its feasibility, and put the feasible actions into the set of feasible actions The action branches expanded under the root node b0 come from For other nodes b, first, based on the parent observation branch z b filter out the set of available actions A a (z b ). Then, for each a ∈ A a (z b ), extract the feature quantity σ(z b , a), use the random number generator to generate a random number uniformly distributed on [0, 1]. If this random number is not greater than P(a|σ(z b , a)), then put a into the set of feasible actions A f (z b ). The action branches expanded under the node b come from A f (z b ); The modified reward function adopts: the modified reward function Γ(h(s), a) ≤ 0, where h(s) is the action and observation history experienced from the start of the task to state s, and it evaluates the executed action according to the historical information; The step S2 adopts: Step S2.1: Update the current belief; Step S2.2: Use the improved belief tree search algorithm to plan the action that the robot needs to execute currently; Step S2.3: The robot executes the planned action and then obtains a new observation; Step S2.4: Update the action feasibility probability under the new observation; if the task is completed, end the task, otherwise repeat and trigger steps S2.1 to S2.4; The step S2.2 adopts: When constructing the belief tree, the action branches expanded under each node are determined according to the improved belief tree search algorithm; when searching the belief tree, the reward value is the sum of the reward function and the modified reward function; The step S2.4 adopts: Let the set of available actions under the current observation \(z\) of the robot be \(A^{(z)}\), and the set of executable actions obtained through motion planning be \(A^{(z)}\). a (z). For each \(a\in A^{(z)}\), f (z), extract the feature quantity \(\sigma(z, a)\), and update the action feasibility probability \(P(a|\sigma(z, a))\) according to the moving average method to a (z). Where N > 0 is the period of the moving average, approximately regarded as using the recent N samples for probability estimation.
2. A task planning system for a robot to pick up and place objects under limited observation, characterized in that, Including: Module M1: Improve the belief tree search algorithm, including: establishing the action feasibility probability, setting the belief tree action branch expansion based on the action feasibility probability, and setting the modified reward function; Module M2: Use the improved belief tree search algorithm for task planning and execution; The establishment of the action feasibility probability adopts: Module M1.1: Under the observation z, put all possible feasible actions into the available action set A a (z); Module M1.2: Action a ∈ A under Observation z a The characteristic quantity σ(z, a) of (z) can reflect the action difficulty. Let the probability that action a is feasible under this characteristic quantity be P(a|σ(z, a)); Module M1.3: Randomly arrange the objects that may appear in the task, obtain the observation z, and record each action a ∈ A a Check whether (z) is feasible and the corresponding feature quantity σ(z,a). Repeat triggering Module M1.2 to Module M1.3, calculate the frequency of the occurrence of feasible actions under each feature quantity σ, and use it as an estimated value of the action feasibility probability P(a|σ(z,a)); The setting of the belief tree action branch expansion based on the action feasibility probability adopts: For the root node b0, first, based on the current observation of the robot filter out the set of available actions Then, for each call motion planning to check its feasibility, and put the feasible actions into the set of executable actions The action branches expanded under the root node b0 come from For other nodes b, first, based on the parent observation branch z b filter out the set of available actions A a (z b ), and then for each a ∈ A a (z b ), extract the feature quantity σ(z b , a), use a random number generator to generate a random number uniformly distributed on [0, 1]. If this random number is not greater than P(a|σ(z b ), a)), then put a into the set of executable actions A f (z b ), and the action branches expanded under the node b come from A f (z b ); The modified reward function adopts: the modified reward function Γ(h(s), a) ≤ 0, where h(s) is the action and observation history experienced from the start of the task to state s, and it evaluates the executed action according to the historical information; The module M2 adopts: Module M2.1: Update the current belief; Module M2.2: Use the improved belief tree search algorithm to plan the action that the robot needs to execute currently; Module M2.3: The robot executes the planned action and then obtains a new observation; Module M2.4: Update the action feasibility probability under the new observation; if the task is completed, end the task, otherwise repeat and trigger modules M2.1 to M2.4; The module M2.2 adopts: When constructing the belief tree, the action branches expanded under each node are determined according to the improved belief tree search algorithm; when searching the belief tree, the reward value is the sum of the reward function and the modified reward function; The module M2.4 adopts: Let the set of available actions under the current observation \(z\) of the robot be \(A_{ a}(z)\), and the set of actionable actions obtained through motion planning be \(A_{ f}(z)\). For each \(a\in A_{ a}(z)\), extract the feature quantity \(\sigma(z,a)\), and update the action feasibility probability \(P(a|\sigma(z,a))\) according to the moving average method to a (z), the set of actionable actions obtained through motion planning is \(A_{ f}\) f (z), for each \(a\in A_{ a}\) a (z), extract the feature quantity \(\sigma(z,a)\), and update the action feasibility probability \(P(a|\sigma(z,a))\) according to the moving average method to Where N > 0 is the period of the moving average, approximately regarded as using the recent N samples for probability estimation.
Citation Information
Patent Citations
Extended POMDP planning method and system based on robot accompanying behavior model
CN112131754A
Extended POMDP planning method and system based on robot adjoint behavior model
CN112131754B
Online planning method based on Kernel sampling strategy in uncertain environment
CN112356031A
An Online Planning Method Based on Kernel Sampling Strategy in Uncertain Environments
CN112356031B
Robot task autonomous planning method and system
CN113190012A