Contact Point Selection Method and Device for Graspability Combining Dual Manipulators
The multi-agent value iteration algorithm optimizes grip points for dual-arm robots, addressing inefficiencies in existing methods by enhancing cooperation and efficiency in tasks like valve turning.
Patent Information
- Application Number
- CN202310045310.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-01-30
AI Technical Summary
In robotic double-arm cooperating operation, how to select the appropriate operating point to improve efficiency and save labor to complete tasks, such as the rotation of the valve, existing methods are limited to specific tasks and fail to fully utilize the synergistic advantages of the double-arm.
The multi-agent iterative framework is adopted to evaluate the grasping point of the dual robotic arms through a value function, and combine spatial accessibility and task-related value to optimize the selection of the grasping point of the two arms.
In the coordinated operation of the robot with both arms, the optimal grasping point is selected to efficiently complete the valve rotation task, giving full play to the coordinated advantages of both arms, and improving operational efficiency and labor saving.
Smart Images

Figure CN116175562B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information science, and in particular to a contact point selection method and device for a dual-arm robot that combines the graspability of a dual manipulator. Background Art
[0002] In the process of a robot's dual-arm cooperative operation to complete complex tasks, the operation can be defined as physically interacting with an object and applying force to the object to move or reshape it. For a dual-arm system, due to the high degrees of freedom of the robot arms, there is usually a redundancy of degrees of freedom, that is, there are a large number of feasible operation points. However, different operation points will greatly affect the efficiency of interacting with the object. For example, for a valve rotation task, as Figure 1 shown, different grasping points on the valve will also have a great impact. Therefore, how to select appropriate operation points for grasping, that is, how to select the poses of the dual-arm ends interacting with the object, is of great significance for a dual-arm robot.
[0003] In previous work, the analysis methods for the reachable space of a robot mainly include geometric method, analytical method, and numerical method. However, from the perspective of humans, in addition to simply analyzing the reachability of the arm end, the operability of the task also needs to be considered. For example, for common arm operation tasks such as screwing the cap of a bottle or turning the steering wheel, people usually select the space in front of the chest as the operation space because in this area, the joint angles of the arm are in a moderate state and can well adapt to the subsequent operations. If the operation is carried out at the farthest reachable position of the arm, it may result in a situation where the movement can no longer continue. For a robot, how to abstract the "operability" of different operation points in the working space of the robot arm is a prerequisite for path planning or task planning. Yoshikawa proposed the concept of "Manipulability" in 1985 and obtained the manipulability index of a redundant manipulator by analyzing the manipulability ellipsoid formed by the singular vectors of the Jacobian matrix of the robot arm. To be more applicable to actual operations, many researchers will discretize the reachable space of the robot into 3D or 6D grids and calculate the manipulability of each grid. For the specific selection of the operation points at the dual-arm ends, previous researchers have used methods such as random sampling or dynamic analysis, but they are all relatively limited to specific tasks.
[0004] When a person performs a collaborative operation with both arms, such as folding clothes with both hands or unscrewing a bottle cap, he or she will adjust his or her own position or the position of the object being operated to make the operation point within the reachable space of his or her arms. In addition, people will choose different operation points for the objects being operated according to different tasks to improve the efficiency of subsequent work. For example, when rotating a valve, when both hands symmetrically grasp the valve, the advantages of both arms can be brought into play, and the valve can be turned more effortlessly. Therefore, when the robot's two arms are operating collaboratively, it is necessary to learn from the relevant movement mechanisms of humans, and by combining the workspace accessibility and task-related value assessment methods, calculate the value functions of different valve grasping points in the robot's two-arm workspace, and ultimately enable the two arms to efficiently complete the work at the selected valve grasping point. Summary of the invention
[0005] In order to solve the above problems, the present invention proposes a contact point selection method and device combined with the graspability of dual robotic arms. The method selects appropriate operating points for the specific dual-arm grasping valve task. The method regards each of the two arms as an intelligent agent, independently selects their own operating point posture, and introduces a multi-agent value iteration framework to evaluate the robot dual-arm grasping points.
[0006] The technical contents of the present invention include:
[0007] A contact point selection method combining dual robotic arm graspability, the method comprising:
[0008] The target operation area is gridded according to the width of the end grippers of the dual robotic arms; the dual robotic arms include: a robotic arm A and a robotic arm B;
[0009] Combine the graspability of robot B in the grid and the current grasping point position of robot A to update the value function V of robot B B With the final value function , and based on the value function V B , calculate the updated grasping point position of robot arm B; wherein, the value function represents the value function of the current state considering spatial accessibility and task relatedness, and the final value function represents the value function of the expected value of the next action combined with the posture state of the previous step;
[0010] Combine the graspability of robot A in the grid and the updated grasping point position of robot B to update the value function V of robot A. A With the final value function , and based on the value function V A , calculate the updated grasping point position of robot arm A;
[0011] Based on the value function V A , the final value function , the value function VB With the final value function , determine whether the updated grasping point positions of robotic arm A and robotic arm B are the optimal positions;
[0012] In the case where the updated grasping point positions of robotic arm A and robotic arm B are not the optimal positions, based on the updated grasping point positions of robotic arm A and robotic arm B, return to the step of updating the value function V of robotic arm B according to the graspability of robotic arm B in the grid and the current grasping point position of robotic arm A B With the final value function ;
[0013] In the case where the updated positions of robotic arm A and robotic arm B are the optimal positions, calculate the grasping poses of robotic arm A and robotic arm B based on the updated grasping point positions of robotic arm A and robotic arm B.
[0014] Further, the step of updating the value function V of robotic arm B by combining the graspability of robotic arm B in the grid and the current grasping point position of robotic arm A B , includes:
[0015] According to the number of sampling points in the grid where the end of robotic arm B is located, calculate the graspability M(d B );
[0016] Update the value function V of robotic arm B B (d B ) = αM(d B ) - βcos(|d A - d B |); where α represents the manipulability weight, β represents the relative position weight of the two arms, d A represents the current grasping point position of robotic arm A, and d B represents the current grasping point position of robotic arm B.
[0017] Further, the step of calculating the updated grasping point position of robotic arm B based on the value function V B , includes:
[0018] The updated grasping point position d' B of robotic arm B = arg max(V A ).
[0019] Further, the step of determining whether the updated grasping point positions of robotic arm A and robotic arm B are the optimal positions based on the value function V A , the final value function , the value function V B and the final value function , includes:
[0020] Calculate the current value where d' A represents the position of the grasping point of the updated robotic arm A, and d' B represents the position of the grasping point of the updated robotic arm B;
[0021] By comparing the current value P with the set value δ, it is determined whether the positions of the grasping points of the updated robotic arm A and robotic arm B are the optimal positions.
[0022] A contact point selection device combining the graspability of dual robotic arms, the device includes:
[0023] A discretization module for meshing the target operation area according to the width of the end grippers of the dual robotic arms; the dual robotic arms include: robotic arm A and robotic arm B;
[0024] A first update module for updating the value function V of robotic arm B by combining the graspability of robotic arm B in the grid and the current grasping point position of robotic arm A B with the final value function and based on the value function V B calculate the position of the updated grasping point of robotic arm B; where the value function represents the value function considering spatial accessibility and the value of the current state related to the task, and the final value function represents the value function of the expected value for the next action by combining the pose state of the previous step;
[0025] A second update module for updating the value function V of robotic arm A by combining the graspability of robotic arm A in the grid and the updated grasping point position of robotic arm B A with the final value function and based on the value function V A calculate the position of the updated grasping point of robotic arm A;
[0026] A position judgment module for judging whether the positions of the updated grasping points of robotic arm A and robotic arm B are the optimal positions based on the value function V A the final value function the value function V B and the final value function judge whether the positions of the updated grasping points of robotic arm A and robotic arm B are the optimal positions; in the case where the positions of the updated grasping points of robotic arm A and robotic arm B are not the optimal positions, based on the positions of the updated grasping points of robotic arm A and robotic arm B, return to updating the value function V of robotic arm B according to the graspability of robotic arm B in the grid and the current grasping point position of robotic arm A B with the final value function When the updated positions of the robotic arm A and the robotic arm B are the optimal positions, calculate the grasping poses of the robotic arm A and the robotic arm B according to the updated grasping point positions of the robotic arm A and the robotic arm B.
[0027] Further, the first update module is configured to:
[0028] Calculate the graspability M(d of the robotic arm B according to the number of sampling points in the grid where the end of the robotic arm B is located B )
[0029] Update the value function V of the robotic arm B B (d B ) = αM(d B ) - βcos(|d A - d B |); where α represents the manipulability weight, β represents the relative position weight of the two arms, d A represents the current grasping point position of the robotic arm A, and d B represents the current grasping point position of the robotic arm B.
[0030] Further, the first update module is further configured to:
[0031] The updated grasping point position d' of the robotic arm B B = arg max(V A )
[0032] Further, the position judgment module is configured to:
[0033] Calculate the current value where d' A represents the updated grasping point position of the robotic arm A, and d' B represents the updated grasping point position of the robotic arm B;
[0034] Judge whether the updated grasping point positions of the robotic arm A and the robotic arm B are the optimal positions by comparing the current value P with the set value δ.
[0035] An electronic device includes a memory and a processor. The memory stores a computer program. The processor, when executing the computer program, implements the contact point selection method combining the graspability of the dual robotic arms described in any one of the above.
[0036] A computer-readable storage medium stores computer program instructions thereon. The computer program instructions, when executed, implement the contact point selection method combining the graspability of the dual robotic arms described in any one of the above.
[0037] Compared with the prior art, the present invention has at least the following advantages.
[0038] 1. Since the end of the robotic arm operates the valve in a grasping state, the "operability" mentioned in the background art is refined into "graspability" in this part, that is, the feasibility of the end of the robotic arm reaching the specified pose and performing the grasping action, so as to give full play to the advantages of the two arms and turn the valve more labor - saving.
[0039] 2. The multi - agent iterative method used in the present invention can give full play to the collaborative effect of the two arms. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of a two - degree - of - freedom robotic arm.
[0041] Figure 2 It is a flowchart of the method of the present invention.
[0042] Figure 3 It is a schematic diagram of coordinate system setting.
[0043] Figure 4 It is a schematic diagram of the selection result of the grasping points of the planar valve image.
[0044] Figure 5 It is the evaluation results of the value of two groups of grasping points. SPECIFIC IMPLEMENTATION METHOD
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only specific embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0046] In the present invention, first, considering the width of the gripper at the end of the robotic arm, the valve operation area is discretized into grids of appropriate size. Second, as Figure 3As shown, for the same grid, it can be regarded as determining the three-dimensional spatial position coordinates of the grasping points, and there are also distinctions in different operating postures. Finally, the values of each grasping posture in each grid are calculated through the multi-agent value iteration method, and the optimal grasping points at the ends of the two arms are selected. Among them, the value iteration algorithm is a single-agent reinforcement learning algorithm. The agent continuously adjusts its own strategy through the value function and finally obtains the optimal strategy to guide its own action selection. For a two-arm system, since in the Cartesian space, the graspability evaluations of each arm for the same grasping point are different. Therefore, the two-arm system can be regarded as a multi-agent system. In a multi-agent system, each agent can learn and optimize its own strategy by interacting with the environment to obtain a reward value, so as to obtain the optimal strategy in different environments. Through continuous iteration, each agent can complete cooperative or competitive tasks. Therefore, in this paper, the multi-agent value iteration method is used between the two arms to determine the optimal valve grasping points for each arm.
[0047] As described above, the key of the present invention lies in designing a multi-agent value iteration model based on the mechanism of the human body and task orientation. Considering the reachability of the robot's two arms and the value function related to the task, the following gives the specific model-related functions and parameter setting methods:
[0048] (1) Given the position d of Arm A in the two arms A , then the value function of Arm B can be expressed as:
[0049] V B (d B ) = αM(d B ) - βcos(|d A - d B |)
[0050] M(p i,j,k ) = C i,j,k / N
[0051] Among them, M(d B ) represents the graspability of position d B for Arm B, that is, the feasibility of the robot arm B reaching the specified position of the valve and using the gripper to hold the valve. It is defined that among N sampling points, there are C i,j,k points located in the sampling points to evaluate the graspability. In one example, the operability weight α and the relative position weight β of the two arms can be set to 0.5 respectively.
[0052] (2) Given the valve rotation direction, the final value function V * of each position can be obtained by the following formula:
[0053] V * (d B) = M(d′ B ) + γV(d′ B )
[0054] where d′ B represents the position according to the valve rotation direction. The previous position of position d B can be expressed as d′ B = d B - Δd. When considering the current valve grasping point, this formula introduces the representation of the operability of the next movement position. Among them, the discount factor γ is used to control the magnitude of the influence of the operability of the next pose on the value of the current pose. In one example, γ can be selected as 0.4.
[0055] (3) According to the above two value function representations, when the positions of both arms are undetermined, a multi-agent iterative method can be used to determine the value function, as shown in the following algorithm. Figure 4 It is a schematic diagram of the selection result of the grasping point of the planar valve image obtained through this algorithm.
[0056]
[0057]
[0058] Specifically, the above algorithm can be described as follows:
[0059] Step 1: Combine the graspability of robotic arm B in the grid and the current grasping point position of robotic arm A to update the value function V B of robotic arm B and the final value function and based on the value function V B , calculate the updated grasping point position of robotic arm B; among them, the value function represents the value function considering spatial reachability and the value of the current state related to the task, and the final value function represents the value function of the expected value of the next action combined with the pose state of the previous step.
[0060] This Step 1 includes the following sub-steps.
[0061] Step 1.1: Calculate the graspability M(d B ) of robotic arm B according to the number of sampling points in the grid where the end of robotic arm B is located;
[0062] Step 1.2: Update the value function V B (d B ) = αM(d B ) - pcos(|d A - d B |); where α represents the operability weight, β represents the relative position weight of the two arms, d A represents the current grasping point position of robotic arm A, dB Indicates the current grasping point position of robotic arm B.
[0063] Step 1.3: Given the valve rotation direction, the final value function V of robotic arm B * (d B ) = M(d′ B ) + γV(d′ B ).
[0064] Step 1.4: Based on the value function V B , calculate the updated grasping point position d′ of robotic arm B B = arg max(V A );
[0065] Step 2: Combine the graspability of robotic arm A in the grid and the updated grasping point position of robotic arm B to update the value function V of robotic arm A A with the final value function and based on the value function V A , calculate the updated grasping point position of robotic arm A.
[0066] Step 3: Based on the value function V A , the final value function , the value function V B and the final value function , determine whether the updated grasping point positions of robotic arm A and robotic arm B are the optimal positions.
[0067] This step 3 includes the following sub-steps.
[0068] Step 3.1: Calculate the current value where d′ A represents the updated grasping point position of robotic arm A, and d′ B represents the updated grasping point position of robotic arm B;
[0069] Step 3.2: By comparing the current value P with the set value δ, determine whether the updated grasping point positions of robotic arm A and robotic arm B are the optimal positions.
[0070] Step 4: In the case where the updated grasping point positions of robotic arm A and robotic arm B are not the optimal positions, based on the updated grasping point positions of robotic arm A and robotic arm B, return to updating the value function V of robotic arm B according to the graspability of robotic arm B in the grid and the current grasping point position of robotic arm A B with the final value function Moreover, when the updated positions of the robotic arm A and the robotic arm B are the optimal positions, the grasping pose of the robotic arm A and the robotic arm B is calculated based on the updated grasping point positions of the robotic arm A and the robotic arm B.
[0071] Experimental results:
[0072] Two baseline methods are adopted in this paper as control experiments, which are respectively: 1. Random selection method: Randomly sample 10 groups of grasping points within the intersection of the feasible space of the two arms and the valve area, and select the group with the highest evaluation value as the grasping points of the two arms; 2. Analysis based on reachability map: Based on the method mentioned in the "Background Art" above, the valve is gridded at intervals of 1 degree, and then the graspability of each grid relative to each arm is calculated by sampling, and the grasping points of each arm are selected accordingly.
[0073] (1) The experimental results of the simulation platform are shown in Table 1:
[0074]
[0075] Table 1
[0076] (2) The experimental results of the physical platform are shown in Table 2:
[0077]
[0078] Table 2
[0079] It can be concluded from the above Table 1 and Table 2 that the multi-agent iterative method proposed in the present invention has better effects both in the simulation environment and the physical environment. Moreover, among the two groups of experiments selected therefrom, the value of the grasping points generated by the multi-agent value iteration method is represented by a heat map as Figure 5 shown. It can be seen from Figure 5 that: (1) For a single arm, the value of the grasping point is affected by the value of the grasping point at the next position; (2) For the two arms, the high-value grasping points of the two arms are usually staggered, that is, as described above, the larger the included angle between the two arms, the more the synergistic effect of the two arms can be exerted.
[0080] After considering the specification and practicing the present disclosure, those skilled in the art will readily think of other embodiments of the present disclosure. This application aims to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A contact point selection method combining the graspability of dual robotic arms, characterized in that The method includes: Meshing the target operation area according to the width of the end grippers of the dual robotic arms; the dual robotic arms include: robotic arm A and robotic arm B; Update the value function V of robotic arm B by combining the graspability of robotic arm B in the grid and the current grasp point position of robotic arm A B with the final value function and based on the value function V B , calculate the updated grasp point position of robotic arm B; wherein, the value function represents the value function of the current state considering spatial reachability and task relevance, and the final value function represents the value function of the expected value of the next action in combination with the pose state of the previous step; Update the value function V of manipulator A by combining the graspability of manipulator A in the grid and the updated grasp point position of manipulator B A with the final value function and calculate the updated grasp point position of manipulator A based on the value function V A ,; Based on the value function V A and the final value function the value function V B and the final value function Determine whether the updated grasping point positions of robotic arm A and robotic arm B are the optimal positions; In the case where the updated grasping point positions of robotic arm A and robotic arm B are not the optimal positions, based on the updated grasping point positions of robotic arm A and robotic arm B, re - execute the combination of the graspability of robotic arm B and the current grasping point position of robotic arm A in the combined grid to update the value function V of robotic arm B B with the final value function When the updated positions of robotic arm A and robotic arm B are the optimal positions, calculate the grasping poses of robotic arm A and robotic arm B based on the updated grasping point positions of robotic arm A and robotic arm B.
2. The method according to claim 1, characterized in that, Updating the value function V of the robotic arm B by combining the graspability of the robotic arm B in the grid and the current grasping point position of the robotic arm A B , including: Calculate the graspability M(d B ) of the robotic arm B based on the number of sampling points in the grid where the end of the robotic arm B is located; Update the value function V of the robotic arm B B (d B ) = αM(d B ) - βcos(|d A - d B |); where α represents the manipulability weight, β represents the relative position weight of the two arms, d A represents the current grasping point position of the robotic arm A, and d B represents the current grasping point position of the robotic arm B.
3. The method according to claim 1, characterized in that, Based on the value function V B , calculating the position of the updated grasping point of the robotic arm B, including: The position d′ of the updated grasping point of the robotic arm B B = arg max(V A ).
4. The method according to claim 1, wherein Based on the value function V A , the final value function The value function V B and the final value function Determining whether the updated grasping point positions of robotic arm A and robotic arm B are optimal positions, includes: Calculate the current value where d' A represents the position of the updated grasping point of robotic arm A, and d' B represents the position of the updated grasping point of robotic arm B; Judge whether the updated grasping point positions of robotic arm A and robotic arm B are the optimal positions by comparing the current value P with the set value δ.
5. A contact point selection device combined with the graspability of dual robotic arms, characterized in that, The device includes: A discretization module for meshing the target operation area according to the width of the end grippers of the dual robotic arms; the dual robotic arms include: robotic arm A and robotic arm B; A first update module, which is used to update the value function V of the robotic arm B by combining the graspability of the robotic arm B in the grid and the current grasping point position of the robotic arm A B and the final value function and based on the value function V B , calculate the updated grasping point position of the robotic arm B; wherein, the value function represents a value function considering spatial reachability and the value of the current state related to the task, and the final value function represents a value function of the expected value for the next action by combining the pose state of the previous step A second update module, configured to update the value function V of robotic arm A by combining the graspability of robotic arm A in the grid and the updated grasp point position of robotic arm B A with the final value function and, based on the value function V A , calculate the updated grasp point position of robotic arm A; A position judgment module, configured to be based on the value function V A , the final value function the value function V B and the final value function judge whether the updated grasping point positions of the robotic arm A and the robotic arm B are optimal positions; in the case where the updated grasping point positions of the robotic arm A and the robotic arm B are not optimal positions, based on the updated grasping point positions of the robotic arm A and the robotic arm B, re - execute combining the graspability of the robotic arm B in the grid and the current grasping point position of the robotic arm A to update the value function V of the robotic arm B B and the final value function In the case where the updated positions of the robotic arm A and the robotic arm B are optimal positions, calculate the grasping pose of the robotic arm A and the robotic arm B according to the updated grasping point positions of the robotic arm A and the robotic arm B.
6. The device according to claim 5, characterized in that, The first update module is used for: Calculate the graspability M(d B ) of the robotic arm B according to the number of sampling points in the grid where the end of the robotic arm B is located; Update the value function V of the robotic arm B B (d B ) = αM(d B ) - βcos(|d A - d B |); where α represents the manipulability weight, β represents the relative position weight of the two arms, d A represents the current grasping point position of the robotic arm A, and d B represents the current grasping point position of the robotic arm B.
7. The device according to claim 5, wherein, The first update module is also used for: The position d′ of the updated grasping point of the robotic arm B B = arg max(V A ).
8. The device according to claim 5, characterized in that, The position judgment module is used for: Calculate the current value where d' A represents the position of the updated grasping point of robotic arm A, and d' B represents the position of the updated grasping point of robotic arm B; Judge whether the updated grasping point positions of robotic arm A and robotic arm B are the optimal positions by comparing the current value P with the set value δ.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the contact point selection method for the graspability of the dual robotic arms according to any one of claims 1 to 4.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed, they implement the contact point selection method for the graspability of the dual robotic arms according to any one of claims 1 to 4.