Motion decision-making method, device and equipment of mechanical arm, medium and product

By constructing a motion decision model based on reinforcement learning, and using the environment perception equipment of the robotic arm to collect point cloud data, the problem of inefficient motion decision-making of the robotic arm is solved, and efficient and reliable action decisions are achieved.

CN120287307APending Publication Date: 2025-07-11RUANTONG TIANSHU INTELLIGENT (NANJING) TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510649604.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

现有技术中机械臂的动作决策方法存在效率低下和计算复杂度高的问题,尤其在逆运动学求解时容易陷入多解或局部最优解。

Method used

By constructing a motion decision model based on reinforcement learning, the environment perception device of the robot arm collects point cloud data of the target object, and combines the motion state information of the robot arm to perform time-step action decision optimization.

Benefits of technology

It improves the efficiency of the action decision-making and the reliability of the decision-making of the robotic arm, ensures the accuracy and stability of the action, and reduces the consumption of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120287307A_ABST
    Figure CN120287307A_ABST
Patent Text Reader

Abstract

The invention discloses an action decision-making method, device and equipment of a mechanical arm, a medium and a product. The method comprises the steps that motion state information of the mechanical arm at the current time step is obtained, and environment information of the current time step is obtained; the environment information is point cloud data which is collected by environment sensing equipment deployed on the mechanical arm and contains a target object; according to the motion state information of the current time step and the environment information of the current time step, determining an action decision of a next time step based on a pre-constructed action decision model; the action decision model is obtained by training an action decision in the process of executing the historical grabbing task by the mechanical arm based on reinforcement learning. According to the technical scheme, the problem that the motion decision-making efficiency of the mechanical arm is low is solved, the motion decision-making efficiency can be effectively improved by constructing the motion decision-making model, and the reliability of motion decision-making is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robotic arm control, and particularly to a method, device, equipment, medium and product for action decision-making of a robotic arm. Background Art

[0002] With the rapid development of industrial automation, robotic arms have been widely used in fields such as production manufacturing, logistics warehousing, and equipment inspection. The action decision-making method of a robotic arm is the core of completing various tasks.

[0003] Currently, the action decision-making method of a robotic arm in the prior art usually sets a camera on the robotic arm to collect surrounding environment information following the movement of the robotic arm. After hand-eye calibration, the end pose of the robotic arm is solved based on inverse kinematics. However, inverse kinematics has certain bottlenecks, such as having multiple solutions or falling into local optimal solutions, and the calculation process is cumbersome and complex, consuming a large amount of computing resources and storage resources. Summary of the Invention

[0004] The present invention provides a method, device, equipment, medium and product for action decision-making of a robotic arm to solve the problem of low efficiency of robotic arm motion decision-making. By constructing an action decision-making model, the action decision-making efficiency can be effectively improved, and the reliability of motion decision-making can be ensured.

[0005] According to one aspect of the present invention, a method for action decision-making of a robotic arm is provided, and the method includes:

[0006] Obtain the motion state information of the robotic arm at the current time step, and obtain the environmental information at the current time step; the environmental information is point cloud data containing a target object collected by an environmental perception device deployed on the robotic arm; the target object is the grasping object of the grasping task being executed by the robotic arm;

[0007] Based on the motion state information at the current time step and the environmental information at the current time step, determine the action decision for the next time step based on a pre-constructed action decision-making model; the action decision-making model is obtained by training the action decisions in the process of the robotic arm executing historical grasping tasks based on reinforcement learning.

[0008] According to another aspect of the present invention, an action decision-making device for a robotic arm is provided, and the device includes:

[0009] An information acquisition module, configured to obtain the motion state information of the robotic arm at the current time step, and obtain the environmental information at the current time step; the environmental information is point cloud data containing a target object collected by an environmental perception device deployed on the robotic arm; the target object is the grasping object of the grasping task being executed by the robotic arm;

[0010] An action decision-making module, configured to determine the action decision for the next time step based on the motion state information of the current time step and the environmental information of the current time step, based on a pre-constructed action decision model; the action decision model is obtained by training the action decisions during the historical grasping tasks performed by the robotic arm based on reinforcement learning.

[0011] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:

[0012] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the action decision method of the robotic arm according to any embodiment of the present invention.

[0013] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the action decision method of the robotic arm according to any embodiment of the present invention when executed.

[0014] According to another aspect of the present invention, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the action decision method of the robotic arm according to any embodiment of the present invention.

[0015] The technical solution of the embodiment of the present invention obtains the motion state information of the robotic arm at the current time step and obtains the environmental information of the current time step; the environmental information is point cloud data containing the target object collected by an environmental perception device deployed on the robotic arm; the target object is the grasping object of the grasping task being performed by the robotic arm; according to the motion state information of the current time step and the environmental information of the current time step, based on a pre-constructed action decision model, determine the action decision for the next time step; the action decision model is obtained by training the action decisions during the historical grasping tasks performed by the robotic arm based on reinforcement learning. This technical solution solves the problem of low efficiency of the motion decision-making of the robotic arm. By constructing an action decision model, the action decision-making efficiency can be effectively improved, and the reliability of the motion decision-making can be ensured.

[0016] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0018] Figure 1 is a flowchart of a method for action decision-making of a robotic arm provided according to Embodiment 1 of the present invention;

[0019] Figure 2 is a flowchart of a method for action decision-making of a robotic arm provided according to Embodiment 2 of the present invention;

[0020] Figure 3 is a schematic structural diagram of a device for action decision-making of a robotic arm provided according to Embodiment 3 of the present invention;

[0021] Figure 4 is a schematic structural diagram of an electronic device for implementing the method for action decision-making of the robotic arm in the embodiments of the present invention. Detailed Embodiments

[0022] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. The acquisition, storage, use, processing, etc. of data in the technical solutions of this application all comply with the relevant regulations of national laws and regulations.

[0024] Embodiment 1

[0025] Figure 1The following is a flowchart of a motion decision-making method for a robotic arm provided in the first embodiment of the present invention. This embodiment is applicable to the robotic arm control scenario, especially the situation where the robotic arm grasps an object. This method can be executed by the motion decision-making device of the robotic arm. The device can be implemented in the form of hardware and / or software and can be configured in an electronic device. As Figure 1 shown, the method includes:

[0026] S110. Obtain the motion state information of the robotic arm at the current time step, and obtain the environmental information at the current time step; the environmental information is point cloud data containing the target object collected by the environmental perception device deployed on the robotic arm; the target object is the grasping object of the grasping task being executed by the robotic arm.

[0027] This solution can be executed by the robotic arm control platform. Environmental perception devices such as cameras and radars can be deployed on each joint and the end effector of the robotic arm to sense the working environment information of the robotic arm. One or more positions on the robotic arm can be deployed with environmental perception devices, and one or more environmental perception devices can be deployed at each position. The deployment position and the number of deployed environmental perception devices can be determined according to the grasping task of the robotic arm.

[0028] The robotic arm can send the motion state information of the current time step to the control platform at a preset frequency. The motion state information can include information such as the angles of each joint of the robotic arm, the speeds of each joint, the movement ranges of each joint, and the position of the end effector. The environmental perception device deployed on the robotic arm can also send the environmental information of the current time step to the control platform at a preset frequency. The environmental information can be point cloud data containing the target object. The target object is the grasping object of the grasping task being executed by the robotic arm, and the point cloud data can be used to represent the distance information between the environmental perception device and the target object.

[0029] S120. Based on the motion state information of the current time step and the environmental information of the current time step, and based on the pre-constructed motion decision-making model, determine the motion decision for the next time step; the motion decision-making model is obtained by training the motion decisions in the process of the robotic arm executing historical grasping tasks based on reinforcement learning.

[0030] After obtaining the motion state information and environmental information of the current time step, the control platform can fuse the operation state information and environmental information of the current time step to obtain the input data of the motion decision-making model, which jointly provides a basis for the motion decision for the next time step. After inputting the motion state information and environmental information of the current time step into the motion decision-making model, the motion decision-making model can output the motion decision for the next time step. The motion decision can include information such as the angles that each joint of the robotic arm needs to rotate, the speeds that each joint of the robotic arm needs to rotate, and the position that the end effector needs to reach.

[0031] The action decision-making model can be trained based on reinforcement learning for the action decisions during the execution of historical grasping tasks by the robotic arm. It can be understood that reinforcement learning can, through interaction with the environment, train the intelligent agent to take the best actions in different motion states of the robotic arm to achieve the purpose of grasping the target object. The action decision-making model can be used to make action decisions step by step in time.

[0032] In the technical solution of the embodiment of the present invention, by obtaining the motion state information of the robotic arm at the current time step and obtaining the environmental information at the current time step; the environmental information is point cloud data containing the target object collected by an environmental perception device deployed on the robotic arm; the target object is the grasping object of the grasping task being executed by the robotic arm; based on the motion state information at the current time step and the environmental information at the current time step, and based on a pre-constructed action decision-making model, determine the action decision for the next time step; the action decision-making model is trained based on reinforcement learning for the action decisions during the execution of historical grasping tasks by the robotic arm. This technical solution solves the problem of low efficiency of the motion decision-making of the robotic arm. By constructing an action decision-making model, the action decision-making efficiency can be effectively improved, and the reliability of the motion decision-making can be ensured.

[0033] Embodiment Two

[0034] Figure 2 It is a flowchart of a method for making action decisions of a robotic arm provided in Embodiment Two of the present invention. In this embodiment, the construction process of the action decision-making model is refined based on the above embodiment. As Figure 2 shown, the method includes:

[0035] S201. Obtain the motion state information of the robotic arm at the first time step and the environmental information at the first time step during the execution of the target grasping task.

[0036] In this solution, the reinforcement learning model can be trained through large-scale target grasping tasks to obtain an action decision-making model that meets the requirements of action decision-making accuracy. It can be understood that a grasping task can be completed through action decisions in multiple time steps. The control platform can, in chronological order, obtain the motion state information of the first time step in the target grasping task and use it as the motion state information of the first time step, and obtain the environmental information of the first time step as the environmental information of the first time step.

[0037] S202. Generate the input data for the first time step according to the motion state information at the first time step and the environmental information at the first time step.

[0038] The control platform can merge the motion state information and environmental information of the first time step to obtain the input data for the first time step.

[0039] S203. Input the input data of the first time step into the pre-constructed reinforcement learning model to obtain the action decision of the second time step.

[0040] The control platform can input the input data of the first time step into the reinforcement learning model, and the reinforcement learning model can output the action decision of the second time step. It can be understood that the second time step is the next time step of the first time step.

[0041] S204. When the robotic arm moves to the position according to the action decision of the second time step, obtain the environmental information of the second time step.

[0042] The robotic arm can move according to the action decision of the second time step output by the reinforcement learning model. When the movement is completed, the environmental information of the second time step is obtained through the environmental perception device deployed on the robotic arm.

[0043] S205. Determine the decision reward according to the environmental information of the second time step, and perform one training on the reinforcement learning model according to the decision reward.

[0044] The control platform can evaluate the current action decision based on the environmental information of the second time step, determine the decision reward of the current action decision, and feedback the decision reward of the current action decision to the reinforcement learning model to train the reinforcement learning model. Specifically, the reinforcement learning model can be a deep neural network model, and the control platform can adjust the weight parameters in the reinforcement learning model using the decision reward of the current action decision to achieve the training purpose.

[0045] In this solution, the decision reward is determined based on distance reward, progress reward, visual reward, motion smoothness penalty, jitter suppression, and safety constraints;

[0046] The distance reward is determined based on the distance between the end effector of the robotic arm and the target object;

[0047] The progress reward is determined based on the change in the distance between the end effector of the robotic arm and the target object between adjacent time steps;

[0048] The visual reward is determined based on the position of the target object in the point cloud data;

[0049] The motion smoothness penalty is determined based on the motions of the joints of the robotic arm;

[0050] The jitter suppression is determined based on the jitter amplitude of the joints of the robotic arm;

[0051] The safety constraint is determined based on whether the end effector of the robotic arm enters the abnormal space, and the abnormal space is a singular point or other space outside the workspace.

[0052] In a feasible solution, the decision reward can be expressed as:

[0053] R = r d + r p + r v - r a - r o - r w + r s ;

[0054] where r d represents the distance reward, r p represents the progress reward, r v represents the visual reward, r a represents the action smoothness penalty, r o represents the jitter suppression, r w represents the safety constraint, r s represents the success reward;

[0055] The distance reward can be expressed as:

[0056] r d = -‖d t ‖2; t represents the time step index, d t represents the distance between the environmental perception device and the target object at time step t.

[0057] The progress reward can be expressed as:

[0058] r p = α(‖d t-1 ‖2 - ‖d t ‖2); α represents a constant coefficient, d t-1 represents the distance between the environmental perception device and the target object at time step t - 1;

[0059] The visual reward can be expressed as:

[0060] β1, β2, and β3 respectively represent different reward values and are constants.

[0061] The action smoothness penalty can be expressed as:

[0062] γ represents a constant coefficient, i represents the joint index in the robotic arm, a i represents the angular deviation value of joint i at the current time step relative to the previous time step, and n represents the number of joints in the robotic arm.

[0063] r o represents the jitter suppression, and δ1, δ2 respectively represent different suppression values and are constants.

[0064]

[0065] ε1 and ε2 respectively represent different constraint values and are constants.

[0066] r s = μI(‖d t ‖2 < σ); μ represents a constant coefficient, σ represents an error threshold, and I(·) represents an indicator function, that is, when ‖d t ‖2 < σ, the function value is 1, otherwise the function value is 0.

[0067] The reward and punishment mechanism of this scheme comprehensively considers various factors and can effectively guide the intelligent agent to quickly learn the grasping strategy. Introducing the basic distance reward to encourage the robotic arm to approach the target can solve the problem of blind exploration; the progress reward encourages the advancement of the task and can avoid falling into local optima; the visual reward can make full use of the point cloud data to improve the recognition and positioning ability of the robotic arm in complex environments; the action smoothness penalty and the jitter penalty can ensure smooth actions and improve the operation stability; the safety constraint can ensure that the robotic arm operates within a safe workspace; the success reward can strengthen the correct behavior of completing the task and improve the task completion rate.

[0068] S206. Determine whether the evaluation period has arrived.

[0069] It is easy to understand that after obtaining the decision reward for this decision, the control platform can judge whether the evaluation period of the intelligent agent has arrived according to the total number of time steps of the executed target grasping task. Among them, the evaluation period can include a preset number of time steps. For example, if the evaluation period includes 2000 time steps, 100 target grasping tasks have been executed, and a total of 1200 time steps have been executed for 100 target grasping tasks, then the current time step has not reached the evaluation period. If the evaluation period has arrived, execute S210; if the evaluation period has not arrived, execute S207.

[0070] S207. Determine whether the termination condition of the target grasping task is satisfied.

[0071] The control platform can judge whether the termination condition of the target grasping task is satisfied. If the termination condition of the target grasping task is not satisfied, execute S208; if the termination condition of the target grasping task is satisfied, execute S209.

[0072] Optionally, the satisfaction of the termination condition of the target grasping task includes that the target grasping task is completed or the cumulative number of time steps for executing the target grasping task reaches a preset quantity threshold.

[0073] The termination condition of the target grasping task can be that the target grasping task is completed, that is, the end effector of the robotic arm grasps the target object. To avoid the influence of abnormal grasping tasks, such as grasping tasks that cannot be completed, the control platform can preset a quantity threshold to limit the maximum number of time steps for the robotic arm to complete the target grasping task. Therefore, the termination condition of the target grasping task can also be that the cumulative number of time steps for executing the target grasping task reaches the preset quantity threshold.

[0074] S208. Update the second time step to the first time step, and update the input data of the first time step according to the environmental information of the first time step and the motion state information of the first time step.

[0075] If the target grasping task is not terminated, the control platform can update the second time step to the first time step, update the input data of the first time step according to the environmental information of the first time step and the motion state information of the first time step, and return to execute S203.

[0076] S209. Update the target grasping task.

[0077] If the termination condition of the target grasping task is satisfied, the control platform can continue to set a new target grasping task and return to execute S201.

[0078] S210. Obtain the decision rewards of each time step within the evaluation period, and determine the action decision model according to the decision rewards of each time step within the evaluation period.

[0079] If the evaluation period arrives, the control platform can obtain the decision rewards of all time steps within the evaluation period and evaluate the reinforcement learning model obtained at the current time step. Specifically, the control platform can count the decision rewards of each time step within the evaluation period, evaluate the usability of the reinforcement learning model obtained at the current time step according to the statistical results, and then determine whether the reinforcement learning model obtained at the current time step can be used as the action decision model according to the evaluation results.

[0080] In a feasible solution, the determining the action decision model according to the decision rewards of each time step within the evaluation period includes:

[0081] Determine the performance evaluation index of the reinforcement learning model according to the decision rewards of each time step within the evaluation period; the performance evaluation index includes at least one of cumulative reward, average reward, reward stability, and convergence speed;

[0082] Determine the action decision model according to the performance evaluation index of the reinforcement learning model.

[0083] In this solution, the control platform can calculate the performance evaluation metrics of the reinforcement learning model based on the decision rewards at each time step within the evaluation period, such as cumulative reward, average reward, reward stability, and convergence speed, and evaluate the usability of the reinforcement learning model obtained at the current time step according to the performance evaluation metrics of the reinforcement learning model. If the performance evaluation metrics meet the preset usability conditions, the reinforcement learning model obtained at the current time step is used as the action decision model.

[0084] For the case where the performance evaluation metrics do not meet the preset usability conditions, the control platform can execute S207 to continue training the reinforcement learning model until the reinforcement learning model obtained at the current time step reaches the usability conditions. The control platform can also discard the reinforcement learning model obtained at the current time step, reset the reinforcement learning model, and retrain a new reinforcement learning model from scratch to obtain an action decision model that meets the preset usability conditions.

[0085] Among them, the cumulative reward can be expressed as: The average reward can be expressed as: The reward stability can be expressed as: The convergence speed can be expressed as: Among them, j represents the time step index within an evaluation period, m represents the number of time steps within an evaluation period, and R j represents the decision reward value at time step j.

[0086] This solution ensures the action decision accuracy of the action decision model by calculating the multi-dimensional performance evaluation metrics of the reinforcement learning model.

[0087] S211. Obtain the motion state information of the robotic arm at the current time step, and obtain the environmental information at the current time step; the environmental information is the point cloud data containing the target object collected by the environmental perception device deployed on the robotic arm; the target object is the grasping object of the grasping task being executed by the robotic arm.

[0088] S212. Based on the motion state information at the current time step and the environmental information at the current time step, and based on the pre-constructed action decision model, determine the action decision for the next time step.

[0089] In a feasible solution, the determining the action decision for the next time step based on the motion state information at the current time step and the environmental information at the current time step, and based on the pre-constructed action decision model includes:

[0090] Generate the input data for the current time step according to the motion state information at the current time step and the environmental information at the current time step; the input data includes the joint angles, joint velocities, end effector position, movable space, singular points of the robotic arm, and the point cloud of the target object.

[0091] Input the input data of the current time step into the motion decision-making model to obtain the motion decision for the next time step; the motion decision includes the joint angles, joint speeds of the robotic arm, and the position of the end effector.

[0092] In this solution, the input data includes the joint angles, joint speeds, end effector position, movable space, singularity, and target object point cloud of the robotic arm. Based on the input data of the current time step, the motion decision-making model can obtain the motion decision for the next time step, which includes the joint angles, joint speeds of the robotic arm, and the position of the end effector.

[0093] This solution uses the target object point cloud as part of the input data of the motion decision-making model, providing rich environmental information for the intelligent agent and helping the intelligent agent make more accurate motion decisions.

[0094] Embodiment III

[0095] Figure 3 It is a schematic structural diagram of a motion decision-making device for a robotic arm provided in Embodiment III of the present invention. As Figure 3 shown, the device includes:

[0096] An information acquisition module 310, configured to acquire the motion state information of the robotic arm at the current time step and acquire the environmental information at the current time step; the environmental information is point cloud data containing the target object collected by the environmental perception device deployed on the robotic arm; the target object is the grasping object of the grasping task being executed by the robotic arm.

[0097] A motion decision-making module 320, configured to determine the motion decision for the next time step based on the motion state information at the current time step and the environmental information at the current time step, based on a pre-constructed motion decision-making model; the motion decision-making model is obtained by training the motion decisions in the process of the robotic arm executing historical grasping tasks based on reinforcement learning.

[0098] In this solution, the device further includes a decision model construction module, configured to:

[0099] Take one of the historical grasping tasks as the target grasping task, and acquire the motion state information of the robotic arm at the first time step and the environmental information at the first time step in the process of executing the target grasping task;

[0100] Generate the input data of the first time step according to the motion state information of the first time step and the environmental information of the first time step;

[0101] Input the input data of the first time step into the pre-constructed reinforcement learning model to obtain the motion decision for the second time step;

[0102] When the robotic arm moves to the designated position according to the action decision at the second time step, obtain the environmental information at the second time step;

[0103] Determine the decision reward based on the environmental information at the second time step, and perform one training on the reinforcement learning model according to the decision reward;

[0104] Update the second time step to the first time step, update the input data at the first time step according to the environmental information and the motion state information at the first time step, and return to execute inputting the input data at the first time step into the pre-constructed reinforcement learning model until the termination condition of the target grasping task is met;

[0105] Update the target grasping task, and return to execute obtaining the motion state information and the environmental information at the first time step during the execution of the target grasping task by the robotic arm until the evaluation period arrives, obtain the decision rewards at each time step within the evaluation period, and determine the action decision model according to the decision rewards at each time step within the evaluation period.

[0106] In a feasible solution, the decision reward is determined based on distance reward, progress reward, visual reward, action smoothness penalty, jitter suppression, and safety constraints;

[0107] The distance reward is determined based on the distance between the end effector of the robotic arm and the target object;

[0108] The progress reward is determined based on the change in the distance between the end effector of the robotic arm and the target object between adjacent time steps;

[0109] The visual reward is determined based on the position of the target object in the point cloud data;

[0110] The action smoothness penalty is determined based on the actions of the joints of the robotic arm;

[0111] The jitter suppression is determined based on the jitter amplitude of the joints of the robotic arm;

[0112] The safety constraint is determined based on whether the end effector of the robotic arm enters an abnormal space, and the abnormal space is a singular point or other space outside the workspace.

[0113] Based on the above solution, the decision model construction module is specifically used for:

[0114] Determine the performance evaluation index of the agent according to the decision rewards at each time step within the evaluation period; the performance evaluation index includes at least one of cumulative reward, average reward, reward stability, and convergence speed;

[0115] Determine the action decision model according to the performance evaluation index of the reinforcement learning model.

[0116] Optionally, the satisfaction of the target grasping task termination condition includes that the target grasping task is completed or the cumulative number of time steps for executing the target grasping task reaches a preset quantity threshold.

[0117] In a preferred solution, the action decision module 320 is specifically configured to:

[0118] Generate input data for the current time step according to the motion state information of the current time step and the environmental information of the current time step; the input data includes the joint angles, joint velocities, end effector positions, movable spaces, singular points, and target object point clouds of the robotic arm;

[0119] Input the input data of the current time step into the action decision model to obtain the action decision for the next time step; the action decision includes the joint angles, joint velocities, and end effector positions of the robotic arm.

[0120] The action decision device of the robotic arm provided by the embodiments of the present invention can execute the action decision method of the robotic arm provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0121] Embodiment 4

[0122] Figure 4 FIG. shows a schematic structural diagram of an electronic device 410 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0123] As Figure 4As shown, the electronic device 410 includes at least one processor 411, and a memory communicatively connected to the at least one processor 411, such as a read-only memory (ROM) 412, a random access memory (RAM) 413, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 411 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 412 or the computer program loaded from the storage unit 418 into the random access memory (RAM) 413. In the RAM 413, various programs and data required for the operation of the electronic device 410 can also be stored. The processor 411, the ROM 412, and the RAM 413 are connected to each other via a bus 414. The input / output (I / O) interface 415 is also connected to the bus 414.

[0124] Multiple components in the electronic device 410 are connected to the I / O interface 415, including: an input unit 416, such as a keyboard, a mouse, etc.; an output unit 417, such as various types of displays, speakers, etc.; a storage unit 418, such as a magnetic disk, an optical disc, etc.; and a communication unit 419, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 419 allows the electronic device 410 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0125] The processor 411 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 411 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 411 executes the various methods and processes described above, such as the action decision method of the robotic arm.

[0126] In some embodiments, the action decision method of the robotic arm can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 418. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 410 via the ROM 412 and / or the communication unit 419. When the computer program is loaded into the RAM 413 and executed by the processor 411, one or more steps of the action decision method of the robotic arm described above can be executed. Alternatively, in other embodiments, the processor 411 can be configured to execute the action decision method of the robotic arm in any other appropriate way (for example, by means of firmware).

[0127] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0128] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable robotic arm motion decision-making devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0129] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).

[0131] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0132] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0133] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0134] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A motion decision-making method for a robotic arm, characterized in that, The method includes: Obtaining the motion state information of the robotic arm at the current time step, and obtaining the environmental information at the current time step; the environmental information is point cloud data containing the target object collected by an environmental perception device deployed on the robotic arm; the target object is the grasping object of the grasping task being executed by the robotic arm. Based on the motion state information at the current time step and the environmental information at the current time step, and based on a pre-constructed action decision model, determining the action decision for the next time step; the action decision model is obtained by training the action decisions during the historical grasping tasks executed by the robotic arm based on reinforcement learning.

2. The method according to claim 1, wherein The construction process of the action decision model includes: Obtaining the motion state information of the robotic arm at the first time step and the environmental information at the first time step during the execution of the target grasping task. Generating the input data for the first time step according to the motion state information at the first time step and the environmental information at the first time step. Inputting the input data for the first time step into a pre-constructed reinforcement learning model to obtain the action decision for the second time step. When the robotic arm moves in place according to the action decision for the second time step, obtaining the environmental information at the second time step. Determining the decision reward according to the environmental information at the second time step, and performing one training on the reinforcement learning model according to the decision reward. Updating the second time step to the first time step, updating the input data for the first time step according to the environmental information at the first time step and the motion state information at the first time step, and returning to execute inputting the input data for the first time step into the pre-constructed reinforcement learning model until the termination condition of the target grasping task is satisfied. Updating the target grasping task, and returning to execute obtaining the motion state information of the robotic arm at the first time step and the environmental information at the first time step during the execution of the target grasping task until the evaluation period arrives, obtaining the decision rewards for each time step within the evaluation period, and determining the action decision model according to the decision rewards for each time step within the evaluation period.

3. The method according to claim 2, wherein The decision reward is determined based on distance reward, progress reward, visual reward, action smoothness penalty, jitter suppression, and safety constraints. The distance reward is determined based on the distance between the end effector of the robotic arm and the target object. The progress reward is determined based on the change in the distance between the end effector of the robotic arm and the target object between adjacent time steps. The visual reward is determined based on the position of the target object in the point cloud data. The action smoothness penalty is determined based on the actions of the joints of the robotic arm. The jitter suppression is determined based on the jitter amplitude of the joints of the robotic arm. The safety constraint is determined based on whether the end effector of the robotic arm enters an abnormal space, and the abnormal space is a singular point or other space outside the workspace.

4. The method according to claim 2, wherein The determining the action decision model according to the decision rewards for each time step within the evaluation period includes: Determining the performance evaluation index of the reinforcement learning model according to the decision rewards for each time step within the evaluation period; the performance evaluation index includes at least one of cumulative reward, average reward, reward stability, and convergence speed. Determining the action decision model according to the performance evaluation index of the reinforcement learning model.

5. The method according to claim 2, characterized in that, The satisfaction of the termination condition of the target grasping task includes that the target grasping task is completed or the cumulative number of time steps for executing the target grasping task reaches a preset threshold.

6. The method according to claim 1, wherein The determination of the action decision for the next time step based on the motion state information of the current time step and the environmental information of the current time step, based on a pre-constructed action decision model, includes: Generating input data for the current time step according to the motion state information of the current time step and the environmental information of the current time step; the input data includes the joint angles, joint velocities, end effector positions, movable spaces, singular points, and target object point clouds of the robotic arm. Inputting the input data of the current time step into the action decision model to obtain the action decision for the next time step; the action decision includes the joint angles, joint velocities, and end effector positions of the robotic arm.

7. An action decision-making device for a robotic arm, characterized in that, The device includes: An information acquisition module for acquiring the motion state information of the robotic arm at the current time step and acquiring the environmental information of the current time step; the environmental information is point cloud data containing the target object collected by an environmental perception device deployed on the robotic arm; the target object is the grasping object of the grasping task being executed by the robotic arm. An action decision module for determining the action decision for the next time step based on the motion state information of the current time step and the environmental information of the current time step, based on a pre-constructed action decision model; the action decision model is obtained by training the action decisions in the process of the robotic arm executing historical grasping tasks based on reinforcement learning.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the action decision method of the robotic arm according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the action decision method of the robotic arm according to any one of claims 1-6 when executed.

10. A computer program product, including a computer program that implements the action decision method of the robotic arm according to any one of claims 1-6 when executed by a processor.

Citation Information

Cited By

  • Mechanical arm dynamic path planning method and system based on simulation learning

    CN121267922A