Implementation Method of End-to-End Robot Reinforcement Learning Batting Strategy Based on Stage Rewards
Through the end-to-end robot reinforcement learning method of phased training goals and learning task rewards, the robustness and cost problems of robotic batting strategies in the existing technology in complex scenarios are solved, and efficient and low-latency table tennis tasks completion and landing control are achieved.
Patent Information
- Application Number
- CN202211639730.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-12-20
AI Technical Summary
The existing robotic batting strategy is difficult to achieve efficient and robust table tennis tasks in complex scenarios, and the coordination of multiple systems leads to communication delays and high costs, and special structural transformation increases the deployment difficulty.
The end-to-end robot reinforcement learning method based on stage reward is adopted, and a single end-to-end strategy system is built through phased training goals and learning task rewards, combining PD servo control and traditional mechanical structures to realize the robot's ball hitting strategy.
It reduces system delay and deployment costs, improves the deployability and robustness of the robot in the real environment, can complete complex table tennis tasks and control table tennis landing points, and adapt to a variety of scenarios.
Smart Images

Figure CN115946137B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot control, and particularly to an end-to-end robot hitting ball strategy implementation method, device and medium based on stage-setting rewards. Background Art
[0002] Currently, the robot hitting ball strategy based on reinforcement learning is mainly implemented by the following methods:
[0003] 1. Adopt a data-driven or model-based method to directly predict the table tennis movement trajectory, and set the hitting position according to the prediction result; reinforcement learning is only responsible for controlling the hitting posture of the robot racket, and its typical representative is the KUKA robot of the University of Tübingen. This method requires multiple different systems (table tennis trajectory prediction, robot joint trajectory planning) to coordinate with each other, which will lead to more communication delays between multiple different systems, and also puts higher requirements on the robustness of each system itself.
[0004] 2. Directly read the position of the table tennis ball, and obtain the hitting position and hitting posture through reinforcement learning. Use robotics to implement trajectory planning and control the robot to reach the target pose at the target moment, and its typical representative is the SIASUN robot. This method reduces the coupling degree of multiple systems to a certain extent, but still requires the trajectory planning system to coordinate with the hitting strategy. Moreover, how to ensure that the target pose generated by the hitting strategy has good continuity and feasibility is also a difficult problem. In addition, the robot developed by SIASUN is a parallel robot, which results in a small working space. Therefore, during the actual demonstration process, a specially made reduced-size table tennis table needs to be used, which increases the cost of deploying the robot in reality and reduces the versatility of the robot.
[0005] 3. Directly read the position of the table tennis ball and directly output the robot joint angles to construct an end-to-end hitting ball strategy model. However, in order to reduce the implementation difficulty of robot reinforcement learning, special modifications need to be made to the robot structure, such as a robot controlled by pneumatic muscles, and its typical representative is the Max Planck Institute for Intelligent Systems. This method reduces the reinforcement learning training difficulty by making special modifications to the robot structure. Although the hitting ball strategy of the robot can be more easily implemented. However, a more complex structure will not only greatly increase the construction cost of the robot entity, but also significantly increase the control difficulty.
[0006] The applicant's prior application CN115120949A discloses a method, system, and storage medium for implementing a flexible hitting strategy for a robot, which discloses setting rewards for four trajectory stages that make up a complete table tennis trajectory. The first trajectory stage and the second trajectory stage are the opponent's serve trajectory stage and the robot's receiving trajectory respectively, and the third trajectory stage and the fourth trajectory stage are the robot's counterattack trajectory and the opponent's receiving trajectory respectively. Specifically, the rewards for the four trajectory stages are as follows: making the sum of the rewards for the first trajectory stage and the second trajectory stage inversely proportional to the distance between the ball and the robot's racket; making the reward for the third trajectory stage inversely proportional to the distance between the ball and the target point. In this patent, the reward setting is relatively simple. In actual reinforcement learning, the obtained hitting strategy is relatively elementary and difficult to handle complex scenarios. Summary of the Invention
[0007] Object of the Invention: In order to overcome the deficiencies in the prior art, the present invention provides an end-to-end robot reinforcement learning hitting strategy implementation method based on stage rewards with more reasonable task setting and task reward setting, which enables the robot to perform complex table tennis tasks.
[0008] Technical Solution: To achieve the above object, the end-to-end robot reinforcement learning hitting strategy implementation method based on stage rewards of the present invention includes:
[0009] Obtain the state information of the robot and the state information of the table tennis ball as observation items for reinforcement learning;
[0010] Perform reinforcement learning for receiving training based on the training objective of the first stage and the learning task reward to obtain a first pre-trained model; wherein, the training objective of the first stage is to make the racket contact the ball.
[0011] Perform reinforcement learning for hitting training based on the first pre-trained model, the training objective of the second stage, and the learning task reward to obtain a second pre-trained model; wherein, the training objective of the second stage is that the ball hit by the racket faces the opponent's table and the table tennis ball can fly over the net.
[0012] Perform reinforcement learning for target point hitting training based on the second pre-trained model, the training objective of the third stage, and the learning task reward to obtain an output result, the output result includes joint parameters of each joint of the robot, and the joint parameters include joint position and joint speed; wherein, the training objective of the third stage is that the landing point of the table tennis ball hit back by the robot is within a preset target area.
[0013] Furthermore, the learning task reward for the first stage is:
[0014]
[0015] The learning task rewards in the second stage are as follows:
[0016]
[0017] The learning task rewards in the third stage are as follows:
[0018]
[0019] Where: τ s is the trajectory state, and τ s = 0, 1, 2, 3 represent the opponent's serve trajectory, the robot's receiving trajectory, the robot's counterattack trajectory, and the opponent's receiving trajectory respectively; τ represents the change of the trajectory state τ s The moments when (τ: 0 → 1), (τ: 1 → 2), (τ: 2 → 3) represent the moment when the table tennis ball collides on the table surface on the robot's side, the moment when the robot's racket successfully catches the table tennis ball, and the moment when the robot successfully hits the ball onto the opponent's table surface respectively;
[0020] r is the sparse reward, and d is the continuous reward;
[0021] r rab is the reward for the robot body to avoid the table tennis ball, which only needs to be considered when τ s = 0; d rb is the reward for the distance between the racket and the ball;
[0022] r rhb is the reward for the robot hitting the ball; v rhb_x is the racket speed towards the opponent's table when the robot hits the ball;
[0023] r rho is the reward for the robot hitting the ball to the opponent's side; r dlt is the distance between the landing point of the ball and the target point when the robot hits the ball onto the opponent's table.
[0024] Furthermore, the robot includes two sliding guide rails and a four-degree-of-freedom robotic arm.
[0025] Furthermore, the state information of the robot includes joint positions, joint velocities, racket positions, racket linear velocities, and racket angular velocities.
[0026] Furthermore, the state information of the table tennis ball includes the table tennis ball position, the table tennis ball linear velocity, the target point position, the trajectory state, and the trajectory state expressed by a one-hot vector.
[0027] Furthermore, the reinforcement learning is based on the PPO algorithm.
[0028] An end-to-end robot reinforcement learning hitting strategy implementation device based on stage rewards, which includes:
[0029] An acquisition module, which is used to acquire the state information of the robot and the state information of the table tennis ball as observation items for reinforcement learning;
[0030] A first training module, which is used to perform reinforcement learning for receiving ball training based on the training objectives of the first stage and the learning task rewards to obtain a first pre-trained model; wherein, the training objective of the first stage is to make the racket contact the ball;
[0031] A second training module, which is used to perform reinforcement learning for playing ball training based on the first pre-trained model, the training objectives of the second stage and the learning task rewards to obtain a second pre-trained model; wherein, the training objective of the second stage is that the ball hit by the racket faces the opponent's table and the table tennis ball can fly over the net;
[0032] A third training module, which is used to perform reinforcement learning for target click and hit training based on the second pre-trained model, the training objectives of the third stage and the learning task rewards to obtain an output result, the output result includes the joint parameters of each joint of the robot, and the joint parameters include joint positions and joint velocities; wherein, the training objective of the third stage is that the landing point of the table tennis ball hit back by the robot is within a preset target area.
[0033] Advantages: The method for implementing an end-to-end robot reinforcement learning hitting strategy based on stage rewards of the present invention has the following advantages: (1) It proposes a table tennis robot reinforcement learning that can be achieved by a single end-to-end policy system, and uses PD servo control to replace the trajectory planning system. This method will have lower system latency and higher robustness compared to the multi-system coordination scheme; (2) It uses a traditional mechanical structure as the robot body instead of special structures such as pneumatic muscles, and the existence of a serial manipulator can also greatly increase the working space of the robot, so there is no need to use a specially customized table tennis table. The above approach can significantly reduce the deployment cost of the robot in the real environment and greatly improve the deployability of the robot; (3) The present invention proposes the definition of trajectory states and the setting of a staged reinforcement learning reward function. This approach can ensure that the robot can not only learn complex table tennis tasks, but also control the landing point of the hit-back table tennis ball within a certain range, so as to better interact with people. Description of the Drawings
[0034] Figure 1 It is a schematic flowchart of a method for implementing an end-to-end robot reinforcement learning hitting strategy based on stage rewards;
[0035] Figure 2 It is a structural diagram of the robot and a schematic diagram of each trajectory stage;
[0036] Figure 3It is the network structure diagram of the PPO algorithm;
[0037] Figure 4 It is the schematic composition diagram of the device for implementing the end-to-end robot reinforcement learning hitting strategy based on stage rewards. Specific implementation manners
[0038] The present invention will be further described below in conjunction with the accompanying drawings.
[0039] Such as Figure 1 The end-to-end robot reinforcement learning hitting strategy implementation method shown, the method includes the following steps S101-S104:
[0040] Step S101, obtain the state information of the robot and the state information of the table tennis ball as the observation items of reinforcement learning;
[0041] Step S102, perform reinforcement learning for receiving ball training based on the training objective of the first stage and the learning task reward to obtain a first pre-trained model; wherein, the training objective of the first stage is to make the racket contact the ball;
[0042] In this step, the receiving ball training is the training of the first stage, and its training objective is to make the racket of the robot hit the table tennis ball hit by the opponent as many times as possible. However, when the robot learns to hit the table tennis ball, since the speed of the racket is not in the direction of the opponent when the racket contacts the table tennis ball, it is often impossible to hit the ball onto the opponent's tabletop. Therefore, it is necessary to introduce the training task of the second stage. The robot plays the ball based on the first pre-trained model, which can effectively improve the probability of hitting the ball.
[0043] Step S103, perform reinforcement learning for hitting ball training based on the first pre-trained model, the training objective of the second stage and the learning task reward to obtain a second pre-trained model; wherein, the training objective of the second stage is that the ball hit by the racket faces the opponent's tabletop and the table tennis ball can fly over the net;
[0044] In this step, the hitting ball training is the training of the second stage, and its training objective requires that the table tennis ball hit by the robot can fly towards the opponent's tabletop and cross the net, which requires that the hitting speed of the racket cannot be too small, that is, the initial speed given to the table tennis ball cannot be too small. The racket should have a large enough speed towards the opponent's tabletop to ensure that the table tennis ball can fly over the net and reach the opponent's tabletop. However, there is often a problem in the training task of the second stage that the speed of the racket face is too large, resulting in most of the hit table tennis balls going out of bounds. Therefore, it is necessary to introduce the training task of the third stage. The robot plays the ball based on the second pre-trained model obtained in this step, which can effectively improve the probability of successfully hitting the ball back to the opponent's side and crossing the net.
[0045] Step S104: Based on the second pre-trained model, the training objective in the third stage, and the learning task reward, perform reinforcement learning for target hitting training to obtain an output result, where the output result includes joint parameters of each joint of the robot, and the joint parameters include but are not limited to joint positions and joint velocities; wherein, the training objective in the third stage is that the landing point of the table tennis ball hit back by the robot is within a preset target area.
[0046] In this step, the target hitting training is the training in the third stage, and the preset target area is the range of the opponent's table. Making the hit table tennis ball fall into this area can ensure that the ball hit back by the robot does not go out of bounds, and thus an effective return ball can be completed.
[0047] In the above steps S101 - S104, the learning to achieve the robot's hitting task is divided into three stages, namely the above-mentioned receiving training, hitting training, and target hitting training. By setting the training objectives and learning task rewards for each stage, it is possible to successively complete being able to receive the ball, being able to return the ball to the opponent's side and successfully cross the net, and making the table tennis ball fall into the opponent's table area, gradually completing the construction of the hitting strategy model.
[0048] Dividing the training task into three stages instead of directly performing the task training in the third stage is because in the initial stage, when the robot cannot play table tennis, it is a small probability event for the robot to successfully receive the table tennis ball or hit the table tennis ball to the opponent's table, which will lead to too sparse rewards for the robot and cause the training to fail.
[0049] In the above method, a end-to-end hitting strategy system is constructed, that is, the input is the state information of the table tennis ball and the state information of the robot, and the output is the joint positions of the table tennis robot. For the table tennis robot training algorithm constructed by this method, since only the influence of a single end-to-end system needs to be considered, the problem of communication delay between multiple systems will not exist. And a single system also has a lower probability of making mistakes compared to multiple systems. To solve the problem that directly giving the pose of the racket by reinforcement learning may result in the trajectory being discontinuous due to exceeding the working space and sudden changes in the target racket pose, the present invention does not consider using the method of trajectory planning to plan the robot joint trajectory. Instead, the present invention uses the method of directly giving the joint positions to control the robot, thus solving the above-mentioned problems.
[0050] Furthermore, the learning task reward in the first stage is:
[0051]
[0052] The learning task reward in the second stage is:
[0053]
[0054] The learning task rewards in the third stage are as follows:
[0055]
[0056] Among them: τ s is the trajectory state. As shown in Figure 2 τ s = 0, 1, 2, 3 represent the opponent's serve trajectory, the robot's ball-catching trajectory, the robot's counterattack trajectory, and the opponent's ball-catching trajectory respectively; τ represents the change of the trajectory state τ s . (τ: 0 → 1), (τ: 1 → 2), (τ: 2 → 3) represent the moment when the table tennis ball collides on the table surface where the robot is located, the moment when the robot's racket successfully catches the table tennis ball, and the moment when the robot successfully hits the ball onto the opponent's table surface respectively;
[0057] r is the sparse reward, and d is the continuous reward;
[0058] r rab is the reward for the robot body to avoid the table tennis ball, which needs to be considered only when τ s = 0; d rb is the distance reward between the racket and the ball;
[0059] r rhb is the reward for the robot hitting the ball; v rhb_x is the racket speed towards the opponent's table when the robot hits the ball;
[0060] r rho is the reward for the robot hitting the ball to the opponent's side; r dlt is the distance between the landing point of the ball and the target point when the robot hits the ball onto the opponent's table.
[0061] In the above reward strategy, not only the trajectory states corresponding to τ s = 0, 1, 2, 3 in the prior application are adopted, but also multiple other parameters are introduced as considerations for task rewards, such as: (τ: 0 → 1), (τ: 1 → 2), (τ: 2 → 3), r rab , d rb , r rhb and other parameters. In addition, a specific formula for task reward setting is constructed. In comparison, the task reward design in the present invention considers more factors and is more reasonable. The data model trained based on the above reward strategy can handle the ball-playing tasks in more complex scenarios.
[0062] Preferably, the robot includes a Cartesian robot 02 composed of two sliding guides and a four-degree-of-freedom robotic arm 01. The illustrated robotic arm has 5 degrees of freedom, but only 4 degrees of freedom are used in actual use, and 1 degree of freedom is locked. The Cartesian robot 02 can move the robotic arm 01 in the horizontal and vertical directions to increase the movement area of the robotic arm 01. The racket is fixed at the execution end of the robotic arm 01. This can effectively increase the working space of the robot. The actually adopted five-degree-of-freedom rigid robotic arm not only has the characteristics of being easy to control, but also can greatly reduce the manufacturing cost of the robot. The two-degree-of-freedom slide rail can greatly increase the working space of the robot, so that its hitting range can cover the entire table tennis table.
[0063] Preferably, the state information of the robot in step S101 includes joint positions (7 parameters. The robot has 6 degrees of freedom that can move, but the actual robotic arm has five degrees of freedom, and the locked degree of freedom is also counted among them. For the same reason, all other joint parameters of the robot involved below are also 7 parameters), joint speeds (7 parameters), racket positions (3 parameters), racket linear speeds (3 parameters), and racket angular speeds (3 parameters).
[0064] Preferably, the state information of the table tennis ball in step S101 includes table tennis ball positions (3 parameters), table tennis ball linear speeds (3 parameters), target point positions (2 parameters), trajectory state τ s (1 parameter) and the trajectory state τ expressed by a one-hot vector s (4 parameters).
[0065] The state information of the above-mentioned robot and the state information of the table tennis ball are collectively referred to as observation items, that is, the observation items are composed of 37 dimensions in total. The above output data are the joint parameters of each joint of the robot, which is a 7-dimensional data.
[0066] Preferably, the reinforcement learning is based on the PPO algorithm. In its network structure, Actor and Critic share parameters before the last layer. Its network structure is generally the D2RL network structure proposed by Google. The overall network structure is as Figure 3 shown (the illustration is of Actor, so the output is Action). This network structure is a prior art and will not be introduced in detail here. Figure 3 The content within the left dashed box in represents a D2RL block, and each D2RL block on the right is in a simplified form. In fact, each D2RL block on the right can be represented in the form of the content within the left dashed box.
[0067] The present invention also discloses an end-to-end robot reinforcement learning hitting strategy implementation device based on stage rewards. The hitting strategy implementation device may include or be divided into one or more program modules. The one or more program modules are stored in a storage medium and executed by one or more processors to complete the present invention and implement the above-mentioned hitting strategy implementation method. The program modules referred to in the embodiments of the present invention refer to a series of computer program instruction segments that can complete specific functions, and are more suitable for describing the execution process of the hitting strategy implementation device in the storage medium than the program itself. The following description will specifically introduce the functions of each program module in this embodiment, as Figure 4 shown, which includes:
[0068] An acquisition module 201, which is used to acquire the state information of the robot and the state information of the table tennis ball as the observation items for reinforcement learning;
[0069] A first training module 202, which is used to perform reinforcement learning for receiving ball training based on the training objectives of the first stage and the learning task rewards to obtain a first pre-trained model; wherein, the training objective of the first stage is to make the racket contact the ball;
[0070] A second training module 203, which is used to perform reinforcement learning for playing ball training based on the first pre-trained model, the training objectives of the second stage and the learning task rewards to obtain a second pre-trained model; wherein, the training objective of the second stage is that the ball hit by the racket faces the opponent's table and the table tennis ball can fly over the net;
[0071] A third training module 204, which is used to perform reinforcement learning for target click hitting training based on the second pre-trained model, the training objectives of the third stage and the learning task rewards to obtain an output result, and the output result includes the joint parameters of each joint of the robot, and the joint parameters include but are not limited to joint positions and joint speeds; wherein, the training objective of the third stage is that the landing point of the table tennis ball hit back by the robot is within a preset target area.
[0072] Other content for implementing the above-mentioned hitting strategy implementation method based on the hitting strategy implementation device has been introduced in detail in the previous embodiments. Reference can be made to the corresponding content in the previous embodiments, and details will not be described here.
[0073] The above is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An implementation method of an end-to-end robot reinforcement learning batting strategy based on stage rewards, characterized in that, The method includes: Obtaining the state information of the robot and the state information of the table tennis ball as observation items for reinforcement learning; Performing reinforcement learning for receiving ball training based on the training objective of the first stage and the learning task reward to obtain a first pre-trained model; wherein, the training objective of the first stage is to make the racket contact the ball; Performing reinforcement learning for hitting ball training based on the first pre-trained model, the training objective of the second stage and the learning task reward to obtain a second pre-trained model; wherein, the training objective of the second stage is that the ball hit by the racket faces the opponent's table and the table tennis ball can fly over the net; Performing reinforcement learning for target point hitting training based on the second pre-trained model, the training objective of the third stage and the learning task reward to obtain an output result, the output result includes the joint parameters of each joint of the robot, and the joint parameters include joint position and joint speed; wherein, the training objective of the third stage is that the landing point of the table tennis ball hit back by the robot is within a preset target area; The learning task reward of the first stage is: ; The learning task reward of the second stage is: ; The learning task reward of the third stage is: ; where: τ s is the trajectory state, and τ s = 0, 1, 2, 3 represent the opponent's serving trajectory, the robot's receiving trajectory, the robot's counterattack trajectory, and the opponent's receiving trajectory respectively; τ represents the change of the trajectory state τ s The moments when (τ: 0 → 1), (τ: 1 → 2), and (τ: 2 → 3) represent the moment when the table tennis ball collides on the table surface on the robot's side, the moment when the robot's racket successfully catches the table tennis ball, and the moment when the robot successfully hits the ball onto the opponent's table surface respectively; r is a sparse reward, and d is a continuous reward; r rab For the robot body to avoid the ping-pong ball reward, it only needs to be considered when τ s = 0; d rb is the distance reward between the racket and the ball; r rhb Reward for the robot hitting the ball; v rhb_x Racket speed towards the opponent's table when the robot hits the ball; r rho Reward for the robot hitting the ball to the opponent's side; r dlt When the robot hits the ball onto the opponent's table, the distance between the landing point of the ball and the target point.
2. The method for implementing an end-to-end robot reinforcement learning batting strategy based on stage rewards according to claim 1, wherein The robot includes two sliding guides and a four-degree-of-freedom robotic arm.
3. The method for implementing an end-to-end robot reinforcement learning batting strategy based on stage rewards according to claim 1, wherein The state information of the robot includes joint position, joint speed, racket position, racket linear velocity, and racket angular velocity.
4. The method for implementing an end-to-end robot reinforcement learning batting strategy based on stage rewards according to claim 1, wherein The state information of the table tennis ball includes table tennis ball position, table tennis ball linear velocity, target point position, trajectory state, and trajectory state expressed by a one-hot vector.
5. The method for implementing an end-to-end robot reinforcement learning batting strategy based on stage rewards according to claim 1, characterized in that, The reinforcement learning is based on the PPO algorithm.
6. An end-to-end robot reinforcement learning batting strategy implementation device based on stage rewards, which is used to implement the end-to-end robot reinforcement learning batting strategy implementation method described in claim 1, and is characterized in that It includes: An acquisition module, which is used to obtain the state information of the robot and the state information of the table tennis ball as observation items for reinforcement learning; A first training module, which is used to perform reinforcement learning for receiving ball training based on the training objective of the first stage and the learning task reward to obtain a first pre-trained model; wherein, the training objective of the first stage is to make the racket contact the ball; A second training module, which is used to perform reinforcement learning for hitting ball training based on the first pre-trained model, the training objective of the second stage and the learning task reward to obtain a second pre-trained model; wherein, the training objective of the second stage is that the ball hit by the racket faces the opponent's table and the table tennis ball can fly over the net; A third training module, which is used to perform reinforcement learning for target point hitting training based on the second pre-trained model, the training objective of the third stage and the learning task reward to obtain an output result, the output result includes the joint parameters of each joint of the robot, and the joint parameters include joint position and joint speed; wherein, the training objective of the third stage is that the landing point of the table tennis ball hit back by the robot is within a preset target area.
Citation Information
Patent Citations
Realization method and system for flexible ball hitting strategy of table tennis robot and storage medium
CN115120949A
Deep reinforcement learning rotation speed prediction method and system for table tennis robot
CN110458281A
Batting method and device for table tennis robot
CN110711368A