A Magnetic-Field-Based Reward Shaping Method for Reinforcement Learning in Manipulator Control
By treating the target object and obstacle as permanent magnets, calculating the magnetic field strength and converting it into a potential energy reward function, the problem of insufficient design of the reward function in the reinforcement learning robotic arm control is solved, and the learning efficiency and convergence speed are improved.
Patent Information
- Application Number
- CN202210705509.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-21
AI Technical Summary
The existing reinforcement learning robot arm control method is insufficient in complex dynamic environments and cannot provide rich orientation information, resulting in low learning efficiency.
A reward shaping method based on magnetic field is designed, by treating target objects and obstacles as permanent magnets, calculating the magnetic field strength and converting it into a potential energy reward function, and depositing it into an empirical playback pool in combination with the DPBA algorithm to train the optimal strategy of the robotic arm.
While ensuring that the optimal strategy remains unchanged, it provides richer orientation information for the robotic arm, which improves the learning efficiency and convergence speed of reinforcement learning algorithms.
Smart Images

Figure CN115179280B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of robot control, and particularly relates to a magnetic field-based reward shaping method for robotic arm control in reinforcement learning. Background Art
[0002] Traditional robotic arm control methods usually require modeling the robotic arm based on kinematic and dynamic equations and solving for the end-effector pose and the angles of each joint. With the increasing complexity and dynamics of industrial application scenarios, the computational complexity of traditional model-based robotic arm control methods is also getting higher and higher, making it unable to adapt to changes in the external environment in a timely manner, and lacking the ability of autonomous learning and generalization for the environment.
[0003] In recent years, reinforcement learning has been widely applied to robotic arm control tasks due to its unique advantages in dealing with sequential decision-making problems. It directly maps the environmental state information obtained by sensors to the actions executed by the robotic arm to achieve end-to-end control, providing a new solution idea for the control problems of complex continuous high-dimensional systems. The optimization goal of reinforcement learning is to find the optimal policy that maximizes the cumulative reward value in a Markov decision process (MDP). Therefore, designing a scientific reward function is particularly important. Regarding the design of the reward function in robotic arm motion control tasks, the existing methods are relatively simple, including the design of orientation reward functions, heuristic reward functions, etc., which cannot provide rich reward signals for the robotic arm in complex dynamic environments and fail to effectively improve the learning efficiency.
[0004] The patent document with the publication number CN113894787A discloses a design method for a heuristic reward function in robotic arm reinforcement learning motion planning, including: establishing a heuristic function for the robotic arm motion planning problem; constructing a heuristic reward function for the robotic arm motion planning according to the heuristic function; determining the parameter values in the heuristic reward function; and training a neural network motion planner for the robotic arm motion planning using the constructed heuristic reward function. This invention sets the heuristic reward function based on the straight-line distance from the end position of the robotic arm to the target position, which cannot provide higher-order reward signals and cannot guarantee the optimality of the learning strategy. Summary of the Invention
[0005] The object of the present invention is to address the problem that the reward function in existing reinforcement learning robotic arm control methods provides limited information in complex dynamic environments, and proposes a magnetic field-based reward shaping method for robotic arm control in reinforcement learning, which can provide richer orientation information about the target object and obstacles for the robotic arm while ensuring that the optimal policy remains unchanged, thereby improving the learning efficiency and convergence speed of the reinforcement learning algorithm.
[0006] The technical solution of the present invention is as follows: A magnetic field-based reward shaping method for robotic arm control in reinforcement learning, characterized by including the following steps:
[0007] S1. Design the task environment, set the relevant parameters of the robotic arm, target object, and obstacle, and set the hyperparameters of the reinforcement learning algorithm;
[0008] S2. Regard the target object as a square permanent magnet of the same shape, determine its magnetization direction and the calculation method of the three-dimensional space magnetic field intensity distribution, and the same applies to the obstacle;
[0009] S3. The robotic arm interacts with the environment, collects training data, and calculates the magnetic field intensity of the end coordinate of the robotic arm in the magnetic fields of the target object and the obstacle according to the next state. After standardization and normalization processing, a magnetic field reward function is obtained;
[0010] S4. Use the DPBA algorithm to convert the magnetic field reward function into a potential energy-based shaping reward function, and store it in the experience replay pool together with the training data;
[0011] S5. Collect a batch of data from the experience replay pool, and use the reinforcement learning algorithm to train the optimal strategy for the robotic arm to avoid obstacles and reach the target object in the dynamic environment.
[0012] Preferably, the step S1 includes the following steps:
[0013] Step 1.1. Design the state observation value of the task environment and the action value of the robotic arm, specifically including:
[0014] a. The environmental state observation value includes the rotation angles of the three joints of the robotic arm, the coordinates of the end of the robotic arm, and the coordinates of the centers of the target object and the obstacle;
[0015] b. The action value of the robotic arm is the rotation angular speed of the three joint motors, that is, the angles rotated by the three joints in the unit time step.
[0016] Step 1.2. Establish a connection with the robotic arm, set the speed and acceleration ranges of the rotation of the three joints; stipulate the random generation method of the target object and the obstacle to ensure that the target object is within the reachable range of the end of the robotic arm, and the target object and the obstacle do not intersect.
[0017] Step 1.3. Set the basic hyperparameters of the reinforcement learning algorithm, including at least: exploration noise, the size of the experience replay pool ; the number of updates K for each training, the size N of the data batch used for each update; the number of layers of the neural network, the number of nodes and activation functions of each layer; the discount factor γ; the policy network μ θ (s) and the value function network Q φ (s,a) The optimizers, learning rates for parameter updates of the target network and the soft update step size τ.
[0018] In step S2, the analytical calculation method for the magnetic field intensity distribution in the three-dimensional space of the square permanent magnet is as follows:
[0019] Assume that the magnetization direction is the positive direction of the z-axis and the magnetization intensity is M c , for a square permanent magnet with lengths l, w, and h along the x-axis, y-axis, and z-axis respectively, the magnetic field intensity components in the x-axis, y-axis, and z-axis directions at any point P(x, y, z) in the three-dimensional space can be expressed as:
[0020]
[0021] where Γ(γ1, γ2, γ3) and are two auxiliary functions, and the expressions are as follows:
[0022]
[0023] where ∈ is a minimum value. Then, the magnetic field intensity at any point in the three-dimensional space of the square permanent magnet can be obtained as:
[0024]
[0025] Preferably, step S3 includes the following steps:
[0026] Step 3.1, initialize the rotation angles of the three joints of the robotic arm to zero and read the coordinates of the end of the robotic arm; randomly set the positions of the target object and the obstacle, and read the coordinates of the centers of the target object and the obstacle in the world coordinate system. Obtain the initial value of the state observation.
[0027] Step 3.2, the robotic arm outputs an action according to the current state observation s and the policy, adds noise to it to obtain a, and after interacting with the environment, obtains the next state s ′ and the original reward value r. Under the condition of ensuring that the rotation angles of the three joints of the robotic arm in the next state are within their corresponding working ranges, control the robotic arm to move to the next state.
[0028] Step 3.3, convert the coordinates of the end of the robotic arm in the next state from the world coordinate system to the magnetic field coordinate systems of the target magnet and the obstacle magnet.
[0029] Assume that the coordinates of the end of the robotic arm in the next state in the world coordinate system are The translation amount of the origin of the target magnetic field coordinate system relative to the origin of the world coordinate system is (T x , T y , T z), the rotation angles of the target magnetic field coordinate system relative to the world coordinate system around the x-axis, y-axis, and z-axis are θ x , θ y , θ z , and its positive direction follows the right-hand screw rule. Then the coordinates of the end of the robotic arm in the target magnetic field coordinate system can be expressed as:
[0030]
[0031] where are the rotation transformation matrices of the coordinate system around the x-axis, y-axis, and z-axis respectively, as follows:
[0032]
[0033] Step 3.4, calculate the magnetic field intensities of the end coordinates of the robotic arm in the target magnetic field and the obstacle magnetic field in the next state, and perform normalization processing on them.
[0034] Assume that there is 1 target and n obstacles in the environment. The magnetic field intensity calculation functions of the target and obstacle magnets are respectively The coordinates of the end of the robotic arm in the magnetic field coordinate systems of the target and obstacle magnets are respectively Then, the magnetic field intensities of the end coordinates of the robotic arm in the target and obstacle magnetic fields can be calculated as H T ,
[0035] Store H T , in the magnetic field intensity playback pool and according to the current mean μ T , and standard deviation σ T , of the magnetic field intensity in it, map the calculated H T , to the standard Gaussian distribution. The normalized magnetic field intensity is expressed as follows:
[0036]
[0037] Step 3.5, calculate the combined magnetic field intensity of the target and obstacle magnets, and perform normalization processing on it to obtain the magnetic field reward function.
[0038] Define the combined magnetic field intensity of the target and obstacle magnets as:
[0039]
[0040] Normalize the combined magnetic field intensity using the Softsign function and define the output result as the magnetic field reward function r M . Specifically as follows:
[0041]
[0042] Preferably, step S4 includes the following steps:
[0043] Step 4.1, define the potential function neural network in the DPBA algorithm as Φ ψ (s,a), whose input is the state observation value and the action value of the robotic arm, and the output is the potential energy value of the current state-action pair. ψ is the parameter of the neural network. The loss function used to update the potential function neural network is defined as:
[0044]
[0045] Among them, y is the gradient-free "label value" of the potential function, and the specific expression is as follows:
[0046] y = -r M +γΦ ψ (s′,a′)
[0047] Among them, r M is the magnetic field reward function obtained in step 3.5, and γ is the discount factor. Update the parameters of the potential function neural network using the gradient descent method as follows:
[0048]
[0049] Among them, η is the learning rate for updating the potential function neural network.
[0050] Step 4.2, according to the parameters Ψ before update and the parameters Ψ ′ , the potential-based shaping reward function can be calculated as follows:
[0051] f M =γΦ ψ′ (s ′ ,a′)-Φ ψ (s,a)
[0052] When the potential function Φ ψ (s,a) is initialized to zero and updated to final convergence in the above manner, the magnetic field reward function can be completely converted into the potential-based shaping reward function, that is: f M =r M .
[0053] Combine the shaping reward f M with the original reward value r obtained in step 3.2, and use (s,a,r + f M, s ′ ) are stored in the experience replay pool as a set of training data for the subsequent training of the reinforcement learning algorithm. According to the optimal policy invariance theorem, the algorithm is based on the reward function r + f M The optimal policy learned is consistent with the optimal policy learned from the original reward function r. Repeat steps 3.2 to step 4.2 until the end of the robotic arm reaches the target object, or the robotic arm touches an obstacle or the ground, or the set maximum number of time steps is experienced.
[0054] Preferably, the step S5 includes the following steps:
[0055] Step 5.1, randomly sample a batch of data (S, A, R + F M , S ′ ) from the experience replay pool, where (s i , a i , r i + f i M , s i+1 ) represents a single training data.
[0056] Step 5.2, calculate the loss function for updating the parameters of the value function network:
[0057]
[0058] where y i is the gradient-free "label value" of the value function, and the specific expression is as follows:
[0059]
[0060] Update the parameters of the value function network using the gradient descent method as follows:
[0061]
[0062] where β is the learning rate for updating the value function network.
[0063] Step 5.3, calculate the loss function for updating the parameters of the policy network:
[0064]
[0065] Update the parameters of the policy network using the gradient descent method as follows:
[0066]
[0067] where α is the learning rate for updating the policy network.
[0068] Step 5.4, softly update the parameters of the target network:
[0069]
[0070] Step 5.5, repeat steps 5.1 to 5.4 for a total of K times to end this round. Repeat steps S3 to S5 until the algorithm converges completely to obtain the optimal policy network for the robotic arm to avoid obstacles and reach the target in a dynamic environment.
[0071] The beneficial effects of the present invention are as follows: Compared with the prior art, a magnetic field-based reward shaping method for robotic arm control in reinforcement learning proposed by the present invention can provide the robotic arm with richer azimuth information about the target and obstacles while ensuring the invariance of the optimal policy, thereby effectively improving the learning efficiency and convergence speed of the reinforcement learning algorithm in a complex dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 is the task scenario diagram of the embodiment of the present invention in a simulation environment, where A is the end of the robotic arm, B is the obstacle, and C is the target;
[0073] Figure 2 is the overall framework diagram of the algorithm of the present invention;
[0074] Figure 3 is the schematic diagram of the magnetic field coordinate system of the square permanent magnet in the embodiment of the present invention;
[0075] Figure 4 is the experimental effect comparison diagram of the magnetic field-based reward shaping method of the present invention and similar algorithms in a simulation environment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0076] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.
[0077] Embodiment: As shown in the attached Figure 1 figures: In this embodiment, a Dobot Magician robotic arm with 3 degrees of freedom is taken as an example. The designed task scenario is to use a reinforcement learning algorithm to control the robotic arm to complete the task of moving the end to the target without hitting the obstacle in a dynamic environment. Among them, the target and the obstacle are cuboids of different sizes, and their positions are randomly changed in each round. The design framework of a magnetic field reward function for robotic arm control in reinforcement learning described in this embodiment is as shown in the attached Figure 2 figures, and at least includes the following steps:
[0078] Step S1, design the task environment, set the relevant parameters of the robotic arm, the target and the obstacle, and set the various hyperparameters of the reinforcement learning algorithm, specifically including the following steps:
[0079] Step 1.1, design the state observation values of the task environment and the action values of the robotic arm, specifically including:
[0080] a. The environmental state observation values include the rotation angles of the three joints of the robotic arm, the coordinates of the end of the robotic arm, and the coordinates of the centers of the target object and the obstacle;
[0081] b. The action values of the robotic arm are the angular velocities of the three joint motors, that is, the angles rotated by the three joints in a unit time step, and their range is limited to [-1°, 1°].
[0082] Step 1.2, establish a connection with the robotic arm, set the speed and acceleration ranges of the rotation of the three joints; specify the random generation method of the target object and the obstacle to ensure that the target object is within the reachable range of the end of the robotic arm, and the target object and the obstacle do not intersect.
[0083] Step 1.3, set the basic hyperparameters according to the adopted reinforcement learning algorithm. In this embodiment, the DDPG algorithm proposed by "Lillicrap, Timothy P., et al. "Continuous control with deep reinforcement learning." arXiv preprint arXiv:1509.02971(2015)." applicable to the continuous state-action space is adopted. The hyperparameters to be set include: exploration noise experience replay pool with a size of 10 6 ; the number of updates K = 20 for each training, the size N = 128 of the data batch used for each update; the number of layers of the neural network is 2, the number of nodes in each layer is 256, the activation function is ReLU, and the parameters of the neural network are randomly initialized; the discount factor γ = 0.99; the policy network μ θ (s) and the value function network Q φ (s,a) The optimizers for parameter updates are Adam, and the learning rates are 3×10 -4 and 10 -3 , the target network and The soft update step size τ = 10 -3 .
[0084] Step S2, regard the target object as a square permanent magnet of the same shape, determine its magnetization direction and the calculation method of the three-dimensional space magnetic field intensity distribution, and the same applies to the obstacle.
[0085] In this embodiment, a square permanent magnet is taken as an example, and its magnetic field intensity distribution is calculated by an analytical method. For permanent magnets of other shapes, the magnetic field intensity distribution can be obtained by a similar analytical method or a simulation method based on physical simulation. The magnetic field coordinate system of the square permanent magnet is as shown in the appendix Figure 3As shown. Assume that the magnetization direction is the positive direction of the z-axis and the magnetization intensity is M c , for a square permanent magnet with lengths of l, w, and h along the x-axis, y-axis, and z-axis respectively, the magnetic field intensity components in the x-axis, y-axis, and z-axis directions at any point P(x, y, z) in three-dimensional space can be expressed as:
[0086]
[0087] where Γ(γ1, γ2, γ3) and are two auxiliary functions, and the expressions are as follows:
[0088]
[0089] where ∈ is a very small value, and in this embodiment, ∈ = 10 -7 . Then, the magnetic field intensity at any point in three-dimensional space of the square permanent magnet can be obtained as:
[0090]
[0091] In this embodiment, the target object size is set to l = 0.03m, w = 0.045m, h = 0.02m, and the obstacle size is set to l = 0.038m, w = 0.047m, h = 0.12m.
[0092] Step S3, the robotic arm interacts with the environment, collects training data, and calculates the magnetic field intensity of the end coordinate of the robotic arm in the magnetic fields of the target object and the obstacle according to the next state, and obtains the magnetic field reward function after standardization and normalization processing, which specifically includes the following steps:
[0093] Step 3.1, initialize the rotation angles of the three joints of the robotic arm to zero, and read the coordinate of the end of the robotic arm; randomly set the positions of the target object and the obstacle, and read the coordinates of the center points of the target object and the obstacle in the world coordinate system. Obtain the initial value of the state observation.
[0094] Step 3.2, the robotic arm outputs an action according to the current state observation value s and the policy network μ θ (s), applies noise to it to obtain a, and obtains the next state s ′ and the original reward value r after interacting with the environment. While ensuring that the rotation angles of the three joints of the robotic arm are within their respective working ranges in the next state, control the robotic arm to move to the next state. In this embodiment, the working ranges of the three joints are [-90°, 90°], [0°, 85°], [-10°, 90°] respectively, and the original reward function is set as follows:
[0095]
[0096] Step 3.3, convert the coordinates of the end of the robotic arm in the next state from the world coordinate system to the magnetic field coordinate systems of the target magnet and the obstacle magnet. In this embodiment, taking the target magnet as an example, the coordinate transformation of the obstacle magnet can be obtained in the same way.
[0097] Assume that the coordinates of the end of the robotic arm in the next state in the world coordinate system are The translation amount of the origin of the target magnetic field coordinate system relative to the origin of the world coordinate system is (T x , T y , T z ), and the rotation angles of the target magnetic field coordinate system relative to the world coordinate system around the x-axis, y-axis, and z-axis are θ x , θ y , θ z , and its positive direction follows the right-hand screw rule. Then the coordinates of the end of the robotic arm in the target magnetic field coordinate system can be expressed as:
[0098]
[0099] where are the rotation transformation matrices of the coordinate system around the x-axis, y-axis, and z-axis respectively, and are specifically as follows:
[0100]
[0101] Step 3.4, calculate the magnetic field intensities of the coordinates of the end of the robotic arm in the next state in the target magnetic field and the obstacle magnetic field, and perform normalization processing on them.
[0102] Assume that there is 1 target and n obstacles (n = 1 in this embodiment) in the environment, and the magnetic field intensity calculation functions of the target and the obstacle magnets are respectively The coordinates of the end of the robotic arm in the magnetic field coordinate systems of the target and the obstacle magnets are respectively Then, the magnetic field intensities of the coordinates of the end of the robotic arm in the target and the obstacle magnetic fields can be calculated as H T ,
[0103] Since the interval ranges of the magnetic field intensities obtained by different calculation methods are often different, it is necessary to normalize the calculated H T , This invention introduces a magnetic field intensity playback pool to store the values of H T , , and according to the mean value μ of the magnetic field intensity in the current T , and the standard deviation σ T , The calculated H T , is mapped to a standard Gaussian distribution as follows:
[0104]
[0105] where is the magnetic field strength after standardization. In this embodiment, the size of the magnetic field strength playback pool is 10 6 .
[0106] Step 3.5, calculate the combined magnetic field strength of the target object and the obstacle magnet, and perform normalization processing on it to obtain the magnetic field reward function.
[0107] In the present invention, the task of the robotic arm is to move the end to the target object while avoiding obstacles. Therefore, the combined magnetic field strength of the target object and the obstacle magnet is defined as:
[0108]
[0109] where the "attraction" of the target object magnet to the end of the robotic arm is equal to the "repulsion" of all n obstacle magnets to the end of the robotic arm. According to the characteristics of the magnetic field strength distribution, the magnetic field strength near the target object will tend to positive infinity, and the magnetic field strength near the obstacle will tend to negative infinity. In order to control the combined magnetic field strength within a reasonable range and maintain its distribution law, the present invention uses the Softsign function to normalize the combined magnetic field strength, and defines the output result as the magnetic field reward function r M . Specifically as follows:
[0110]
[0111] Step S4, use the DPBA algorithm to convert the magnetic field reward function into a potential-based shaping reward function, and store it in the experience replay pool together with the training data. The DPBA algorithm is proposed in the literature "Harutyunyan A, Devlin S, Vrancx P, et al. Expressing arbitrary reward functions as potential-based advice[C] / / Proceedings of the AAAI sConference on Artificial Intelligence. 2015, 29(1).", which can convert any reward function given by an expert into an expression form that satisfies potential-based reward shaping to meet the optimal policy invariance theorem. Specifically, it includes the following steps:
[0112] Step 4.1, define the potential function neural network in the DPBA algorithm as Φ ψ (s,a), whose input is the state observation value and the action value of the robotic arm, and the output is the potential energy value of the current state-action pair. ψ is the parameter of the neural network. In this embodiment, the potential function neural network has two hidden layers, with 256 nodes in each layer, and the activation function is ReLU for both layers; the optimizer for parameter update is Adam, and the learning rate is 10 -4 . The loss function used to update the potential function neural network is defined as follows:
[0113]
[0114] where y is the gradient-free "label value" of the potential function, and the specific expression is as follows:
[0115] y = -r M + yΦ ψ (s′,a′)
[0116] where r M is the magnetic field reward function obtained in Step 3.5, and γ is the discount factor. Update the parameters of the potential function neural network using the gradient descent method as follows:
[0117]
[0118] where η is the learning rate for updating the potential function neural network.
[0119] Step 4.2, according to the parameters Ψ before update and the parameters Ψ ′ after update, the potential-based shaping reward function can be calculated as follows:
[0120] f M = γΦ ψ′ (s ′ ,a′) - Φ ψ (s,a)
[0121] When the potential function Φ ψ (s,a) is initialized to zero and updated to final convergence in the above manner, the magnetic field reward function can be completely converted into the potential-based shaping reward function, that is: f M = r M .
[0122] Combine the shaping reward f M with the original reward value r obtained in Step 3.2, and store (s,a,r + f M ,s ′ ) as a set of training data in the experience replay pool for the training of subsequent reinforcement learning algorithms. According to the optimal policy invariance theorem, the algorithm uses the reward function r + f MThe learned optimal policy is consistent with the optimal policy learned by the original reward function r. Repeat steps 3.2 to 4.2 until the end of the robotic arm reaches the target object, or the robotic arm touches an obstacle or the ground, or the set maximum number of time steps (set to 200 in this embodiment) is experienced.
[0123] Step S5: Collect a batch of data from the experience replay pool and use the reinforcement learning algorithm to train the optimal policy for the robotic arm to avoid obstacles and reach the target object in a dynamic environment, which specifically includes the following steps:
[0124] Step 5.1: Randomly sample a batch of N groups of data (S, A, R+F M , S ′ ) from the experience replay pool, where (s i , a i , r i +f i M , s i+1 ) represents a single training data.
[0125] Step 5.2: Calculate the loss function for updating the parameters of the value function network:
[0126]
[0127] where y i is the gradient-free "label value" of the value function, and the specific expression is as follows:
[0128]
[0129] Update the parameters of the value function network using the gradient descent method as follows:
[0130]
[0131] where β is the learning rate for updating the value function network.
[0132] Step 5.3: Calculate the loss function for updating the parameters of the policy network:
[0133]
[0134] Update the parameters of the policy network using the gradient descent method as follows:
[0135]
[0136] where α is the learning rate for updating the policy network.
[0137] Step 5.4: Soft update the parameters of the target network:
[0138]
[0139] Step 5.5, repeat Steps 5.1 to 5.4 for a total of K times to end this round. Repeat Steps S3 to S5 until the algorithm converges completely, and obtain the optimal policy network for the robotic arm to avoid obstacles and reach the target object in a dynamic environment.
[0140] In this embodiment, the method disclosed in the present invention is compared with similar algorithms in the attached Figure 1 simulation environment for training effects, and the results are compared as attached Figure 4 shown. It can be seen from the figure that the success rate of each round in the learning process of the reward shaping method based on magnetic field is significantly higher than that of the original reward and the reward shaping method based on distance. Therefore, it can be seen that the method disclosed in the present invention can effectively improve the learning efficiency of the reinforcement learning algorithm in the robotic arm motion control task.
[0141] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the creative concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A magnetic field-based reward shaping method for robotic arm control in reinforcement learning, characterized in that, It includes the following steps: S1. Design the task environment, set the relevant parameters of the robotic arm, the target object, and the obstacle, and set the hyperparameters of the reinforcement learning algorithm; S2. Regard the target object as a square permanent magnet of the same shape, determine its magnetization direction and the calculation method of the three-dimensional space magnetic field intensity distribution. The same applies to the obstacle; S3. The robotic arm interacts with the environment to collect training data. Initialize the rotation angles of the three joints of the robotic arm to zero; randomly set the positions of the target object and the obstacle to obtain the initial value of the state observation; The robotic arm outputs an action according to the current state observation value s and the policy, adds noise to it to obtain a, and after interacting with the environment, obtains the next state s ′ and the original reward value r; control the robotic arm to move to the next state; And calculate the magnetic field intensity of the end coordinate of the robotic arm in the magnetic fields of the target object and the obstacle according to the next state. After standardization and normalization processing, obtain the magnetic field reward function; S4. Use the DPBA algorithm to convert the magnetic field reward function into a potential-based shaping reward function. Define the potential function neural network in the DPBA algorithm as Φ(s,a), where ψ is the parameter of the neural network. Calculate the potential-based shaping reward function according to the parameter Ψ before update and the parameter Ψ after update. Combine the shaping reward function with the original reward value r and store them together with the training data in the experience replay pool. ψ (s,a), where ψ is the parameter of the neural network; According to the parameter Ψ before update and the parameter Ψ ′ after update, calculate the potential-based shaping reward function; Combine the shaping reward function with the original reward value r and store them together with the training data in the experience replay pool; S5. Sample a batch of data from the experience replay pool, and use the reinforcement learning algorithm to train the optimal strategy for the robotic arm to avoid obstacles and reach the target object in the dynamic environment.
2. The method for magnetic field-based reward shaping in the control of a robotic arm for reinforcement learning according to claim 1, wherein The step S1 includes the following steps: Step 1.
1. Design the state observation value of the task environment and the action value of the robotic arm, specifically including: a. The environmental state observation value includes the rotation angles of the three joints of the robotic arm, the coordinate of the end of the robotic arm, and the coordinates of the center points of the target object and the obstacle; b. The action value of the robotic arm is the rotation angular velocity of the three joint motors, that is, the angles rotated by the three joints in the unit time step; Step 1.
2. Establish a connection with the robotic arm, set the speed and acceleration ranges of the rotation of the three joints; stipulate the random generation method of the target object and the obstacle to ensure that the target object is within the reachable range of the end of the robotic arm, and the target object and the obstacle do not intersect; Step 1.3, set the basic hyperparameters of the reinforcement learning algorithm, including at least: exploration noise, the size of the experience replay pool D R ; the number of updates K per training, the size N of each data batch used for each update; the number of layers of the neural network, the number of nodes in each layer, the activation function; the discount factor γ; the policy network μ θ (s) and the value function network Q φ (s,a) the optimizer for parameter update, the learning rate, the target network and the soft update step size τ of 3. The method for magnetic field-based reward shaping in the control of a robotic arm for reinforcement learning according to claim 2, wherein In the step S2, the analytical calculation method of the magnetic field intensity distribution of the square permanent magnet in three-dimensional space is as follows: Assume that the magnetization direction is the positive direction of the z-axis and the magnetization intensity is M c , for a square permanent magnet with lengths of l, w, and h along the x-axis, y-axis, and z-axis respectively, the magnetic field intensity components in the x-axis, y-axis, and z-axis directions at any point P(x, y, z) in three-dimensional space are expressed as: where, Γ(γ1,γ2,γ3) and are two auxiliary functions, and the expressions are as follows: Where, ∈ is a minimum value; thus, the magnetic field intensity at any point in three-dimensional space of the square permanent magnet is obtained as:
4. The method for magnetic field-based reward shaping in the control of a robotic arm for reinforcement learning according to claim 3, characterized in that, The step S3 includes the following steps: Step 3.
1. Initialize the rotation angles of the three joints of the robotic arm to zero, and read the coordinate of the end of the robotic arm; randomly set the positions of the target object and the obstacle, and read the coordinates of the center points of the target object and the obstacle in the world coordinate system to obtain the initial value of the state observation; Step 3.2: The robotic arm outputs an action according to the current state observation value s and the policy, adds noise to it to obtain a, and after interacting with the environment, obtains the next state s ′ and the original reward value r; under the condition of ensuring that the rotation angles of the three joints of the robotic arm in the next state are within their respective working ranges, control the robotic arm to move to the next state; Step 3.
3. Convert the coordinate of the end of the robotic arm in the next state from the world coordinate system to the magnetic field coordinate systems of the target object magnet and the obstacle magnet; Assume that the coordinates of the end of the robotic arm in the world coordinate system in the next state are The translation of the origin of the target magnetic field coordinate system relative to the origin of the world coordinate system is (T x , T y , T z ), and the rotation angles of the target magnetic field coordinate system relative to the world coordinate system around the x-axis, y-axis, and z-axis are θ x , θ y , θ z . The positive direction follows the right-hand screw rule. Then the coordinates of the end of the robotic arm in the target magnetic field coordinate system are expressed as: Among them, are respectively the rotation transformation matrices of the coordinate system around the x-axis, y-axis, and z-axis, which are specifically as follows: Step 3.
4. Calculate the magnetic field intensity of the coordinate of the end of the robotic arm in the magnetic fields of the target object and the obstacle in the next state, and perform standardization processing on it: Assume that there is 1 target object and n obstacle objects in the environment. The magnetic field strength calculation functions of the target object and the obstacle object magnets are respectively The coordinates of the end of the robotic arm in the magnetic field coordinate systems of the target object and the obstacle object magnets are respectively Therefore, the magnetic field strengths of the coordinates of the end of the robotic arm in the magnetic fields of the target object and the obstacle object are calculated as respectively Store in the magnetic field strength playback pool and, based on the current mean value of the magnetic field strength and standard deviation map the calculated onto a standard Gaussian distribution. The magnetostatic field strength after standardization is expressed as follows: Step 3.
5. Calculate the combined magnetic field intensity of the target object and the obstacle magnet, and perform normalization processing on it to obtain the magnetic field reward function: Define the combined magnetic field intensity of the target object and the obstacle magnet as: Normalize the combined magnetic field intensity using the Softsign function and define the output result as the magnetic field reward function r M , as follows:
5. The method for magnetic field-based reward shaping in the control of a robotic arm for reinforcement learning according to claim 4, wherein The step S4 includes the following steps: Step 4.1, define the potential function neural network in the DPBA algorithm as Φ ψ (s,a), whose input is the state observation value and the action value of the robotic arm, and the output is the potential energy value of the current state-action pair. ψ is the parameter of the neural network. The loss function used to update the potential function neural network is defined as: Where, y is the gradient-free "label value" of the potential function, and the specific expression is as follows: y = -r M +γΦ ψ (s′, a′) where r M is the magnetic field reward function obtained in step 3.5, γ is the discount factor; the parameters of the potential function neural network are updated by the gradient descent method as follows: Where, η is the learning rate for updating the potential function neural network; Step 4.2, according to the parameter Ψ before update and the parameter Ψ after update ′ , calculate the potential energy-based shaping reward function as follows: f M = γΦ ψ′ (s ′ , a′) - Φ ψ (s, a) When the potential function Φ ψ (s,a) is initialized to zero and updated to convergence in the above manner, the magnetic field reward function is fully converted into a potential-based shaping reward function, i.e.: f M = r M ; Combine the shaping reward f M with the original reward value r obtained in step 3.2, and use (s, a, r + f M , s ′ ) as a set of training data and store it in the experience replay pool for the training of subsequent reinforcement learning algorithms; According to the optimal policy invariance theorem, the optimal policy learned by the algorithm with the reward function r + f M is consistent with the optimal policy learned by the original reward function r; Repeat steps 3.2 to 4.2 until the end of the robotic arm reaches the target object, or the robotic arm touches an obstacle or the ground, or the set maximum number of time steps is experienced.
6. The method for magnetic field-based reward shaping in the control of a robotic arm for reinforcement learning according to claim 5, wherein The step S5 includes the following steps: Step 5.1, randomly sample a batch of data (S, A, R+F M , S ′ ) from the experience replay pool, where represents a single training data; Step 5.
2. Calculate the loss function for updating the parameters of the value function network: where y i is the gradient-free "label value" of the value function, and is specifically expressed as follows: Update the parameters of the value function network using the gradient descent method as follows: Where, β is the learning rate for updating the value function network; Step 5.
3. Calculate the loss function for updating the parameters of the policy network: Update the parameters of the policy network using the gradient descent method as follows: Among them, α is the learning rate for updating the policy network; Step 5.4, softly update the parameters of the target network: Step 5.5, repeat Steps 5.1 to 5.4 for a total of K times to end this round; repeat Steps S3 to S5 until the algorithm converges completely to obtain the optimal policy network for the robotic arm to avoid obstacles and reach the target object in a dynamic environment