Visual servo control method and equipment for robotic arm based on target-condition reinforcement learning
Through the reinforcement learning of target conditions, the visual servo control strategy is dynamically generated, which solves the problem of insufficient flexibility in complex environments of traditional methods and realizes efficient control of robotic arms in complex environments.
Patent Information
- Application Number
- CN202510596524.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Traditional visual servo control methods lack flexibility and adaptability in complex and changing task scenarios and environments, and cannot dynamically adjust control strategies based on real-time visual information, resulting in poor control effects and high development costs.
A method based on target condition reinforcement learning is adopted, and a security reinforcement learning model is constructed using deep deterministic strategy gradient algorithm (DDPG), Lagrangian constraints, post-empirical playback algorithm (HER) and safety-critic network (safety-critic). The feature points of the QR code are captured by the robotic arm camera, and the control strategy is generated dynamically to ensure that the robotic arm completes the visual servo task in a complex environment.
It improves the flexibility and adaptability of the robotic arm in complex environments, reduces the cost of manual design and commissioning, and can maintain good control effect under target object position and posture changes and environmental interference.
Smart Images

Figure CN120134324B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual servo control, and in particular to a method and device for visual servo control of a robotic arm based on target condition reinforcement learning. Background Art
[0002] Visual servo control plays a vital role in the field of robotic arms. It aims to precisely control the motion of robotic arms through visual feedback, enabling them to complete complex tasks such as grasping, assembly, and welding. Traditional visual servo control methods are mostly based on fixed control strategies and predefined rules.
[0003] Existing technologies often lack flexibility and adaptability when faced with complex and ever-changing task scenarios and environments. Unable to dynamically adjust control strategies based on real-time visual information and task objectives, significant changes in the target object's position or posture, or interference from the environment, can significantly reduce control effectiveness and even render the task impossible. Furthermore, traditional methods require extensive manual design and debugging, resulting in high development costs and low efficiency.
[0004] The prior art, publication number CN113146616A, discloses a visual servo control method for a four-degree-of-freedom robotic arm. First, the entire arm mode of the robotic arm is set to visual servo mode. During each control cycle of the visual servo mode, the validity of the visual measurement pose data is determined. If the data is invalid for multiple consecutive cycles, the robotic arm stops moving, the entire arm mode is switched to servo standby mode, and the joint control mode is switched to position servo mode. During each control cycle of the visual servo mode, if the visual measurement pose data is valid, the joint control mode is in velocity control mode, and the planned four-dimensional velocity VW_POR of the end is calculated and output. The planned joint angular velocity and planned joint angular position are then obtained through inverse kinematics and output as control instructions to control the angular velocity and angular position of the joint in the next control cycle. This method requires a large amount of parameter information and is computationally intensive, resulting in a long robotic arm response time. Summary of the Invention
[0005] Purpose of the invention: In response to the above-mentioned existing technologies, a method and device for visual servo control of a robotic arm based on target-condition reinforcement learning are proposed; in complex and changeable scenarios, this method utilizes target-condition reinforcement learning to dynamically generate control strategies based on real-time visual information and task objectives, thereby improving the flexibility and adaptability of visual servo control, reducing manual design and debugging costs, and being able to demonstrate good performance in environments with interference factors.
[0006] Technical solution: To achieve the above technical objectives, the present invention provides a method for controlling a robotic arm visual servo based on target condition reinforcement learning. A camera is set at the end of the robotic arm and a robotic arm environment is constructed. The steps are as follows:
[0007] S1. Set four pixel points with known coordinates in the image captured by the robot arm camera as fixed target points;
[0008] S2. Set a QR code on the object as the recognition target for the robot arm visual servo task. The recognition target is fixed in position after initialization.
[0009] S3: The robotic arm obtains four moving feature points by photographing the QR code with a camera. Using the secure reinforcement learning model, the pixel coordinates of the moving feature points and the fixed target point are set as the achieved target and the desired target in the target-conditional reinforcement learning, respectively.
[0010] S4. The safety reinforcement learning model includes the deep deterministic policy gradient algorithm DDPG, Lagrange constraints, the hindsight experience replay algorithm HER, and the safety critic network safety-critic. The DDPG algorithm includes an actor network and a critic network.
[0011] S5. After the safety reinforcement learning model is trained using the historical interaction data of the robotic arm environment, the current state and desired goal are directly input into the actor network in DDPG. The safety reinforcement learning model can then directly output the target speed movement of the robotic arm joints based on the information of the identified target captured by the camera, control the robotic arm to lock onto the target with the QR code, and control the movement of the robotic arm so that the moving feature points move to the pixel coordinates of the four fixed target points in the captured image, thus completing the visual servo control task.
[0012] Furthermore, before the robotic arm camera captures a fixed target point, it is necessary to determine the camera image direction as follows:
[0013] The camera's screen direction is the vector in the camera coordinate system that starts from the center of the shooting screen and points to the top of the screen. The camera's screen direction is always fixed in the camera coordinate system and is set to ;
[0014] The direction of the camera image in the world coordinate system will change with the movement of the robot arm. By calculating the representation of the camera image direction in the world coordinate system To ensure the correct refresh of the camera image; obtain the relative position information of each joint of the robotic arm, vector The calculation process is as follows:
[0015] ;
[0016] In the above formula Represents the coordinate system around Axis rotation The rotation matrix of the angle, Represents the coordinate system around Axis rotation The rotation matrix of the angle, Represents the coordinate system around Axis rotation The rotation matrix of the angle;
[0017] ;
[0018] The above formula is the rotation matrix between the coordinate systems of each robot joint, where For joints Coordinate system relative to joint The rotation matrix of the coordinate system, For joints The rotation angle, Represent the world coordinate system, the robot arm end gripper coordinate system and the camera coordinate system respectively;
[0019] Representation of the camera image direction in the world coordinate system as follows:
[0020] ;
[0021] Where, is the rotation matrix of the camera coordinate system relative to the gripper coordinate system at the end of the robotic arm, Represents the rotation matrix of the manipulator end gripper coordinate system relative to the manipulator joint 8 coordinate system, Indicates the camera image direction in the camera coordinate system.
[0022] Furthermore, the safety reinforcement learning model performs HER and standardization operations on the robot arm's environment historical interaction data: actions, states, rewards, achieved goals, and expected goals to complete preprocessing, and reconstructs the reward data to alleviate the reward sparsity problem of the robot arm environment. A safety-critic network is constructed using a cost function. The safety-critic network inputs the robot arm's environment historical interaction cost data stored in the replay buffer, and the critic network in the DDPG network inputs the robot arm's environment historical interaction reward data in the replay buffer. The regularization term and Lagrange constraint term in the actor network in the DDPG network are used to restrict the robot arm's movement to ensure the safety of the learning process. The Lagrange constraint is implemented by introducing the output value of the safety-critic network into the actor network in the DDPG network.
[0023] The actor network, critic network, and safety-critic network of DDPG all adopt a fully connected structure, including input layer, hidden layer, and output layer. The actor network uses the state of the robot environment and the expected target. As input, it outputs the action vector, which is used to form the regularization term loss through the two-norm regularization. The safety-critic network takes the robot's environmental state and action as input, outputs the cost value, and forms the cost value term loss with the cost value. The critic network takes the robot's environmental state, action and expected goal as input. As input, the action value is output, and it is combined with the constant -1 to form the action value loss term; finally, the Adam optimizer is used to update the actor network parameters according to the gradients calculated by actor network loss 1, actor network loss 2, and actor network loss 3, so as to control the robotic arm environment and complete the visual servoing task as quickly as possible while ensuring that the QR code does not exceed the camera screen.
[0024] Furthermore, the method for detecting the moving feature points of the QR code is as follows:
[0025] Assume that the state space of the robot environment contains four moving feature points. The pixel coordinates of the four moving feature points are obtained by detecting the moving feature point detection function cv2.aruco.detectMarkers() in OpenCV.
[0026] The pixel coordinates of the four moving feature points of the QR code are expressed as an 8-dimensional target space: ,in For the The pixel coordinates of the feature points, the target space represents the moving range space of the four moving feature points of the QR code detected in the picture, and the target reached at time t in the target condition reinforcement learning and expected goals are all in the target space. Designed for the current moment in the robotic arm environment The pixel coordinates of the four moving feature points of the QR code, The pixel coordinates of the four fixed target points in the camera image that are assigned to the positions of the moving feature points during environment initialization.
[0027] Furthermore, the robotic arm environment uses a seven-axis robotic arm. The states and actions of the robotic arm environment are as follows:
[0028] The state of the robot arm environment is represented by a state space, including the position and velocity of the robot arm end, the gripper width, the seven joint angles and velocities, and the range space of the target reached. The robot arm state in this range space serves as the input of the actor network, the critic network, and the safety-critic network to complete the training of the safety reinforcement learning model. The robot arm state space is a 29-dimensional vector, represented as:
[0029] ;
[0030] in, Represents the end information of the seven-axis robot arm: end position, end speed, end gripper width, a total of 7 dimensions; Represents the joint angles of the seven-axis robotic arm, a total of 7 dimensions; Indicates the joint speed of the seven-axis robotic arm, a total of 7 dimensions; Represents the pixel coordinates of the four moving feature points of the QR code, a total of 8 dimensions;
[0031] The actions of the robotic arm environment are identified in the action space, which is the range of motion speeds of the seven joints and one gripper of the robotic arm. In the action space of the robotic arm environment, the actions are used as inputs to the actor network, critic network, and safety-critic network to complete the training of the safety reinforcement learning model. The action space of the robotic arm environment is an 8-dimensional vector, expressed as:
[0032] ;
[0033] This includes the target speed of each joint of the robotic arm , target speed of end gripper movement .
[0034] Furthermore, the safety reinforcement learning model generates rewards from the historical interaction data of the robot arm and the environment through a reward function, which is a sparse reward function:
[0035] ;
[0036] In the above formula ,in for The goal of the moment has been reached, for The desired goal at the moment, That is, the two-norm distance between the two, is a hyperparameter distance threshold. Only when the distance between the feature point and the target point in the picture is less than the threshold , the reward will be 0, otherwise the reward will be -1; this kind of reward is sparse, and the HER algorithm is used to preprocess the data to alleviate the sparse reward and increase the sample efficiency;
[0037] The safety-critic network inputs the state and action of the robot arm and evaluates the degree of danger of the robot arm's action, where the cost function is:
[0038] ;
[0039] Where, for A flag variable indicating whether the environment is terminated at the moment. When the environment is terminated, that is, the QR code on the object surface exceeds the camera screen, the robot arm is given an environmental cost of 0 to punish it for taking an action that terminates the environment at that moment. Otherwise, a cost of -1 is given.
[0040] Furthermore, the specific process of HER operation and standardization operation is as follows:
[0041] (1) HER operation:
[0042] 1) Select a trajectory data generated by the interaction between the robot arm and the environment from the playback buffer. The trajectory data is the interaction data from t=0 to the end of the interaction with the environment, which contains multiple moments.
[0043] 2) Randomly select a moment in the trajectory data and At the moment, the interaction data at the corresponding moment are and ,in Indicates the status of the robot arm. Indicates action, represents the reward function, Indicates desired goal, Indicates that the goal has been achieved;
[0044] 3) In the time data Depend on In the time data Instead, that is ;
[0045] 4) Based on the reward function Recalculate Rewards of the moment ;
[0046] 5) The data after the HER operation is ;
[0047] (2) Standardized operations:
[0048] Robotic arm status Contains different types of information such as the robot's three-dimensional position, speed, gripper width, and pixel coordinates. The following formula is used to calculate the state and expected goals Perform standardized operations and set the robot arm status and expected goals To splice:
[0049] , ;
[0050] Where, is the state of the robotic arm, For the desired goal, are the mean and standard deviation of the robot arm state, is the mean and standard deviation of the desired target, and the splicing operation is in front, Behind.
[0051] Furthermore, the training process of the secure reinforcement learning model is as follows:
[0052] Safety-critic network utilization cost information Construct an assessment of the riskiness of an action and introduce the target strategy after cost:
[0053] ;
[0054] in are the time discount factor and cost threshold of the reward, They are the policy state, action, reward function, cost, and the distribution of state and action generated by the policy; the meaning of the target policy is that the QR code remains in the camera image when the safety constraint is met, that is, the environment step exceeds the corresponding threshold , the expected cumulative reward of the robotic arm is maximized; the loss function of the safety-critic network is as follows:
[0055] ;
[0056] The data Sampling in playback buffer , which is composed of the robot's current state, action, cost, expected goal and state at the next moment. is the time discount factor of the cost, strategy It is a deterministic policy for the DDPG network, which takes the state and the desired goal as input and outputs a deterministic action; and are the target network parameters and network parameters of the safety-critic network, Updated through soft updates; the output of the safety-critic network is the cost value Its input is the state and action, which represents the assessment of the dangerousness of the action, and completes the dangerousness assessment of the robot arm's action causing the QR code to go beyond the screen; the cost value output by the safety-critic network is added to the actor network update of the DDPG network to ensure that the actor network update takes into account the dangerousness of the action.
[0057] Furthermore, based on the critic network and actor network update method in the DDPG network, a regularization term and a safety-critic network are added to the actor network. The specific contents are as follows:
[0058] (1) Critic Network Update:
[0059] ;
[0060] in is the reward at the current moment, and are the parameters of the critic network and the critic target network respectively. The output of the critic network is the action value , whose input is state, action and desired goal, which represents the estimation of action value;
[0061] (2) Update of actor network. The actor network update function in the original DDPG only contains the action value term , modify and add action regular items , cost value item , to ensure that updates to the actor network take into account the magnitude and dangerousness of actions, as follows:
[0062] ;
[0063] in are the actor network parameters, For actor networks with state and expected goals Deterministic actions generated for input, is the two-norm regularization term for generating actions for the actor network, They are the weight parameters of the action regularization term and the cost-value term respectively. The actor network loss function consists of three parts: the action value term output by the critic network, the action regularization term, and the cost-value term output by the safety-critic network. The action value term ensures that the update of the actor network parameters is in the direction of increasing the action value, the action regularization term ensures that the amplitude of the action output by the actor network is reduced, and the cost-value term ensures that the update of the actor network parameters is in the direction of reducing the cost value.
[0064] A computer device includes a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute a method for realizing visual servo control of a robotic arm based on target condition reinforcement learning.
[0065] Beneficial Effects: This invention effectively improves the flexibility and adaptability of the robotic arm in complex and changing mission scenarios and environments. When the position and posture of the target object undergo significant changes and when a large number of environmental interference factors are present, target-condition reinforcement learning is used to dynamically generate control strategies based on real-time visual information and mission objectives, improving the flexibility and adaptability of visual servo control while maintaining effective control of the robotic arm. This invention can effectively reduce development costs, manual design, and debugging costs, and has broad practical application. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 1 is a flow chart of a method for controlling the movement of a robotic arm using a visual servo control method of a robotic arm based on target condition reinforcement learning according to the present invention;
[0067] Figure 2 is a model training flow chart of a manipulator visual servo control method based on target condition reinforcement learning in an embodiment of the present invention;
[0068] Figure 3 This is a model structure diagram of the actor network update in the DDPG network used in the present invention;
[0069] Figure 4 is a motion graph of four moving feature points captured by a camera when the robotic arm moves in an embodiment of the present invention;
[0070] Figure 5 is a graph showing the angles of each joint when the robot arm moves in an embodiment of the present invention;
[0071] Figure 6 2 is a schematic diagram of the structure of a robotic arm in an embodiment of the present invention;
[0072] Figure 7 is a motion graph of four moving feature points captured by a camera when the robotic arm is subject to environmental interference in an embodiment of the present invention;
[0073] Figure 8 This is a curve diagram of the angles of each joint of the robotic arm when the robotic arm is subject to environmental interference in an embodiment of the present invention. DETAILED DESCRIPTION
[0074] The present invention will be further explained below with reference to the accompanying drawings.
[0075] like Figure 1As shown, the present invention discloses a method for visual servo control of a robotic arm based on target-conditional reinforcement learning: the traditional visual servo task is designed as a target-conditional reinforcement learning task, that is, the robotic arm is trained to move the feature point in the camera (the camera is fixed at the end of the robotic arm) to the fixed target point, wherein the pixel coordinates of the feature point and the fixed target point are respectively set as the achieved goal and the desired goal, and the Deep Deterministic Policy Gradient (DDPG) algorithm and the Hindsight Experience Replay (HER) algorithm are used for online learning. In order to avoid the problem of low learning efficiency caused by feature points exceeding the camera field of view during training, regularization terms and Lagrangian constraints are introduced to restrict the action to ensure the safety of the training process. At the same time, a visual servo control simulation platform was built with the robotic arm as the research object to verify the algorithm, as shown in the attached figure. Figure 6 As shown in [1], the algorithm uses the positions of feature points in the camera image as input to control the movement of the robotic arm and lock onto a QR code attached to the surface of the object, completing the visual servoing task. Results show that the proposed method significantly improves the number of iterations and robustness compared to traditional pseudo-inverse control methods.
[0076] S1. Set any four pixels with known coordinates in the image captured by the robot camera as fixed target points;
[0077] S2. Set a QR code on the object as the recognition target for the robot arm visual servo task. The recognition target is fixed in position after initialization.
[0078] S3: The robotic arm obtains four moving feature points by shooting the QR code with a camera. The robotic arm is controlled to move the moving feature points to the pixel coordinates of four fixed target points in the shooting image, thereby controlling the robotic arm to achieve the visual servoing task.
[0079] The method for detecting the moving feature points of a QR code is as follows: Assume that the state space of the QR code in the robotic arm environment contains four moving feature points. The pixel coordinates of the four moving feature points are obtained using the open-source OpenCV moving feature point detection function cv2.aruco.detectMarkers(). During training, the robotic arm's movements must be restricted to ensure that the four moving feature points of the QR code in the camera image are always clear. If the QR code is incomplete or not fully displayed, the moving feature point detection function will automatically report an error and the current training process will be forced to terminate.
[0080] The pixel coordinates of the four moving feature points of the QR code are expressed as an 8-dimensional target space goal: ,in For the The pixel coordinates of the feature points, the target space represents the moving range space of the four moving feature points of the QR code detected in the picture, and the target reached at time t in the target condition reinforcement learning and expected goals are all in the target space. Designed for the current moment in the robotic arm environment The pixel coordinates of the four moving feature points of the QR code, The pixel coordinates of the four fixed target points in the camera image that are assigned to the positions of the moving feature points during environment initialization.
[0081] The robotic arm environment uses a seven-axis robotic arm. The state of the robotic arm environment is represented by a state space, including the position and speed of the robotic arm end, the gripper width, the seven joint angles and speeds, and the range space of the target reached. The robotic arm state in this range space serves as the input of the actor network, the critic network, and the safety-critic network to complete the training of the safety reinforcement learning model. The robotic arm state space is a 29-dimensional vector, represented as:
[0082] ;
[0083] in, Represents the end information of the seven-axis robot arm: end position, end speed, end gripper width, a total of 7 dimensions; Represents the joint angles of the seven-axis robotic arm, a total of 7 dimensions; Indicates the joint speed of the seven-axis robotic arm, a total of 7 dimensions; Represents the pixel coordinates of the four moving feature points of the QR code, a total of 8 dimensions;
[0084] The actions of the robotic arm environment are identified in the action space, which is the range of motion speeds of the seven joints and one gripper of the robotic arm. In the action space of the robotic arm environment, the actions are used as inputs to the actor network, critic network, and safety-critic network to complete the training of the safety reinforcement learning model. The action space of the robotic arm environment is an 8-dimensional vector, expressed as:
[0085] ;
[0086] This includes the target speed of each joint of the robotic arm , target speed of end gripper movement ;
[0087] S4, setting the pixel coordinates of the moving feature point and the fixed target point as the reached target ag and the desired target dg in the target condition reinforcement learning respectively;
[0088] The safety reinforcement learning model generates rewards for the historical interaction data of the robot arm and the environment through a reward function, which is a sparse reward function.
[0089] ;
[0090] In the above formula ,in for The goal of the moment has been reached, for The desired goal at the moment, That is, the two-norm distance between the two, is a hyperparameter distance threshold. Only when the distance between the feature point and the target point in the picture is less than the threshold , the reward will be 0, otherwise the reward will be -1; this kind of reward is sparse, and the HER algorithm is used to preprocess the data to alleviate the sparse reward and increase the sample efficiency;
[0091] The safety-critic network inputs the state and action of the robot arm and evaluates the degree of danger of the robot arm's action, where the cost function is:
[0092] ;
[0093] Where, for A flag variable indicating whether the environment is terminated at the moment. When the environment is terminated, that is, the QR code on the object surface exceeds the camera screen, the robot arm is given an environmental cost of 0 to punish it for taking an action that causes the environment to terminate at that moment. Otherwise, a cost of -1 is given.
[0094] S5. Build a safety reinforcement learning model. The safety reinforcement learning model includes a deep deterministic policy gradient algorithm (DDPG) network, Lagrange constraints, a hindsight experience replay algorithm (HER), and a safety critic network (safety-critic). The DDPG algorithm includes an actor network and a critic network.
[0095] The safety reinforcement learning model preprocesses the robot arm's environment historical interaction data (actions, states, rewards, achieved goals, and desired goals) through HER and standardization operations, reconstructing reward data to alleviate the reward sparsity problem of the robot arm's environment. The safety-critic network inputs the robot arm's environment historical interaction cost data stored in the cache, and the critic network in DDPG inputs the robot arm's environment historical interaction reward data in the cache. The regularization term and Lagrange constraint term in the actor network in DDPG are used to restrict the robot arm's movement to ensure the safety of the learning process. The Lagrange constraint is implemented by introducing the output value of the safety-critic network into the actor network in the DDPG network.
[0096] The actor network, critic network and safety-critic network in the DDPG network all adopt a fully connected structure, including input layer, hidden layer and output layer. The hidden layer of the actor network, critic network and safety-critic network all adopt a 256*4 structure. (29 dimensions), desired goal (8-dimensional) as input, output 8-dimensional action vector, the action vector is weighted by L2 regularization and regularization term The first term that constitutes the actor network loss , the safety-critic network is based on the state (29 dimensions), action (8 dimensions) as input, output cost value (1 dimension), with cost value and corresponding cost weight Constitutes the second term of actor network loss , the critic network is based on the state (29 dimensions), action (8 dimensions) and desired goals (8 dimensions) as input, output action value, and together with the constant -1 constitute the third term of the actor network loss Finally, the Adam optimizer is used to update the actor network parameters according to the gradient of the loss calculation, and the robot arm is controlled to complete the visual servo task as quickly as possible while ensuring that the QR code does not exceed the screen. Figure 3 As shown;
[0097] The training process of the safety reinforcement learning model is as follows: the safety-critic network utilizes cost information Construct an assessment of the riskiness of an action and introduce the target strategy after cost:
[0098] ;
[0099] in are the time discount factor and cost threshold of the reward, They are strategy, state, action, reward function, cost, and the distribution of state and action generated by the strategy; the meaning of the target strategy is to meet the safety constraints, that is, the environment step of the QR code remaining in the camera screen exceeds the corresponding threshold , the expected cumulative reward of the robotic arm is maximized; the loss function of the safety-critic network is as follows:
[0100] ;
[0101] The data Sampling in playback buffer , which is composed of the robot's current state, action, cost, expected goal and state at the next moment. is the time discount factor of the cost, strategy It is a deterministic policy for the DDPG network, which takes the state and the desired goal as input and outputs a deterministic action; and are the target network parameters and network parameters of the safety-critic network, Updated through soft updates; the output of the safety-critic network is the cost value , whose input is state and action, representing the assessment of the dangerousness of the action, and completing the dangerousness assessment of the robot arm's action causing the QR code to go beyond the screen; the cost value output by the safety-critic network is added to the actor network update of the DDPG network to ensure that the actor network update takes into account the dangerousness of the action;
[0102] Based on the critic network and actor network update method in the DDPG network, a regularization term and a safety-critic network are added to the actor network. The specific contents are as follows:
[0103] (1) Critic Network Update:
[0104] ;
[0105] in is the reward at the current moment, and are the parameters of the critic network and the critic target network respectively. The output of the critic network is the action value , whose input is state, action and desired goal, representing the estimation of action value;
[0106] (2) Update of actor network. The actor network update function in the original DDPG only contains the action value term , modify and add action regular items , cost value items To ensure that updates to the actor network take into account the magnitude and dangerousness of actions, we modify it as follows:
[0107] ;
[0108] in are the actor network parameters, For actor networks with state and expected goals Deterministic actions generated for input, is the two-norm regularization term for generating actions for the actor network, are the weight parameters of the action regularization term and the cost-value term, respectively. The actor network loss function consists of three parts: the action value term output by the critic network, the action regularization term, and the cost-value term output by the safety-critic network. The action value term ensures that the actor network parameters are updated in the direction of increasing the action value, the action regularization term ensures that the amplitude of the action output by the actor network is reduced, and the cost-value term ensures that the actor network parameters are updated in the direction of reducing the cost value.
[0109] S6. After the safety reinforcement learning model is trained, the current state and desired goal are directly input into the actor network. The safety reinforcement learning model directly outputs the target speed action of the robot arm joint, controls the robot arm to lock onto the target with the QR code, and completes the visual servo control task.
[0110] A computer-readable storage medium stores a computer program, which is suitable for being loaded by a processor and executing the target condition reinforcement learning-based manipulator visual servo control method.
[0111] A computer device includes a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute a robot arm visual servo control method based on target condition reinforcement learning.
[0112] like Figure 2 As shown, the training process of the robot arm visual servo control method based on target condition reinforcement learning in this embodiment includes the following steps:
[0113] Step 1: Build a visual servo robotic arm environment, which includes a robotic arm;
[0114] Step 2: Data collection or importing saved data. During this data collection phase, the robot interacts with the environment for 1000 episodes. The specific steps for each episode are as follows:
[0115] Step 2.1: Determine whether the environment is terminated or expired. If so, exit the episode; otherwise, continue.
[0116] Step 2.2: Randomly generate actions to interact with the environment for one step;
[0117] Step 2.3: Store data (action, state, reward, cost, whether it is terminated) in the replay buffer;
[0118] Step 2.4: Return to step 2.1;
[0119] Step 3: Training the safety reinforcement learning model of this method; during the training phase, the robot arm interacts with the environment for 20,000 episodes. The specific steps are as follows:
[0120] Step 3.1: Determine whether the number of iterations in the training phase exceeds the maximum number of iterations. If so, the training ends and the trained model parameters are obtained. Otherwise, continue;
[0121] Step 3.2: The robot completes one episode of interaction with the environment, i.e., steps 2.1-2.4, except that the actions in step 2.2 are generated by the robot based on the current actor network in the DDPG network.
[0122] Step 3.3: Randomly sample actions, states, rewards, and costs from the stored data and perform data preprocessing, which includes HER operations and normalization operations;
[0123] Step 3.4: Safety-critic network update;
[0124] Step 3.5: Critic network update;
[0125] Step 3.6: Actor network update;
[0126] Step 3.7: If the number of iterations is a multiple of 100, evaluate the actor network. If the number of iterations is a multiple of 1000, save all network parameters. Otherwise, continue.
[0127] Step 3.8: Return to step 3.1.
[0128] In step 2, whether to use the buffer pool variable use_buffer determines whether to enter the data collection phase at the beginning of training. The purpose of this operation is to avoid the time-consuming operation of performing data collection during multiple training sessions and to locally save the data collected in a certain training session for use in subsequent training sessions.
[0129] In step 2.2, we impose certain limits on the randomly generated actions to ensure that the length of the collected trajectory is not too short (the environment does not terminate too early). The limiting operations are as follows:
[0130] ;
[0131] In the formula The range is The uniform distribution of Randomly generated actions.
[0132] After performing the HER operation on the sampled data in step 3.3, the mean and variance of the state and the expected target need to be updated. The update formula is as follows:
[0133] ;
[0134] in are the expectation of the random variable, the variance of the random variable, Data , very small amount (take ), are the updated data mean and variance.
[0135] like Figure 4 、 Figure 5 and Figure 6 As shown in the figure, the effect of using the trained safety reinforcement learning model to control the robotic arm to complete the task in an offline state is shown. Figure 4 The moving feature points in the blue box are the original positions of the two-dimensional moving feature points in the picture. The red dots are the real-time positions of the four moving feature points of the QR code during the movement process. The green dots are the fixed target pixel points in the picture. Figure 5 The end position of the robot arm converges to a stable state within 10 steps, and the number of simulation steps also corresponds to Figure 3 The number of red points in . Figure 6 This is the motion diagram of the environmental robotic arm, where the red sphere represents the real-time position of the camera during the motion. Figure 6-8 This is the effect of adding interference to the robot arm in the environment when it completes the task and enters the locked state, where the interference is to add an offset to the QR code position. , and add an offset to the pose .
[0136] Figure 7The blue box in the figure represents the position of the QR code in the picture after the QR code position and posture are disturbed, corresponding to Figure 8 At the moment when the simulation step in the image is equal to 50 (the moment when interference is added), the safety reinforcement learning model controls the robot arm again to complete the visual servo task, so that the four corners of the QR code in the image move to the target position. Figure 7 The red dots in the image are the real-time detection points of the QR code corner points moving in the picture. The number of red dots is less than 5 (for a single corner point), corresponding to Figure 8 The number of simulation steps required for the joint angle of the manipulator to lock the target again after the interference is less than 5, which reflects the anti-interference ability of this scheme in controlling the manipulator to complete the visual servoing task.
Claims
1. A visual servo control method for a robotic arm based on target-condition reinforcement learning, wherein a camera is set at the end of the robotic arm and a robotic arm environment is constructed, characterized by: S1. Set four pixel points with known coordinates in the image captured by the robot arm camera as fixed target points; S2. Set a QR code on the object as the recognition target for the robot arm visual servo task. The recognition target is fixed in position after initialization. S3: The robotic arm obtains four moving feature points by photographing the QR code with a camera. Using the secure reinforcement learning model, the pixel coordinates of the moving feature points and the fixed target point are set as the achieved target and the desired target in the target-conditional reinforcement learning, respectively. S4. The secure reinforcement learning model includes a deep deterministic policy gradient algorithm, Lagrangian constraints, a hindsight experience replay algorithm, and a secure critic network. The deep deterministic policy gradient algorithm includes an actor network and a critic network. S5. After the safety reinforcement learning model is trained using the historical interaction data of the robotic arm environment, the current state and desired goal are directly input into the actor network in the deep deterministic policy gradient algorithm. The safety reinforcement learning model can then directly output the target speed movement of the robotic arm joint through the information of the identified target captured by the camera, control the robotic arm to lock onto the target with the QR code, and control the movement of the robotic arm so that the mobile feature point moves to the pixel coordinates of the four fixed target points in the captured image, thus completing the visual servo control task.
2. The method for manipulator visual servo control based on target condition reinforcement learning according to claim 1, characterized in that: Before the robotic arm camera shoots a fixed target point, you need to determine the camera image direction as follows: The camera's screen direction is the vector in the camera coordinate system that starts from the center of the shooting screen and points to the top of the screen. The camera's screen direction is always fixed in the camera coordinate system and is set to ; The direction of the camera image in the world coordinate system will change with the movement of the robot arm. By calculating the representation of the camera image direction in the world coordinate system To ensure the correct refresh of the camera image; obtain the relative position information of each joint of the robotic arm, vector The calculation process is as follows: ; In the above formula Represents the coordinate system around Axis rotation The rotation matrix of the angle, Represents the coordinate system around Axis rotation The rotation matrix of the angle, Represents the coordinate system around Axis rotation The rotation matrix of the angle; ; The above formula is the rotation matrix between the coordinate systems of each robot joint, where For joints Coordinate system relative to joint The rotation matrix of the coordinate system, For joints The rotation angle, Represent the world coordinate system, the robot arm end gripper coordinate system and the camera coordinate system respectively; Representation of the camera screen direction in the world coordinate system as follows: ; Where, is the rotation matrix of the camera coordinate system relative to the gripper coordinate system at the end of the robotic arm, Represents the rotation matrix of the manipulator end gripper coordinate system relative to the manipulator joint 8 coordinate system, Indicates the camera image direction in the camera coordinate system.
3. The method for manipulator visual servo control based on target condition reinforcement learning according to claim 1, characterized in that: The safety reinforcement learning model preprocesses the robot arm's environment historical interaction data: actions, states, rewards, achieved goals, and expected goals, using a post-experience replay algorithm and standardization operations to reconstruct the reward data to alleviate the reward sparsity problem of the robot arm environment. A safety critic network is constructed using a cost function, and the safety critic network is fed with the robot arm environment historical interaction cost data stored in the replay buffer. The critic network in the deep deterministic policy gradient algorithm is fed with the robot arm environment historical interaction reward data in the replay buffer. The regularization term and Lagrangian constraint term in the actor network in the deep deterministic policy gradient algorithm are used to restrict the robot arm's movement to ensure the safety of the learning process. The Lagrangian constraint is implemented by introducing the output value of the safety critic network into the actor network in the deep deterministic policy gradient algorithm. Among them, the actor network, critic network and safety critic network of the deep deterministic policy gradient algorithm all adopt a fully connected structure, including an input layer, a hidden layer and an output layer; the actor network takes the state of the robot arm environment and the expected goal as input, and outputs an action vector. The action vector constitutes a regularization term loss through the two-norm regularization. The safety critic network takes the robot arm environment state and action as input, outputs a cost value, and the cost value constitutes a cost-value term loss. The critic network takes the robot arm environment state, action and expected goal as input, outputs an action value, and constitutes an action-value term loss with a constant -1; finally, the adaptive moment estimation optimizer is used to update the actor network parameters according to the gradient calculated based on the regularization term loss, cost-value term loss and action-value term loss, so as to control the robot arm environment and complete the visual servoing task as quickly as possible while ensuring that the QR code does not exceed the camera screen.
4. The method for manipulator visual servo control based on target condition reinforcement learning according to claim 3, characterized in that: The method for detecting the moving feature points of a QR code is as follows: Assume that the state space of the robot environment contains four moving feature points. The pixel coordinates of the four moving feature points are obtained by detecting the moving feature point detection function cv2.aruco.detectMarkers() in OpenCV. The pixel coordinates of the four moving feature points of the QR code are expressed as an 8-dimensional target space goal: ,in For the The pixel coordinates of the feature points, the target space represents the moving range space of the four moving feature points of the QR code detected in the picture, and the target reached at time t in the target condition reinforcement learning and expected goals are all in the target space. Designed for the current moment in the robotic arm environment The pixel coordinates of the four moving feature points of the QR code, The pixel coordinates of the four fixed target points in the camera image that are assigned to the positions of the moving feature points during environment initialization.
5. The method for manipulator visual servo control based on target condition reinforcement learning according to claim 4, characterized in that: The robotic arm environment uses a seven-axis robotic arm. The states and actions of the robotic arm environment are as follows: The state of the robot arm environment is represented by a state space, including the position and velocity of the robot arm end, the gripper width, the seven joint angles and velocities, and the range space of the target reached. The robot arm state in this range space serves as the input of the actor network, the critic network, and the safety critic network to complete the training of the safety reinforcement learning model. The robot arm state space is a 29-dimensional vector, represented as: ; in, Represents the end information of the seven-axis robot arm: end position, end speed, end gripper width, a total of 7 dimensions; Represents the joint angles of the seven-axis robotic arm, a total of 7 dimensions; Indicates the joint speed of the seven-axis robotic arm, a total of 7 dimensions; Represents the pixel coordinates of the four moving feature points of the QR code, a total of 8 dimensions; The actions of the robotic arm environment are identified in the action space, which is the range of motion speeds of the seven joints and one gripper of the robotic arm. In this action space, the actions are used as inputs to the actor network, critic network, and safety critic network to complete the training of the safety reinforcement learning model. The action space of the robotic arm environment is an 8-dimensional vector, expressed as: ; This includes the target speed of each joint of the robotic arm , target speed of end gripper movement .
6. The method for manipulator visual servo control based on target condition reinforcement learning according to claim 1, characterized in that: The safety reinforcement learning model generates rewards for the robot arm's historical interaction data with the environment through a reward function, which is a sparse reward function: ; In the above formula ,in for The goal of the moment has been reached, for The desired goal at the moment, That is, the two-norm distance between the two, is a hyperparameter distance threshold. Only when the distance between the feature point and the target point in the picture is less than the threshold , you will get a reward of 0, otherwise you will get a reward of -1; This kind of reward is sparse, and the HER algorithm is used to preprocess the data to alleviate the sparse reward and increase sample efficiency; The safety critic network inputs the state and action of the robot arm and evaluates the dangerousness of the robot arm's action, where the cost function is: ; Where, for A flag variable indicating whether the environment is terminated at the moment. When the environment is terminated, that is, the QR code on the object surface exceeds the camera screen, the robot arm is given an environmental cost of 0 to punish it for taking an action that terminates the environment at that moment. Otherwise, a cost of -1 is given.
7. The method for manipulator visual servo control based on target condition reinforcement learning according to claim 6, characterized in that: The specific process of post-experience replay algorithm operation and standardization operation is as follows: (1) Post-experience replay algorithm operation: 1) Select a trajectory data generated by the interaction between the robot arm and the environment from the playback buffer. The trajectory data is the interaction data from t=0 to the end of the interaction with the environment, which contains multiple moments. 2) Randomly select a moment in the trajectory data and At the moment, the interaction data at the corresponding moment are and ,in Indicates the status of the robot arm. Indicates action, represents the reward function, Indicates desired goal, Indicates that the goal has been achieved; 3) In the time data Depend on In the time data Instead, that is ; 4) Based on the reward function Recalculate Rewards of the moment ; 5) The data after the post-experience replay algorithm operation is ; (2) Standardized operations: The robot state contains different types of information such as the robot's 3D position, speed, gripper width, and pixel coordinates. The following formula is used to normalize the state and the desired target and then concatenate the robot state with the desired target: , ; Where, is the state of the robotic arm, For the desired goal, are the mean and standard deviation of the robot arm state, are the mean and standard deviation of the expected target, and the splicing operation is state first and expected target second.
8. The method for manipulator visual servo control based on target condition reinforcement learning according to claim 6, characterized in that: The training process of the secure reinforcement learning model is as follows: Security Critic Network Exploitation Cost Information Construct an assessment of the riskiness of an action and introduce the target strategy after cost: ; in are the time discount factor and cost threshold of the reward, They are strategy, state, action, reward function, cost, and the distribution of state and action generated by the strategy; the meaning of the target strategy is to meet the safety constraints, that is, the environment step of the QR code remaining in the camera screen exceeds the corresponding threshold , the expected cumulative reward of the robotic arm is maximized; the loss function of the security critic network is as follows: ; The data Sampling in playback buffer , which is composed of the robot's current state, action, cost, expected goal and state at the next moment. is the time discount factor of the cost, strategy It is a deterministic policy for the deep deterministic policy gradient algorithm, which takes the state and the desired goal as input and outputs a deterministic action; and are the target network parameters and network parameters of the security critic network, Updated through soft updates; the output of the security critic network is cost value Its input is state and action, which represents the evaluation of the dangerousness of the action, and completes the dangerousness evaluation of the robot arm action causing the QR code to go out of the screen; the cost value output by the safety critic network is added to the actor network update of the deep deterministic policy gradient algorithm to ensure that the update of the actor network takes into account the dangerousness of the action.
9. The method for manipulator visual servo control based on target condition reinforcement learning according to claim 8, characterized in that: Based on the critic network and actor network update method in the deep deterministic policy gradient algorithm, a regularization term and a safety critic network are added to the actor network. The details are as follows: (1) Reviewer Network Update: ; in is the reward at the current moment, and are the parameters of the critic network and the critic target network respectively. The output of the critic network is the action value , whose input is state, action and desired goal, representing the estimation of action value; (2) Update of the actor network. The actor network update function in the original deep deterministic policy gradient algorithm only contains the action value term , modify and add action regular items , cost value items , to ensure that the update of the actor network takes into account the magnitude and dangerousness of the action, specifically expressed as follows: ; in are the actor network parameters, For actor network with status and expected goals Deterministic actions generated for input, is the two-norm regularization term for generating actions for the actor network, are the weight parameters of the action regularization term and the cost value term respectively. The actor network loss function consists of three parts: the action value term output by the critic network, the action regularization term, and the cost value term output by the security critic network. The action value term ensures that the update of the actor network parameters is in the direction of increasing the action value, the action regularization term ensures that the amplitude of the action output by the actor network is reduced, and the cost value term ensures that the update of the actor network parameters is in the direction of reducing the cost value.
10. A computer device, characterized in that: It includes a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the robot arm visual servo control method based on target condition reinforcement learning as described in any one of claims 1-9.
Citation Information
Patent Citations
Visual servo control method for four-degree-of-freedom mechanical arm
CN113146616A
Monocular vision servo screw hole positioning method and system
CN115890657A