A robot visual servo motion control method based on deep reinforcement learning

By using a deep reinforcement learning hybrid visual servo controller trained in a virtual environment, the problems of stability and motion performance in robot visual servo assembly and positioning are solved, achieving stable execution of visual servo tasks and improving robot motion performance.

CN117021066BActive Publication Date: 2025-10-28ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310621091.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-26
Publication Date
2025-10-28
Estimated Expiration
2043-05-26

AI Technical Summary

Technical Problem

Existing robot vision servoing assembly and positioning methods struggle to balance stability and motion performance when handling complex tasks, especially in effectively handling constraints during vision servoing. Furthermore, traditional Q-learning methods are not applicable to situations where the state space or action space is continuous and high-dimensional.

Method used

We designed a hybrid visual servo controller based on deep reinforcement learning. By training the hybrid visual servo controller in a virtual environment, combining the advantages of PBVS and IBVS, we optimized the robot's motion strategy using a deep reinforcement learning model, constructed a hybrid visual servo model, and deployed the controller in a real environment.

Benefits of technology

It improves the stability of visual servoing tasks and the motion performance of robots, avoids losses in real-world environments, and ensures the effective execution of visual servoing tasks under image and physical constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117021066B_ABST
    Figure CN117021066B_ABST
Patent Text Reader

Abstract

This invention discloses a robot visual servo motion control method based on deep reinforcement learning. It includes the following steps: First, determining the robot's visual servo assembly and positioning task, along with the corresponding optimization objectives and constraints; next, constructing a hybrid visual servo controller based on deep reinforcement learning; then, training the hybrid visual servo controller in a virtual environment to obtain a trained hybrid visual servo controller, which is then deployed to a real environment to control the robot to perform the actual assembly and positioning task. This invention utilizes a virtual twin environment and deep reinforcement learning to perform offline training of the hybrid visual servo controller, ensuring the safety of the training process, avoiding unnecessary damage to the real robot, and allowing the trained controller to be directly deployed to real-world work scenarios. This achieves improved robot motion performance while ensuring the stability of the robot's visual servo task, demonstrating significant engineering practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a robot servo motion control method in the field of industrial robot motion control, and in particular to a motion planning and control method for the vision servo assembly and positioning process of industrial robots. Background Technology

[0002] Robot assembly localization is a crucial step in autonomous assembly operations, involving moving the parts to be assembled to their appropriate positions and aligning them with reference parts. When planning robot assembly localization tasks, the robot's motion path points and corresponding motion commands are typically determined based on the pose data of the reference parts in the work environment. However, in practical applications, the position of the reference parts may deviate from the planned position, leading to situations where the robot, after reaching the assembly localization point according to the taught program, is unable to complete subsequent assembly tasks. To overcome these shortcomings, researchers have begun using visual servo control technology to achieve dynamic assembly localization for robots, achieving good results in some application cases.

[0003] With the assistance of visual servoing control technology, robots possess greater autonomy and are capable of undertaking more flexible assembly tasks. However, a series of constraints in the visual servoing process can affect the stability and convergence of the task. Generally speaking, constraints in visual servoing fall into two main categories: image / camera constraints and robot / physical constraints. Image / camera constraints are mainly caused by the limitations of the vision system, including camera field of view constraints and image Jacobi singularities. Robot / physical constraints arise from the robot and physical space, including robot joint constraints, kinematic dynamics constraints, and collision avoidance constraints. To improve the stability of visual servoing, many researchers have introduced path planning methods to improve the visual servoing process, enabling robots to cope with various constraints and achieve the expected performance. These methods can be broadly classified into the following categories: 1) image space-based path planning, 2) global path planning, and 3) optimization-based path planning. These path planning methods for visual servoing address the constraint problems from different perspectives, improving the robustness of the visual servoing process. However, these methods still struggle to handle complex tasks with numerous constraints, and they rarely consider the kinematic and dynamic performance of the robot during operation.

[0004] In recent years, reinforcement learning has gradually emerged and been applied to robotics to achieve complex tasks, such as autonomous grasping, path tracking, and precision assembly of shafts and holes. Simultaneously, researchers have begun to explore its application in visual servoing. Compared to planning-based methods that rely on model analysis, learning-based methods automatically explore an optimized robot motion strategy through trial and error with the environment. Currently, most literature uses Q-learning for motion planning in robot visual servoing, such as camera field-of-view constraint control and adaptive servo gain adjustment. However, Q-learning can only handle problems in discrete spaces and is not suitable for situations where the state space or action space is continuous and high-dimensional. Therefore, using Q-learning for visual servoing motion planning still has certain limitations. Summary of the Invention

[0005] To address the motion planning problem in robot visual servoing assembly and positioning, this invention designs a Deep Reinforcement Learning-based Hybrid Visual Servoing (DRL-HVS) controller and achieves motion planning and control of visual servoing through offline training in a virtual environment.

[0006] The technical solution of the present invention is as follows:

[0007] (1) Determine the robot's vision servo assembly and positioning task and the corresponding optimization objectives and constraints;

[0008] (2) Based on the visual servo assembly and positioning task and the corresponding optimization objectives and constraints, a hybrid visual servo controller based on deep reinforcement learning is constructed.

[0009] (3) Train the hybrid visual servo controller based on deep reinforcement learning in a virtual environment to obtain the trained hybrid visual servo controller.

[0010] (4) Deploy the trained hybrid vision servo controller into the real environment to control the robot to perform actual assembly and positioning tasks.

[0011] In (1), the visual servo assembly positioning task specifically includes:

[0012] Determine the robot's starting position and desired position. Under the premise of achieving the optimization objective and satisfying the constraints, use visual servoing to drive the robot to achieve assembly positioning from the starting position to the desired position. The control process of driving the robot is expressed by the following formula:

[0013]

[0014] v c=[v x ,v y ,v z ,ω x ,ω y ,ω z ] T

[0015] e = ff *

[0016] Among them, v c v represents the six-degree-of-freedom motion velocity of the camera in the camera coordinate system. x ,v y ,v z Let ω represent the velocity components of the camera along the x, y, and z axes of the camera coordinate system. x ,ω y ,ω z Let x, y, and z represent the angular velocity components of the camera along the x, y, and z axes of the camera coordinate system, respectively. Let T denote the matrix transpose, and λ be the servo gain, satisfying λ∈[0,1]. Here, e represents the pseudo-inverse of the characteristic Jacobian matrix, f is the visual feature error, and f and f' are the estimated values. * These represent the visual features extracted from the image when the camera is at its current position and the desired position, respectively.

[0017] In (1), the optimization objective is to successfully complete the assembly and positioning task under the constraints and to achieve the optimal motion performance of the robot.

[0018] In (1), the constraints include camera field of view constraints, robot joint constraints, and robot speed constraints. Specifically, the camera field of view constraint means that the target object cannot leave the camera's field of view. Specifically, the robot joint constraint means that the angle of each joint of the robot cannot exceed its limit during the movement. Specifically, the robot speed constraint means that the speed of the robot's end effector cannot exceed the preset upper limit.

[0019] Specifically, (2) refers to:

[0020] Based on the visual servoing assembly and positioning task and the corresponding optimization objectives and constraints, a hybrid visual servoing model is constructed. Based on the hybrid visual servoing model, the DDPG agent is fused to obtain a deep reinforcement learning model for hybrid visual servoing, which is used to adjust the parameters of the hybrid visual servoing model. The hybrid visual servoing model and the deep reinforcement learning model together form a hybrid visual servoing controller based on deep reinforcement learning.

[0021] The formula for the hybrid vision servoing model is as follows:

[0022]

[0023]

[0024] e H =[e 3D ,e 2D ] T

[0025]

[0026] Among them, v c ′(t) represents the corrected camera motion speed, v c (t) represents the original camera motion velocity obtained from the hybrid visual servoing model, v c (0) represents the camera motion velocity obtained by the hybrid visual servoing model at startup, e H For the mixed error matrix, e 3D Indicates the visual error in the PBVS method; e 2D This represents the visual error in the IBVS method. L represents the mixed characteristic Jacobian matrix H The estimated value, λ is the servo gain, and H represents the weight matrix. Let represent the weighted pseudo-inverse matrix of the mixed feature Jacobian matrix estimates, where n is the number of 2D feature points in the IBVS method. is the attenuation factor, and h represents the weight value of the PBVS method, satisfying h∈[0,1].

[0027] In (2), the state space, action space, and reward function of the deep reinforcement learning model for hybrid visual servoing are set as follows:

[0028] Wherein, the state space s t The formula is as follows:

[0029] s t =[e 3D ,v c ,q,d s ]

[0030] Among them, e 3D and v c Let represent the camera pose error and camera velocity, respectively; q represents the current joint angle value of the robot; d s Indicates the danger distance of a 2D feature point in the image plane;

[0031] Action space a t The formula is as follows:

[0032] a t =[λ,h]

[0033] The formula for the return function is as follows:

[0034] R succeed =τ+E+η

[0035]

[0036]

[0037]

[0038]

[0039] Among them, R succeed Let τ represent the positive reward, E represent the exercise efficiency index, η represent the energy consumption index, and R represent the stability index. failed K represents a negative reward. max Let K be the maximum number of moves in a round, and K represent the number of moves in the current round. This represents the error of the 2D feature points in the current state. q represents the initial 2D feature point error, ω is the joint weight coefficient vector, and q d q represents the desired position. init Indicates the initial position, q k This represents the position of the robot at step k, || represents taking the absolute value of the difference in positions; ReLU() represents the rectified linear function, J e This represents the maximum allowed joint jerk value for a real robot, max(|J t |) represents the robot's maximum joint jerk value during the entire task.

[0040] Specifically, (3) refers to:

[0041] (3.1) Construct a work scene in the virtual environment that is consistent with the real environment, and synchronize the visual servo assembly and positioning task, as well as the corresponding optimization objectives and constraints, to the virtual environment;

[0042] (3.2) The hybrid visual servo controller continuously interacts with the robot in the virtual environment to generate experience data. The hybrid visual servo controller is trained for several rounds during the interaction process until the cumulative reward of the deep reinforcement learning model in the hybrid visual servo controller tends to maximize, and the trained hybrid visual servo controller is obtained.

[0043] In (3.2), at each time step, for the agent of the deep reinforcement learning model, the current agent's Actor network obtains the state space s at time t. t Calculate the corresponding action space a t And send it to the hybrid visual servo model, the hybrid visual servo model according to the received motion space a tAnd hybrid error matrix-driven robotic operations in virtual environments; agents receive rewards r t And the state s at the next moment t+1 Then, the state space s of the current time step is... t Action space a t Rewards r t and the state s at the next moment t+1 The empirical data is stored in the experience pool as a set of data; the current agent randomly draws empirical data from the experience pool, calculates the gradient, and updates the parameters of the Critic network.

[0044] Specifically, (4) refers to:

[0045] In the real environment, the robot's state information is transmitted to the trained hybrid vision servo controller through a data interface. The trained hybrid vision servo controller calculates the robot's next motion speed based on the current state information and sends it to the robot, thereby driving the robot to move until the vision servo assembly and positioning task is completed.

[0046] The beneficial effects of this invention are as follows:

[0047] (1) The hybrid visual servoing model based on deep reinforcement learning can combine the advantages of PBVS and IBVS. The parameter adjustment strategy of the hybrid visual servoing model can be continuously optimized using empirical data according to the specified task, so as to ensure the stability of the visual servoing task and improve the motion performance of the robot.

[0048] (2) Using a virtual twin environment to perform offline training of the DRL-HVS controller can not only ensure the safety of the training process, but also avoid unnecessary damage to the real robot and extend its service life. Moreover, the trained controller can be directly deployed to the real operation scenario. Attached Figure Description

[0049] Figure 1 This is a schematic diagram of the overall framework of the method of the present invention;

[0050] Figure 2 This is a schematic diagram of the robot vision servoing task of the present invention;

[0051] Figure 3 This is a schematic diagram of the DRL-HVS controller of the present invention;

[0052] Figure 4 This is a schematic diagram illustrating the definition of the danger distance of the feature points in this invention;

[0053] Figure 5 These are the real and virtual experimental scenarios in the embodiments;

[0054] Figure 6This is the adaptive output of the DRL-HVS controller parameters in the embodiment when performing the task;

[0055] Figure 7 These are the feature point motion trajectories and camera speed changes of the DRL-HVS controller in the embodiment;

[0056] Figure 8 This is a schematic diagram of the hardware and software configuration of the visual servo experimental system in this embodiment;

[0057] Figure 9 This describes the deployment of the trained DRL-HVS controller into a real-world environment. (a) is the robot's initial position (i.e., starting position), (b) is the image from the robot's camera, (c) is the robot's desired position, (d) is the motion trajectory of the feature points throughout the visual servoing process, (e) is the change in the robot's joint velocity in the simulation environment, and (f) is the change in the robot's joint velocity in the real environment. Detailed Implementation

[0058] To better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific examples.

[0059] like Figure 1 As shown, the present invention includes the following steps:

[0060] (1) Determine the robot's vision servo assembly and positioning task and the corresponding optimization objectives and constraints;

[0061] In (1), the visual servo assembly positioning task is specifically as follows:

[0062] Determine the robot's starting and desired positions. Under the premise of achieving the optimization objective and satisfying constraints, design or specify several appropriate 3D points as feature points (such as the bounding box vertices of the CAD model) on the CAD model of the target part. As the camera position is continuously updated, these 3D feature points are projected onto the image plane to become 2D feature points. Visual servoing is used to drive the robot to achieve assembly positioning from the starting position to the desired position. The robot's visual servoing task is illustrated as follows: Figure 2 As shown, the formula for the control process of driving the robot is as follows:

[0063]

[0064] v c =[v x ,v y ,v z ,ω x ,ω y ,ω z ω] T

[0065] e = ff *

[0066] Among them, v c v represents the six-degree-of-freedom motion velocity of the camera in the camera coordinate system. x ,v y ,v z Let ω represent the velocity components of the camera along the x, y, and z axes of the camera coordinate system. x ,ω y ,ω z Let x, y, and z represent the angular velocity components of the camera along the x, y, and z axes of the camera coordinate system, respectively. Let T denote the matrix transpose, and λ be the servo gain, satisfying λ∈[0,1]. Let f be the estimate of the pseudo-inverse matrix of the characteristic Jacobian matrix (also known as the interaction matrix), e represent the visual feature error, and f and f'''''''''''''''''''''''''''''''''""""","e' ... * These represent the visual features extracted from the image when the camera is at its current position and desired position, respectively. These visual features can be either image features in the image space (corresponding to image-based visual servoing (IBVS)) or camera pose relative to the target object (corresponding to position-based visual servoing (PBVS)). In this invention, the camera pose relative to the target object is represented by a six-dimensional vector x = [t]. x ,t y ,t z ,θ x ,θ y ,θ z ] T This indicates that the first three digits represent the relative translation component, and the last three digits represent the relative rotation component in Euler angles. The corresponding homogeneous transformation matrix is: Obj M C .

[0067] In (1), the optimization objective is to successfully complete the assembly and positioning task under the constraints and achieve optimal robot motion performance. Robot motion performance includes three aspects: 1) improving robot motion efficiency to shorten the work cycle; 2) reducing robot energy consumption; 3) reducing the impact on robot joints to ensure the smoothness of robot motion.

[0068] In (1), based on the actual task, the main constraints to be considered in the visual servoing process include camera field of view constraints, robot joint constraints and robot speed constraints. Specifically, the camera field of view constraint means that the target object cannot leave the camera's field of view, otherwise the visual features in the image will not be extracted, resulting in the failure of the visual servoing task. Specifically, the robot joint constraint means that the angle of each joint of the robot cannot exceed its limit during the movement. Specifically, the robot speed constraint means that the speed of the robot end cannot exceed the preset upper limit.

[0069] (2) Based on the visual servoing assembly and positioning task and the corresponding optimization objectives and constraints, a hybrid visual servoing (DRL-HVS) controller based on deep reinforcement learning is constructed.

[0070] (2) Specifically:

[0071] Based on the visual servoing assembly and positioning task and the corresponding optimization objectives and constraints, a hybrid visual servoing model is constructed. Based on the hybrid visual servoing model, the DDPG agent is fused to obtain a deep reinforcement learning model for hybrid visual servoing, which is used to adjust the parameters of the hybrid visual servoing model. The hybrid visual servoing model and the deep reinforcement learning model together form a hybrid visual servoing controller based on deep reinforcement learning.

[0072] To eliminate sudden speed changes during the startup phase of visual servoing and ensure continuous robot speed during actual operation, the formula for the hybrid visual servoing model is as follows:

[0073]

[0074]

[0075] e H =[e 3D ,e 2D ] T

[0076]

[0077] Where, v′ c (t) represents the corrected camera motion speed, v c (t) represents the original camera motion velocity obtained from the hybrid visual servoing model, v c (0) represents the camera motion speed obtained by the hybrid visual servoing model at startup, a correction term. The existence of this allows the camera's movement speed to increase from zero. H For the mixed error matrix, e 3D This represents the visual error in the PBVS method, specifically the pose error between the camera's current position and its desired position; e 2D This represents the visual error in the IBVS method, specifically the 2D feature point error between the camera's current position and the desired position. L represents the mixed characteristic Jacobian matrix H The estimated value satisfies L H =[L 3D ,L 2D ] T L 3D Indicates pose error e 3D The corresponding characteristic Jacobian matrix, L2D The error e represents the 2D feature point error. 2D The corresponding characteristic Jacobian matrix. λ is the servo gain, and H represents the weight matrix. Let represent the weighted pseudo-inverse matrix of the mixed feature Jacobian matrix estimates, where n is the number of 2D feature points in the IBVS method. The attenuation factor is generally taken as... h represents the weight value of the PBVS method, a scalar used to adjust the proportion of the PBVS and IBVS methods in the hybrid visual servoing, satisfying h∈[0,1]. Obviously, when h takes the value 0 or 1, the hybrid visual servoing model degenerates into the IBVS method or the PBVS method, respectively.

[0078] Visual servoing is assumed to be a discrete process consisting of several time steps with very short intervals, and the speed of the robot is constant between any two time steps. Therefore, the visual servoing process can be regarded as a Markov decision process. Figure 3 This is a schematic diagram of the DRL-HVS controller, where the Deep Deterministic Policy Gradient (DDPG) agent is an Actor-Critic reinforcement learning agent trained using deterministic policy gradient theory. The state space, action space, and reward function of the deep reinforcement learning model for hybrid vision servoing are set as follows:

[0079] Wherein, the state space s t The formula is as follows:

[0080] s t =[e 3D ,v c ,q,d s ]

[0081] Wherein, the state space s t It is a 20-dimensional vector, e 3D and v c Let represent the camera pose error and camera speed, respectively; q represents the current joint angle value of the robot, satisfying q = [q1, q2, q3, q4, q5, q6], where q1, q2, q3, q4, q5, and q6 are the 6-axis joint angle values ​​of the robot; d s Denotes the danger distance of a 2D feature point in the image plane, satisfying d s =[d x ,d y ], d x ,d y These represent the danger distances of 2D feature points in the x and y directions of the image plane, respectively. For example... Figure 4 As shown, the danger distance d sDefined as the ratio of the furthest distance Δd beyond the safety zone boundary among all feature points to the warning distance δ, i.e., d s =Δd / δ, d s ∈[0,1]. If all feature points are located within the safe zone, their danger distance value is 0. Therefore, the positional information of feature points in the image plane can be described by a fixed-dimensional vector, avoiding changes in the dimension of the state space due to variations in the number of feature points. Thus, the state space designed in this invention is independent of the target object and adaptable to different application scenarios.

[0082] Action space a t It is a two-dimensional vector composed of two key parameters in the hybrid vision servoing model, as shown in the following formula:

[0083] a t =[λ,h]

[0084] During training, each round has only two termination outcomes: successfully reaching the target position or unexpectedly terminating due to failure to meet constraints. If the visual feature error is less than a set threshold, the robot is considered to have reached the desired position, and the round is considered successful. An unexpected termination occurs if any of the following conditions are met: 1) reaching the maximum number of steps for the round; 2) the feature point is outside the camera's field of view; 3) exceeding the robot's joint limits; or 4) exceeding the robot's maximum speed. Therefore, at each time step during training, the agent is given the following reward r. t :

[0085]

[0086] When a round successfully converges, its cumulative reward is a positive value R summed from the motion efficiency index τ, the energy consumption index E, and the stability index η. succeed Specifically, when the visual servoing task is running normally, a fixed negative feedback of -1 / K is provided at each time step. max When the robot reaches the desired position after K steps, the agent will receive a positive reward related to the energy consumption index E and the motion stability index η. Typically, robot energy consumption is calculated by summing the products of the joint torques and corresponding angular velocities at each moment. For ease of calculation, this invention approximates this by using the cumulative rotation angles of each joint throughout the entire task. The motion stability index η is a function related to the robot's joint jerk. As the differential of acceleration in the robot's joint space, the joint jerk is an important indicator reflecting the robot's stability during motion. A larger joint jerk indicates more severe vibration during robot movement, and excessive vibration can seriously affect the robot's lifespan.

[0087] For rounds that terminate unexpectedly, different levels of penalty should be imposed based on the degree of task completion, with the cumulative reward being a negative value R. failed For example, the closer the robot is to the desired position when the task ends, the smaller the penalty should be. This invention selects the ratio of the 2D feature point error between the current state and the initial state. This is used to measure the degree of completion of the visual servoing task. The 2D feature point error can be calculated using the following formula:

[0088]

[0089] Where row and col represent the height and width of the image plane, respectively, p i Represents the 2D feature point at the current position. Let || denote the 2D feature point at the desired location, and || denote the L2 norm.

[0090] Specifically, the formula for the return function is as follows:

[0091]

[0092]

[0093]

[0094]

[0095]

[0096] Among them, R succeed Let R represent the positive reward, τ represent the motion efficiency index, satisfying τ∈[0,1], E represent the energy consumption index, satisfying E∈[0,1], η represent the stationarity index, satisfying η∈[0,1], and R failed K represents a negative reward. max Let K be the maximum number of moves in a round, and K represent the number of moves in the current round. This represents the error of the 2D feature points in the current state. q represents the initial 2D feature point error, where ω is the joint weight coefficient vector. Considering the principle that joints closer to the robot base consume more energy, ω is chosen as [4,4,2,2,1,1]. The normalized energy consumption index can be used to measure robot energy consumption under different task scenarios. d q represents the desired position. init Indicates the initial position, q k This represents the position of the robot at step k, where || represents the absolute value of the difference in positions; the numerator of the energy exponent E in the formula represents the position of the robot from its initial position q using joint movements. init Move to the desired position qd The cumulative rotation angle of each joint is given by the formula, and the denominator represents the cumulative rotation angle of each joint in the current round. ReLU() represents the rectified linear function, satisfying ReLU(·)=max(·,0), and is used to ignore rounds with large joint jerk values ​​in the early stages of training. e This represents the maximum allowed joint jerk value for a real robot, max(|J t |) represents the robot's maximum joint jerk value during the entire task.

[0097] (3) Train the hybrid visual servo controller based on deep reinforcement learning in a virtual environment to obtain the trained hybrid visual servo controller.

[0098] (3) Specifically:

[0099] (3.1) Construct a work scene in the virtual environment that is consistent with the real environment, and synchronize the visual servo assembly and positioning task, as well as the corresponding optimization objectives and constraints, to the virtual environment;

[0100] (3.2) After the initialization of the deep reinforcement learning model is completed, the hybrid visual servo controller interacts with the robot in the virtual environment to generate experience data. The hybrid visual servo controller performs several rounds of training during the interaction process until the cumulative reward of the agent of the deep reinforcement learning model in the hybrid visual servo controller tends to maximize the positive value, and the trained hybrid visual servo controller is obtained.

[0101] In (3.2), at each time step, for the agent of the deep reinforcement learning model, the current agent's Actor network obtains the state space s based on time t. t Calculate the corresponding action space a t And send it to the hybrid visual servo model, the hybrid visual servo model according to the received motion space a t The hybrid error matrix is ​​used to calculate the corrected camera motion velocity and drive the robot's operation in the virtual environment; the agent receives a reward r. t And the state s at the next moment t+1 Then, the state space s of the current time step is... t Action space a t Rewards r t and the state s at the next moment t+1 As a set of empirical data, it is packaged into a unit (s) t ,a t ,r t ,s t+1The data is stored in the experience pool; the current agent randomly extracts experience data from the experience pool, calculates the gradient, and updates the parameters of the Critic network. At the same time, the parameters of the Actor network are also updated at the set frequency.

[0102] (4) Deploy the trained hybrid vision servo controller into the real environment (physical space) to control the robot to perform actual assembly and positioning tasks.

[0103] (4) Specifically:

[0104] In the real environment, the robot's state information is transmitted to the trained hybrid vision servo controller through a data interface. The trained hybrid vision servo controller calculates the robot's next motion speed based on the current state information and sends it to the robot, thereby driving the robot to move until the vision servo assembly and positioning task is completed.

[0105] In practice, the robot's state information in the real environment is first collected from the real environment using sensors (such as cameras, encoders, etc.) and transmitted to the data processor. The tracking algorithm in the data processor can estimate the position and orientation of the current camera relative to the target object in real time, and then calculate other visual features in the image. These visual features, together with the robot's motion data, are then standardized into state information in a certain format and sent to the trained hybrid vision servo controller.

[0106] This invention provides a visual servoing motion control method based on deep reinforcement learning, which mainly consists of two stages: virtual training and real deployment. Combined with... Figure 5 The GSK RB03A1 robot and the AprilTag marker “TAG36_11” shown are used as the target objects. The specific implementation process is as follows:

[0107] A virtual training

[0108] (1) Figure 5 The real-world task scenario shown in (a) is synchronized to Figure 5 The virtual environment shown in (b) includes data such as camera intrinsics, hand-eye matrices, and the relative pose of the target object with respect to the robot base. The DRL-HVS controller was developed using the Reinforcement Learning Toolbox in MATLAB. The kinematic parameters of the GSK RB03A1 robot and the hyperparameters of the DDPG algorithm are shown in Table 1.

[0109] Table 1 shows the kinematic parameters and hyperparameters of the reinforcement learning algorithm for the GCS RB03A1 robot.

[0110]

[0111] Table 2 shows the initial and desired positions of the four feature points of the AprilTag marker on the image plane. Traditional PBVS and IBVS methods are inadequate for this visual servoing task: in PBVS, the feature points will leave the camera's field of view; while in IBVS, the task will fail because the robot's speed exceeds a set threshold. A DRL-HVS controller was built for this visual servoing task, and motion planning was performed through virtual training. After approximately 3000 training iterations, the agent finally converged to a motion scheme that met the design objectives.

[0112] Table 2 shows the initial and desired positions of feature points on the image plane.

[0113]

[0114] Figure 6 The diagram shows the curves of parameters λ and h changing over time during task execution by the controller after training. Generally, the robot's movement speed is proportional to the error of the visual features. Since the error of the visual features is relatively large at the beginning of the visual servoing task, the value of λ will be relatively small at this stage to prevent the robot's movement speed from exceeding the limit. Figure 6 (a)); As the robot approaches the target object, the visual feature error gradually decreases, and the value of λ gradually increases and eventually approaches 1. At this point, the increase in the value of λ is beneficial to the rapid convergence of the visual servoing task. On the other hand, the value of h continuously increases during this process to ensure the smooth completion of the visual servoing task. Figure 6 (b)). As can be seen from the figure, both parameter variation curves are relatively smooth, ensuring the robot's stability throughout the entire motion process. The DRL-HVS controller's task execution performance in the virtual environment is as follows: Figure 7 (a) and Figure 7 As shown in (b).

[0115] B. Actual Deployment

[0116] The trained DRL-HVS controller is then... Figure 8 The hardware and software configurations were ported to a real-world environment. The robot's initial position and the images from the camera are shown below. Figure 9 (a) and Figure 9 As shown in (b), after the visual servoing task is started, the robot reaches the following position: Figure 9 The desired position is shown in (c), and the motion trajectory of the feature points throughout the entire visual servoing process is as follows: Figure 9 As shown in (d). From Figure 9 In (d), it can be observed that the actual trajectory of the feature point is similar to... Figure 9The simulation results in (a) are basically consistent, and the overall motion of the robot is relatively smooth. This shows that the trained controller can be successfully deployed in real-world work scenarios, ensuring the effective execution of visual servoing tasks under image and physical constraints. Figure 9 (e) and Figure 9 Figure (f) shows the joint velocity curves of the robot during task execution in both virtual and real environments. It can be observed that the actual velocity curve lags behind the simulation curve to some extent, but the maximum velocity value is slightly larger than the simulation value. The implication of this phenomenon is that when setting the threshold values ​​for kinematic and dynamic parameters before virtual training, the set values ​​should be slightly smaller than the actual allowable values.

[0117] The basic principles and main features of the present invention have been described in detail above with reference to the accompanying drawings. Using the above invention, motion planning can be effectively achieved for visual servo assembly and positioning tasks of industrial robots. Although the above embodiments only used the AprilTag marker as the target object to achieve motion planning for visual servo tasks, the present invention can also achieve motion planning for visual servo tasks on actual mechanical parts, ensuring the stability of the visual servo task and improving the robot's motion performance. The scope of protection of the present invention is defined by the appended claims, and any modifications made based on the claims of the present invention are within the scope of protection of the present invention.

Claims

1. A robot vision servo motion control method based on deep reinforcement learning, characterized in that, Includes the following steps: (1) Determine the robot's vision servo assembly and positioning task and the corresponding optimization objectives and constraints; (2) Based on the visual servo assembly and positioning task and the corresponding optimization objectives and constraints, a hybrid visual servo controller based on deep reinforcement learning is constructed. Specifically, (2) refers to: Based on the visual servoing assembly and positioning task and the corresponding optimization objectives and constraints, a hybrid visual servoing model is constructed. Based on the hybrid visual servoing model, the DDPG agent is fused to obtain a deep reinforcement learning model for hybrid visual servoing, which is used to adjust the parameters of the hybrid visual servoing model. The hybrid visual servoing model and the deep reinforcement learning model together form a hybrid visual servoing controller based on deep reinforcement learning. The state space, action space, and reward function of the deep reinforcement learning model for hybrid vision servoing are set as follows: Wherein, the state space s t The formula is as follows: s t =[e 3D ,v c ,q,d s ] Among them, e 3D and v c Let represent the visual error and camera velocity in the PBVS method, respectively; q represents the current joint angle value of the robot; d s Indicates the danger distance of a 2D feature point in the image plane; Action space a t The formula is as follows: a t =[λ,h] The formula for the return function is as follows: R succeed =τ+E+η Among them, R succeed Let τ represent the positive reward, E represent the exercise efficiency index, η represent the energy consumption index, and R represent the stability index. failed K represents a negative reward. max Let K be the maximum number of moves in a round, and K represent the number of moves in the current round. This represents the error of the 2D feature points in the current state. q represents the initial 2D feature point error, ω is the joint weight coefficient vector, and q d q represents the desired position. init Indicates the initial position, q k q represents the position of the robot at step k. k+1 This represents the position of the robot at step k+1, || represents taking the absolute value of the difference in positions; ReLU() represents the rectified linear function, J e This represents the maximum allowed joint jerk value for a real robot, max(|J t |) represents the robot's maximum joint jerk value during the entire task; λ is the servo gain, and h represents the weight value of the PBVS method, satisfying h∈[0,1]; (3) Train the hybrid visual servo controller based on deep reinforcement learning in a virtual environment to obtain the trained hybrid visual servo controller; at each time step, for the agent of the deep reinforcement learning model, the current agent's Actor network is based on the state space s obtained at time t. t Calculate the corresponding action space a t And send it to the hybrid visual servo model, the hybrid visual servo model according to the received motion space a t And hybrid error matrix-driven robotic operations in virtual environments; agents receive rewards r t And the state s at the next moment t+1 Then, the state space s of the current time step is... t Action space a t Rewards r t and the state s at the next moment t+1 The empirical data is stored in the experience pool as a set of data; the current agent randomly draws empirical data from the experience pool, calculates the gradient, and updates the parameters of the Critic network. (4) Deploy the trained hybrid vision servo controller into the real environment to control the robot to perform actual assembly and positioning tasks.

2. The robot visual servo motion control method based on deep reinforcement learning according to claim 1, characterized in that, In (1), the visual servo assembly positioning task specifically includes: Determine the robot's starting position and desired position. Under the premise of achieving the optimization objective and satisfying the constraints, use visual servoing to drive the robot to achieve assembly positioning from the starting position to the desired position. The control process of driving the robot is expressed by the following formula: v c =[v x ,v y ,v z ,ω x ,ω y ,ω z ] T e=f-f * Among them, v c v represents the six-degree-of-freedom motion velocity of the camera in the camera coordinate system. x ,v y ,v z Let ω represent the velocity components of the camera along the x, y, and z axes of the camera coordinate system. x ,ω y ,ω z Let x, y, and z represent the angular velocity components of the camera along the x, y, and z axes of the camera coordinate system, respectively. Let T denote the matrix transpose, and λ be the servo gain, satisfying λ∈[0,1]. Here, e represents the pseudo-inverse of the characteristic Jacobian matrix, f is the visual feature error, and f and f' are the estimated values. * These represent the visual features extracted from the image when the camera is at its current position and the desired position, respectively.

3. The robot visual servo motion control method based on deep reinforcement learning according to claim 1, characterized in that, In (1), the optimization objective is to successfully complete the assembly and positioning task under the constraints and to achieve the optimal motion performance of the robot.

4. The robot visual servo motion control method based on deep reinforcement learning according to claim 1, characterized in that, In (1), the constraints include camera field of view constraints, robot joint constraints, and robot speed constraints. Specifically, the camera field of view constraint means that the target object cannot leave the camera's field of view. Specifically, the robot joint constraint means that the angle of each joint of the robot cannot exceed its limit during the movement. Specifically, the robot speed constraint means that the speed of the robot's end effector cannot exceed the preset upper limit.

5. The robot visual servo motion control method based on deep reinforcement learning according to claim 1, characterized in that, The formula for the hybrid vision servoing model is as follows: And H =[and 3D ,And 2D ] T Where, v′ c (t) represents the corrected camera motion speed, v c (t) represents the original camera motion velocity obtained from the hybrid visual servoing model, v c (0) represents the camera motion velocity obtained by the hybrid visual servoing model at startup, e H For the mixed error matrix, e 3D Indicates the visual error in the PBVS method; e 2D This represents the visual error in the IBVS method. L represents the mixed characteristic Jacobian matrix H The estimated value, λ is the servo gain, and H represents the weight matrix. Let represent the weighted pseudo-inverse matrix of the mixed feature Jacobian matrix estimates, where n is the number of 2D feature points in the IBVS method. is the attenuation factor, and h represents the weight value of the PBVS method, satisfying h∈[0,1].

6. The robot visual servo motion control method based on deep reinforcement learning according to claim 1, characterized in that, Specifically, (3) refers to: (3.1) Construct a work scene in the virtual environment that is consistent with the real environment, and synchronize the visual servo assembly and positioning task, as well as the corresponding optimization objectives and constraints, to the virtual environment; (3.2) The hybrid visual servo controller continuously interacts with the robot in the virtual environment to generate experience data. The hybrid visual servo controller is trained for several rounds during the interaction process until the cumulative reward of the deep reinforcement learning model in the hybrid visual servo controller tends to maximize, and the trained hybrid visual servo controller is obtained.

7. The robot visual servo motion control method based on deep reinforcement learning according to claim 1, characterized in that, Specifically, (4) refers to: In the real environment, the robot's state information is transmitted to the trained hybrid vision servo controller through a data interface. The trained hybrid vision servo controller calculates the robot's next motion speed based on the current state information and sends it to the robot, thereby driving the robot to move until the vision servo assembly and positioning task is completed.