Mechanical arm shaft hole assembling method and system based on AER-DDPG algorithm
By optimizing the experience replay and path planning of the robot arm through the AER-DDPG algorithm, the problems of task continuity and training efficiency of the robot arm in the shaft-hole assembly task were solved, and the complete process from grasping to installation was completed autonomously, which improved the generalization ability and flexibility of the assembly system.
Patent Information
- Application Number
- CN202510895341.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-19
AI Technical Summary
The existing robotic arm shaft-hole assembly technology lacks task continuity and is difficult to autonomously complete the complete task from grasping to installation. In addition, the existing DDPG algorithm has low training efficiency and slow convergence speed during reinforcement learning, and cannot quickly adapt to complex tasks.
The AER-DDPG algorithm is adopted to optimize experience storage and playback by introducing a dual-value network structure and an adaptive experience replay mechanism. Combined with tree search path planning and YOLOv8 visual recognition, autonomous shaft-hole assembly of the robotic arm is achieved.
It improves the training efficiency and autonomy of the robotic arm in shaft-hole assembly tasks, enhances its adaptability to complex environments, reduces manual intervention, and improves production efficiency and assembly success rate.
Smart Images

Figure CN120663319A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot automated assembly, and in particular to a robot arm shaft hole assembly method and system based on an AER-DDPG algorithm. Background Art
[0002] Shaft-to-hole assembly is one of the most typical automated assembly tasks in industrial manufacturing, widely used in industries such as aerospace, automotive manufacturing, and precision instruments. The core of this task depends on the robot arm's ability to accurately grasp and assemble the workpiece, which places high demands on the robot's control accuracy and environmental adaptability.
[0003] Currently, mainstream approaches for automated assembly tasks involving industrial robots include teach-through programming and model-based policy control. However, these methods rely on manual instruction or require precise environmental information, requiring the operator to pre-set the robot's motion path or movement strategy. While these methods work well in known environments, they are unable to adapt to changing tasks and have poor generalization capabilities, especially in complex, unknown environments.
[0004] To address these issues, researchers have proposed a robotic arm control method based on a deep reinforcement learning algorithm. By combining it with the DDPG (Deep Deterministic Policy Gradient) algorithm, the robotic arm's adaptability to different tasks and environments has been improved to a certain extent. However, the existing DDPG algorithm uses uniform random sampling in the experience replay mechanism during reinforcement learning, which cannot effectively distinguish the experience requirements of different training stages. This results in low training efficiency and slow convergence, making it difficult to quickly learn a stable strategy, especially when faced with complex tasks such as shaft-hole assembly. Furthermore, existing robotic arm assembly technology lacks task continuity and often focuses only on the assembly stage. It assumes that the workpiece has been correctly grasped beforehand, but lacks a complete task plan from grasping to installation, making it impossible for the robotic arm to autonomously complete the entire shaft-hole assembly task. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a robot arm shaft hole assembly method and system based on the AER-DDPG algorithm. On the basis of the traditional DDPG algorithm, the proposed AER (Adaptive Experience Replay)-DDPG algorithm optimizes the experience storage and replay mechanism, adopts a dual-objective value network structure, improves the training efficiency of the deep reinforcement learning algorithm, and realizes the autonomy and continuity of the robot arm in the shaft hole assembly task.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a robot arm shaft hole assembly method based on the AER-DDPG algorithm, comprising: According to the current position of the robot arm and the position of the workpiece, a tree search-based path planning strategy is used to generate joint trajectories to control the robot arm to grasp the workpiece; A deep reinforcement learning model is built based on the AER-DDPG algorithm, which inputs the state space to generate action instructions, controlling the robotic arm to move the workpiece to the installation position and complete the assembly; Among them, the deep reinforcement learning model includes a dual value network and an adaptive experience replay mechanism; the dual value network adopts two independent target critic networks, and suppresses over-estimation by taking the minimum value to calculate the target Q value; the adaptive experience replay mechanism divides the experience pool into a successful experience pool and a failed experience pool, and dynamically adjusts the sampling ratio of the two types of experience based on time difference.
[0007] In a second aspect, the present invention provides a robotic arm shaft hole assembly system based on the AER-DDPG algorithm, comprising: The workpiece grasping module is configured to generate joint trajectories based on the current posture of the robot arm and the posture of the workpiece using a tree search-based path planning strategy to control the robot arm to grasp the workpiece; The shaft-hole assembly module is configured to build a deep reinforcement learning model based on the AER-DDPG algorithm, input the state space to generate action instructions, and control the robot arm to move the workpiece to the installation position and complete the assembly; Among them, the deep reinforcement learning model includes a dual value network and an adaptive experience replay mechanism; the dual value network adopts two independent target critic networks, and suppresses over-estimation by taking the minimum value to calculate the target Q value; the adaptive experience replay mechanism divides the experience pool into a successful experience pool and a failed experience pool, and dynamically adjusts the sampling ratio of the two types of experience based on time difference.
[0008] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the robot arm shaft hole assembly method based on the AER-DDPG algorithm described in the first aspect.
[0009] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the robot arm shaft hole assembly method based on the AER-DDPG algorithm described in the first aspect are implemented.
[0010] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention constructs a complete "grasping-assembly" process for shaft-hole assembly tasks, solving the problem of task discontinuity and allowing the robotic arm to complete the entire process autonomously. In the AER-DDPG algorithm, the dual value network takes the minimum value to suppress Q-value overestimation and improve the accuracy of value assessment; adaptive experience replay dynamically adjusts the success and failure experience sampling based on time difference, breaking through the limitations of uniform sampling, accelerating convergence, and improving training efficiency. This allows the robotic arm to quickly learn stable assembly strategies, effectively handle complex shaft-hole assembly tasks, and enhance practicality and intelligence.
[0011] (2) The present invention integrates YOLOv8 visual recognition and robotic arm motion control modules, which can obtain the 6D pose of workpieces of different sizes and shapes. Combined with the reinforcement learning dynamic adjustment strategy, the robotic arm is controlled to realize the complete shaft hole assembly process from grasping to insertion and installation. This solves the problem that traditional robotic arm assembly methods are mostly targeted at fixed workpieces and fixed trajectories, which are difficult to adapt to different scenarios. It improves the generalization ability and flexibility of the robot assembly system, reduces manual intervention, and improves production efficiency.
[0012] (3) The present invention adopts the AER mechanism to dynamically adjust the experience replay strategy. In the early stage of training, it focuses on learning failure experience to improve the strategy's error avoidance ability; in the later stage of training, it learns more successful experience, accelerates strategy optimization, and improves training efficiency. It solves the problem that the original DDPG uses a fixed ratio of experience sampling, which easily leads to slow reinforcement learning convergence and low sample utilization.
[0013] (4) The present invention introduces a dual critic network structure and selects the minimum value of the two target networks when calculating the target Q value, thereby effectively avoiding the Q value over-estimation problem caused by the single critic network structure of the original DDPG algorithm, making the policy learning more stable, avoiding the policy deviation caused by the wrong high Q value estimation, and solving the problem of the original DDPG algorithm causing the training process to oscillate due to inaccurate target Q value estimation, making the reinforcement learning process smoother and improving the reliability of the final policy.
[0014] (5) The present invention has the capability to perform continuous tasks. The shaft-hole assembly task includes two steps: grasping the workpiece and installing the workpiece. The existing technology can complete a single task, but is not competent for the continuous task. The method provided by the present invention can locate the shaft hole position through visual recognition and control the robotic arm through a deep reinforcement learning algorithm to complete the continuous tasks of grasping and installing.
[0015] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.
[0017] Figure 1 A main flow chart of a manipulator shaft hole assembly method based on the AER-DDPG algorithm provided in an embodiment of the present invention; Figure 2 A detailed flow chart of a manipulator shaft hole assembly method based on the AER-DDPG algorithm provided in an embodiment of the present invention; Figure 3 AER-DDPG algorithm schematic diagram provided by an embodiment of the present invention; Figure 4 AER adaptive sampling function diagram provided by an embodiment of the present invention; Figure 5 A diagram showing the task completion effect of a robotic arm shaft hole assembly method based on the AER-DDPG algorithm provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0019] Example 1 like Figure 1 As shown, this embodiment discloses a robot arm shaft hole assembly method based on the AER-DDPG algorithm, comprising the following steps: S1: Based on the current position of the robot arm and the workpiece position, a tree search-based path planning strategy is used to generate joint trajectories to control the robot arm to grasp the workpiece; S2: Build a deep reinforcement learning model based on the AER-DDPG algorithm, input the state space to generate action instructions, and control the robotic arm to move the workpiece to the installation position and complete the assembly; Among them, the deep reinforcement learning model includes a dual value network and an adaptive experience replay mechanism; the dual value network uses two independent target critic networks to suppress over-estimation by taking the minimum value to calculate the target Q value; the adaptive experience replay mechanism divides the experience pool into a successful experience pool and a failed experience pool, and dynamically adjusts the sampling ratio of the two types of experience based on time difference.
[0020] Next, combine Figure 2 , a robot arm shaft hole assembly method based on the AER-DDPG algorithm disclosed in this embodiment is described in detail.
[0021] The robot arm shaft hole assembly task described in this embodiment includes first grabbing the target workpiece, and then moving the target workpiece to the shaft hole position (installation position) to achieve assembly.
[0022] Before S1, first, based on the environmental image of the task scene obtained by the depth camera, the YOLOv8 network is used to obtain the 6D pose of the workpiece and 3D coordinates of the installation location .
[0023] Among them, YOLOv8 is used to obtain the workpiece pose and installation position, which specifically includes the following steps: 1. Collect a certain number of images of a specific shaft-hole assembly task scenario, use the labelImg tool to annotate the images to create a dataset for the specific task, and use the dataset to train the YOLOv8 model to obtain a trained YOLOv8 model.
[0024] 2. Obtain the RGB image within the field of view through the depth camera fixed on the wrist of the end effector of the robotic arm ( ) and the depth image ( ).
[0025] Use the trained YOLOv8 model to detect the target in the task scene and obtain the diagonal pixel coordinates of the bounding box of the target object (workpiece and shaft hole) , and calculate the center coordinates of the target object:
[0026] Get the height value of the target object through the depth image: .
[0027] 3. Using the intrinsic parameter matrix of the depth camera:
[0028] in, and Respectively represent the equivalent focal length of the camera in the x (image horizontal) and y (image vertical) directions, Represents the coordinates of the principal point on the image plane.
[0029] Convert the pixel coordinates of the target object to image coordinates:
[0030] The coordinates of the target object in the camera coordinate system are:
[0031] 4. Use the manually calibrated camera extrinsic matrix to convert the coordinates in the camera coordinate system to the coordinates in the Cartesian coordinate system of the robot end:
[0032] in, is the rotation matrix, is the translation matrix.
[0033] Finally, the coordinates of the workpiece position and installation position relative to the end effector of the robotic arm are obtained .
[0034] 5. Determine the gripping posture based on the physical constraints of the robot arm in real-world task scenarios. In the shaft-hole assembly task, the shaft body is a symmetrical cylindrical shaft or square shaft. Therefore, in order to maintain the stability of the gripping process and the subsequent installation of the socket, the gripping posture of the robot arm is consistent with the posture of the workpiece. , the rotation angle along the z-axis of the end effector of the robot arm Maintain the initial angle of the robot arm.
[0035] This embodiment uses a depth camera and a YOLOv8 network to accurately acquire the workpiece's 6D pose and the 3D coordinates of its installation location. By using an image annotation training model, transforming the coordinates using the depth image and the camera's intrinsic parameter matrix, and mapping them to the robotic arm's coordinate system using the extrinsic parameter matrix, this method determines the grasping posture based on physical constraints. This enables precise acquisition and conversion of environmental perception to pose, providing a reliable foundation for robotic arm shaft-hole assembly, ensuring gripping and assembly accuracy and stability, addressing the challenge of acquiring workpiece pose in mission scenarios, and improving the efficiency and success rate of autonomous robotic arm assembly tasks. In S1, the obtained workpiece pose is used as the target pose of the robot end effector. Motion planning is implemented through ROS combined with OMPL (Open Source Motion Planning Library). The specific steps include: S101. The current posture of the end effector of the robot arm and the acquired posture of the workpiece are used as the starting point and the end point of the robot arm movement respectively.
[0036] S102. Using a tree search-based path planning strategy, starting from the starting point and the end point, continuously expanding the tree to connect the starting state and the target state until the two trees are connected together to form a feasible path, that is, obtaining the path sequence of the end effector of the robot arm without time information. .
[0037] S103, perform kinematic inverse analysis on each node on the path to obtain the joint state of the robot arm; use the interpolation method to generate a set of joint state sequences of the robot arm on the path, and these state sequences constitute the joint trajectory of the robot arm, that is, the path sequence of each joint angle of the robot arm .
[0038] S104: Perform motion smoothing and time parameterization on the obtained joint trajectory to obtain a trajectory containing time information. ;Sent to the robot controller through franka_ros for the robot to execute, detect the current pose and target pose of the robot end effector, and if the robot can successfully execute the trajectory and execute the grasping instruction, feedback is given to complete the target grasping.
[0039] It should be understood that the kinematic inverse solution, motion smoothing processing and time parameterization operation are all existing technical means that can be routinely implemented by those skilled in the art according to actual needs.
[0040] This embodiment uses the current pose of the robot's end effector and the workpiece as its starting and ending points, employing a tree search strategy to synchronously expand the tree from both ends until it connects. This effectively explores feasible paths, copes with complex environmental constraints, and generates a path sequence suitable for the task. Kinematic inversion and interpolation generate joint trajectories, which are then smoothed and parameterized over time to ensure a continuous, smooth trajectory containing temporal information. This ensures precise execution and improves the success rate of grasping by the robot. This addresses the issues of traditional planning, which are prone to local optimality and rough trajectories, and effectively supports the entire process of shaft-hole assembly tasks.
[0041] In S2, a deep reinforcement learning model is built based on the AER-DDPG algorithm, and the state space is input to generate action instructions to control the robotic arm to move the workpiece to the installation position and complete the assembly.
[0042] (1) Construction of basic elements of the model First construct the state space and action space: The state space consists of multiple parameters, including the gripping position, gripping posture, installation position, and current speed of the robot end effector. The state space can be expressed as: ,in is the position reached by the end effector of the current robot arm (achieved goal), is the grasping posture of the end effector of the robotic arm, is the desired goal reached by the end effector of the robotic arm, is the current speed of the end effector of the robot arm.
[0043] Taking the posture of the robot arm's end effector when the grasping task is completed as part of the state space can ensure that the robot arm maintains the grasping posture and smoothly transitions from grasping to assembly when completing tasks continuously, avoiding excessive movements and excessive speeds.
[0044] The action space consists of the control instructions executed by the robot arm, including the robot end effector in the Cartesian coordinate system. The displacement on the axis can be specifically expressed as: .
[0045] Secondly, determine the termination conditions for each training round during model training: 1. The robotic arm successfully completes its intended task, including grasping or assembly operations. 2. A collision occurs during the movement of the robotic arm; 3. Reach the maximum number of training steps; Finally, design the model's reward function. The reward function guides the algorithm to optimize the action strategy. The reward function consists of three parts: task completion reward, target proximity reward, and collision penalty. The calculation formula is as follows:
[0046] in, 、 、 are the reward coefficients, 、 、 They are task completion rewards, target proximity rewards, and collision penalties.
[0047] The task completion reward is a large positive reward obtained when the end effector reaches the target position after the robot arm performs the action. The calculation formula is as follows:
[0048] in, is the number of iterations in the current round, The maximum number of iterations in a round. The greater the number of iterations to reach the target position in each round, the smaller the positive reward value obtained. This guides the robotic arm to reach the target position faster during training and improves the efficiency of task completion.
[0049] The target proximity reward is determined by calculating the Euclidean distance between the current position of the end effector and the target position after the robot arm performs the action, as shown in the following formula:
[0050] in, is the current position of the end effector of the robot arm, is the target position that the end effector of the robot arm needs to reach, is a certain scaling factor. During the training process, as the end effector of the robotic arm approaches the target position, the reward gradually increases.
[0051] The collision penalty is a negative penalty given when a collision occurs during the movement of the robot arm. Collisions need to be avoided during training, so a large negative reward value needs to be set as a penalty to guide the robot arm to avoid collisions:
[0052] This embodiment employs a multi-dimensional reward function. Task completion rewards incentivize the robot arm to achieve its goals efficiently, accelerating convergence; target proximity rewards gradually increase as distance decreases, promoting continued approach; and collision penalties prevent dangerous maneuvers. These three factors work together to ensure both efficient and accurate task completion and improved robot arm motion safety, enabling stable learning and precise execution in tasks like shaft-hole assembly.
[0053] (2) Algorithm Core Architecture The deep reinforcement learning model is built based on the AER-DDPG algorithm. Its principle is as follows Figure 3 shown.
[0054] 1. Composition of the Deep Reinforcement Learning Model Based on the AER-DDPG Algorithm The DDPG algorithm that introduces AER (Adaptive Experience Replay) adopts an Actor (strategy)-Critic (value) architecture, including an Actor network, a Critic network, and an adaptive experience replay pool.
[0055] (1) The Actor network uses a deterministic strategy to select an action based on the current state. That is, given a state as input, the output is a certain action rather than a probability distribution of the action.
[0056] In this embodiment, the Actor network is a four-layer, fully connected network structure. The input layer represents the actual dimension of the state. The first and second hidden layers contain 400 and 300 neurons, respectively. The output layer represents the actual dimension of the action, which is three-dimensional here, representing the displacement of the robot's end effector along the x, y, and z axes. The activation function is ReLU.
[0057] (2) The Critic network is responsible for evaluating the actions generated by the Actor network. It provides feedback to the Actor network by estimating the value function of the action, helping it improve its strategy.
[0058] In this embodiment, the Critic network is a dual-objective value network, including and , are all four-layer fully connected network structures. The input layer is the actual dimension of the state plus the dimension of the output action. The first hidden layer and the second hidden layer contain 400 and 300 neurons respectively. The output layer is the action Q value, and the activation function is ReLU.
[0059] (3) The adaptive experience replay pool stores the experience gained from the interaction between the robot arm and the environment according to success experience and failure experience, and uses the adaptive experience replay (AER) mechanism for dynamic sampling during the training process.
[0060] 2. Dual-Objective Value Network During the training process of the original DDPG algorithm, a single critic network is used, and parameter optimization is performed by calculating the target Q value through maximization operations, which can easily lead to over-estimation of the Q value and affect the stability of training.
[0061] Considering that overestimation usually stems from the accumulation of positive deviations in the value function during training, the minimum value selection can effectively cut off this deviation and make the target Q value closer to the lower bound of the true value.
[0062] Therefore, this embodiment introduces a dual-objective value network into the AER-DDPG algorithm, that is, two independent critic networks and , each network has different parameters , are randomly initialized at the beginning of training. and is the output of the target network (the delayed updated replica network). This mechanism suppresses the possible overestimation values in the two networks and forces the Q-value estimation to converge in a conservative direction, thereby reducing the overestimation bias.
[0063] Specifically, when calculating the target Q value, the smaller value of the two target critic networks is selected for calculation:
[0064] in, represents the target Q value, Indicates immediate reward, represents the discount factor for future rewards; and Represent the output values of the target value network, Indicates the next state, represents the target policy network, represents the parameters corresponding to the target policy network, 、 Represent the target network 、 Parameters.
[0065] The parameter update of the Critic value network is achieved by gradient descent of its loss function, that is, the corresponding mean square error:
[0066] in, Indicates the current state, Indicates the action taken in the current state.
[0067] The two independent value networks can ensure the independence of parameter updates, thereby reducing the overestimation of Q values, preventing the strategy from falling into an erroneous overestimation state, improving the stability of training, and enhancing the speed of convergence.
[0068] 3. Adaptive Experience Replay Pool During the training process, TD-Error (time difference) represents the difference between the current actual Q value and the target Q value, which reflects the degree to which the agent has learned the current strategy.
[0069] Based on the concept of TD-Error, this embodiment proposes an improved adaptive experience replay pool to dynamically adjust the experience sampling preference based on the agent's learning level of the policy, allowing the model to efficiently utilize useful information and reduce ineffective exploration at different training stages: A high TD-Error indicates that the agent's Q-value estimation is inaccurate and it is still exploring (strategy instability). Considering that the model is still inexperienced with the current task, more failure experiences should be sampled to help the agent learn how to avoid errors. A low TD-Error indicates that the Q-value estimation is accurate and the strategy is converging (strategy stability). Considering that the model has already learned a good strategy, more successful experiences should be sampled to strengthen the model's optimal strategy.
[0070] Specifically, the TD-error calculation formula for each training step is as follows:
[0071] in, Represents the time difference between the two value networks , and perform mean operation on it to obtain , as a tuning parameter for the adaptive sampling strategy , thereby reducing the randomness of single network errors and making the results more stable.
[0072] Furthermore, a Sigmoid sampling ratio adjustment function is designed:
[0073] Among them, k is a hyperparameter that controls the variation of the sampling ratio, which is generally (for example ). 、 are the probabilities of sampling from the failure experience pool and the success experience pool respectively. , that is, when TD-error is high, ;when , that is, when TD-error is low, , the function is S-shaped (Sigmoid), such as Figure 4As shown, this ensures that the sampling ratio changes smoothly between [0,1].
[0074] Sampling a batch (bacth_size) of data from the experience replay pool yields .
[0075] This embodiment breaks through the limitation of "uniform sampling" of the traditional experience replay pool by integrating TD-Error and Sigmoid sampling functions. TD-Error reflects the policy learning state, and the Sigmoid sampling function dynamically adjusts the sampling probability based on the TD-error results. When TD-error is high, the sampling of failed experiences is increased to help the intelligent agent avoid errors. When TD-error is low, the sampling of successful experiences is emphasized to strengthen the optimal strategy, effectively improving training efficiency and stability.
[0076] (3) Model training process 1. Initialization (1) Initialize the Actor network , the Actor network is responsible for outputting the action of the robotic arm (displacement of the end effector of the robotic arm), and the parameters are express; Initialize Critic Network 1 , Critic Network 2 , the critic network evaluates the Q value of the state-action pair (i.e., the quality of the action), and the parameters are express.
[0077] (2) Initialize the target network: Initialize the target network using the parameters of the online network, i.e. , , .
[0078] (3) Initialize the experience replay pool , including the failure experience pool and successful experience pool , used to store the status, actions, rewards, etc. during the training process, that is, .
[0079] (4) Initialization exploration noise , used to improve exploration capabilities.
[0080] (5) Setting the TD-error threshold Serves as a baseline value for adaptive experience replay sampling adjustments.
[0081] 2. Training process For each training episode from 1 to E, the action noise needs to be initialized Used to improve motion exploration capabilities and observe the initial state of the environment For each iteration t from 1 to T, perform the following steps: (1) According to the current strategy and explore noise Select Action .
[0082] (2) Execute action , get rewards , the environment shifts to a new state .
[0083] (3) Experience storage: Determine whether the task is completed in the current time step (the end effector of the robotic arm reaches the specified installation position). If successful, set the success flag to 1 and store it in the successful experience pool. ; If it fails, set the success flag to 0 and store it in the failure experience pool .
[0084] (4) Based on adaptive experience replay (AER), the experience pool is proportional to ( ) Sample N groups of data .
[0085] (5) Calculate the target Q value: .
[0086] (6) Update the Critic network by minimizing the target loss: .
[0087] (7) Update the Actor network and optimize the strategy by maximizing the Q value of the Critic network. The objective function is: ; Parameter update is done by gradient ascent, as shown in the following formula: ; in, Represents the objective function Actor network parameters The gradient of ; N represents the number of samples sampled from the experience pool each time; are the parameters of the Critic network; represents the output of the Critic network (Q network), which evaluates the value of performing action a in state s; Represents the gradient of Q value to action a; Represents the output of the Actor network, Indicates the Actor network output action on its own parameters gradient.
[0088] (8) Update the parameters of the target network using soft update, which is achieved through the following process: ; in, , , are the parameters of the online network, , , are the parameters of the target network, is the update coefficient.
[0089] (9) Repeat the above steps. If the termination conditions are met: task completion or collision occurs, the current round ends; otherwise, the process continues until the maximum number of steps T is reached.
[0090] (10) Adjust exploration noise according to training progress , an exponential decay method is used to gradually reduce the exploration noise. The noise is the largest at the beginning of training, ensuring that the agent tries enough strategies. It tends to 0 in the later stage of training, ensuring that the agent executes the optimal strategy.
[0091] Continue training until E rounds of training are completed to obtain a converged model and deploy it.
[0092] The trained deep reinforcement learning model based on the AER-DDPG algorithm is used to generate action instructions and control the robotic arm to complete the assembly task.
[0093] This implementation builds a model based on the AER-DDPG algorithm. It uses a dual-critic network to evaluate action quality, combined with soft updates to the target network to ensure training stability. Adaptive Experience Replay (AER) divides the pool of successful and failed experience into smaller pools, dynamically adjusting sampling based on TD-error to improve experience utilization efficiency. Furthermore, it explores exponential noise attenuation to balance early exploration with later convergence. This multi-step collaboration effectively addresses the shortcomings of traditional DDPG algorithms, enabling the robot arm to more quickly learn stable and optimal strategies for shaft-hole assembly tasks, improving training efficiency and assembly performance.
[0094] In order to verify the effectiveness of the method provided in this embodiment, this embodiment sets up an environment for testing on the gazebo simulation platform. The communication architecture is shown in the figure. The task completion effect is shown in the figure. Figure 5 shown.
[0095] This paper integrates visual recognition, motion planning, and the AER-DDPG algorithm to construct a complete "grasping-assembly" process, resolving the task discontinuity problem of traditional methods and enabling the robotic arm to adapt to different workpieces and scenarios. Building on the DDPG algorithm, the AER mechanism is introduced to dynamically adjust the success and failure experience sampling based on TD-error. The dual value network minimizes Q-value overestimation, improving training stability and efficiency. YOLOv8 is used to obtain the pose and tree search trajectory planning to ensure precise grasping and assembly by the robotic arm. This overcomes the limitations of uniform sampling and a single network, reduces manual intervention, and enhances the generalization capabilities of the robotic system, allowing the robotic arm to efficiently learn stable strategies and quickly converge on complex shaft-hole assembly tasks, thereby improving production efficiency and assembly success rate.
[0096] Example 2 This embodiment provides a robotic arm shaft-hole assembly system based on the AER-DDPG algorithm, including: The workpiece grasping module is configured to generate joint trajectories based on the current posture of the robot arm and the posture of the workpiece using a tree search-based path planning strategy to control the robot arm to grasp the workpiece; The shaft-hole assembly module is configured to build a deep reinforcement learning model based on the AER-DDPG algorithm, input the state space to generate action instructions, and control the robot arm to move the workpiece to the installation position and complete the assembly; Among them, the deep reinforcement learning model includes a dual value network and an adaptive experience replay mechanism; the dual value network adopts two independent target critic networks, and suppresses over-estimation by taking the minimum value to calculate the target Q value; the adaptive experience replay mechanism divides the experience pool into a successful experience pool and a failed experience pool, and dynamically adjusts the sampling ratio of the two types of experience based on time difference.
[0097] Example 3 This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the robot arm shaft hole assembly method based on the AER-DDPG algorithm as described in the first embodiment above are implemented.
[0098] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the robot arm shaft hole assembly method based on the AER-DDPG algorithm as described in the first embodiment above are implemented.
[0099] The steps or modules involved in Examples 2 to 4 above correspond to those in Example 1. For detailed implementations, please refer to the relevant description of Example 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media that includes one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to perform any method of the present invention.
[0100] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A robot arm shaft hole assembly method based on AER-DDPG algorithm, characterized in that: include: According to the current position of the robot arm and the position of the workpiece, a tree search-based path planning strategy is used to generate joint trajectories to control the robot arm to grasp the workpiece; A deep reinforcement learning model is built based on the AER-DDPG algorithm, which inputs the state space to generate action instructions, controlling the robotic arm to move the workpiece to the installation position and complete the assembly; Among them, the deep reinforcement learning model includes a dual value network and an adaptive experience replay mechanism; the dual value network adopts two independent target critic networks, and suppresses over-estimation by taking the minimum value to calculate the target Q value; the adaptive experience replay mechanism divides the experience pool into a successful experience pool and a failed experience pool, and dynamically adjusts the sampling ratio of the two types of experience based on time difference.
2. The method for assembling a shaft hole of a robot arm based on the AER-DDPG algorithm according to claim 1, wherein: It also includes a pre-trained YOLOv8 model to identify the workpiece posture and installation position in the task scene visual information; the installation position is the axis hole position.
3. The method for assembling a shaft hole of a robot arm based on the AER-DDPG algorithm according to claim 1, wherein: The joint trajectory is generated by adopting a tree search-based path planning strategy according to the current posture of the robot arm and the posture of the workpiece, specifically including: The current posture of the end effector of the robot arm and the posture of the workpiece are used as the starting point and end point of the robot arm movement respectively; Using tree search, based on the starting point and the end point, the tree is continuously expanded to connect the starting state and the target state until the two trees are connected together to form a feasible path; The kinematics of each node on the path is inversely solved to obtain the joint state of the robot arm; an interpolation method is used to generate a set of joint state sequences of the robot arm on the path, and the state sequences constitute the joint trajectory of the robot arm.
4. The method for assembling a shaft hole of a robot arm based on the AER-DDPG algorithm according to claim 1, wherein: The deep reinforcement learning model is constructed based on the AER-DDPG algorithm, the state space is input to generate action instructions, and the robot arm is controlled to move the workpiece to the installation position and complete the assembly. Specifically, it includes: Construct the state space based on the gripping position, gripping posture, installation position and current speed of the end effector of the robot arm; define the action space based on the displacement of the end effector of the robot arm on the xyz axis in the Cartesian coordinate system; Determine the termination conditions for each training round during model training; Determine the reward function that includes task completion reward, goal proximity reward, and collision penalty; Based on the state space, action space, termination condition and reward function, the model is trained until convergence, and the trained model is used to generate motion to control the robotic arm to complete the assembly task.
5. The method for assembling a shaft hole of a robot arm based on the AER-DDPG algorithm according to claim 1, wherein: The dual value network is used to select the smaller value of the two value networks when calculating the target Q value: ; in, represents the target Q value, Indicates immediate reward, represents the discount factor for future rewards; and Represent the output values of the target value network, Indicates the next state, represents the target policy network, represents the parameters corresponding to the target policy network, 、 Represent the target network 、 Parameters.
6. The method for assembling a shaft hole of a robot arm based on the AER-DDPG algorithm according to claim 1, wherein: The adaptive experience replay mechanism divides the experience pool into a successful experience pool and a failed experience pool, and dynamically adjusts the sampling ratio of the two types of experience based on time difference, specifically including: ; in, Represents the time difference between the two value networks , and perform mean operation on it to obtain , as a tuning parameter for the adaptive sampling strategy ; Indicates immediate reward, represents the discount factor for future rewards; and Represent the output values of the target value network, Represents the output of the current value network; 、 Represent the current state and the next state respectively. represents the target policy network, represents the parameters of the current policy network, Represents the parameters corresponding to the target policy network; Adjust the experience sampling tendency based on the Sigmoid sampling ratio adjustment function: ; ; in, It is a hyperparameter that controls the variation of the sampling ratio; 、 are the probabilities of sampling from the failure experience pool and the success experience pool respectively; is the default value. When , increase the failure experience sampling; when , increase the sampling of successful experiences.
7. A robotic arm shaft hole assembly system based on the AER-DDPG algorithm, characterized in that: include: Workpiece The grasping module is configured to generate joint trajectories based on the current posture of the robot arm and the posture of the workpiece using a tree search-based path planning strategy to control the robot arm to grasp the workpiece; The shaft-hole assembly module is configured to build a deep reinforcement learning model based on the AER-DDPG algorithm, input the state space to generate action instructions, and control the robot arm to move the workpiece to the installation position and complete the assembly; Among them, the deep reinforcement learning model includes a dual value network and an adaptive experience replay mechanism; the dual value network adopts two independent target critic networks, and suppresses over-estimation by taking the minimum value to calculate the target Q value; the adaptive experience replay mechanism divides the experience pool into a successful experience pool and a failed experience pool, and dynamically adjusts the sampling ratio of the two types of experience based on time difference.
8. The robot arm shaft hole assembly system based on the AER-DDPG algorithm according to claim 7, characterized in that: It also includes a target recognition module for identifying the workpiece posture and installation position in the task scene visual information based on a pre-trained YOLOv8 model; the installation position is the axis hole position.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the robot arm shaft hole assembly method based on the AER-DDPG algorithm as described in any one of claims 1 to 6 are implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the robot arm shaft hole assembly method based on the AER-DDPG algorithm according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Robot screw hole assembling method and system based on force feedback and displacement detection
CN121374657A
Real-time deviation rectifying method, device, equipment, medium and product for end slope mining cave track
CN121596754A