Multi-drill-boom dynamic task collaborative optimization method and device based on deep reinforcement learning
Through the collaborative optimization method of multi-drilling arm dynamic task based on deep reinforcement learning, the drilling arm joint motion parameters are optimized in real time, and the problem that the existing technology cannot optimize drilling arm operation tasks in real time is solved, and the safe and efficient execution of drilling tasks under complex conditions is achieved.
Patent Information
- Application Number
- CN202510181188.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-23
AI Technical Summary
The existing technology has limitations in the collaborative planning of multi-drilling arm drilling tasks under complex conditions, and it is impossible to optimize drilling arm operation tasks in real time, resulting in the task execution not being carried out as planned, affecting construction safety and efficiency.
The multi-drilling arm dynamic task collaborative optimization method is adopted based on deep reinforcement learning. By constructing a deep space-time graph convolution network and a multi-drilling arm deep reinforcement learning model, the working state of the drill arm is sensed in real time, and the drill arm joint motion parameters are optimized to ensure that the end of the drill arm is in a better position.
Real-time collaborative optimization of drilling tasks in complex dynamic environments is achieved, the safety and efficiency of drilling tasks are improved, and construction progress and safety are ensured.
Smart Images

Figure CN120031328A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automated drilling equipment, and in particular to a multi-drill arm dynamic task collaborative optimization method and equipment based on deep reinforcement learning. Background Art
[0002] At present, the collaborative planning of multi-drill arm drilling tasks under complex conditions is the key link to ensure the safe, efficient and accurate construction of drilling rigs. In terms of collaborative planning of multi-drill arm drilling tasks, current research focuses on hole sequence planning, path planning and trajectory optimization.
[0003] Among them, the purpose of hole sequence planning is to determine the drilling order of each drill arm, which can be regarded as a traveling salesman problem (TSP). At present, intelligent algorithms such as genetic algorithm, ant colony algorithm, hybrid immune algorithm, etc. are often used to solve it. Multi-drill arm motion planning includes path planning and trajectory optimization, which aims to find the optimal path from the starting point to the target point without collision and interference between the drill arms and with the highest efficiency. Planning algorithms such as RRT, A*, artificial potential field method and D* are often used to solve it.
[0004] There are still some limitations in the existing research. First, the hole sequence selection and motion planning in the drilling task planning are often mutually influential processes, and the current research is usually divided into two stages: hole sequence planning and motion planning. Secondly, the existing task planning model is only suitable for static task planning, and cannot meet the requirements of safe and efficient collaborative planning of drilling tasks in complex and uncertain dynamic environments such as geological conditions changes and mechanical equipment failures, which will lead to the task execution not being carried out as planned. Therefore, a method that can optimize the drilling arm operation task in real time is needed to improve the safety and efficiency of the drilling task and ensure the construction progress and safety. Summary of the invention
[0005] In order to solve the technical problems existing in the known technologies, the present invention provides a multi-drill arm dynamic task collaborative optimization method and device based on deep reinforcement learning.
[0006] The technical solution adopted by the present invention to solve the technical problems existing in the known technology is:
[0007] A multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning, the method establishes a kinematic model of the drill arm of a multi-arm rock drilling rig based on the DH method, calculates and solves the position of the drill arm end according to the motion parameters of the drill arm joint, so as to determine the spatial range of the collaborative operation of each drill arm; constructs a deep spatiotemporal graph convolution network, which is used to extract the posture and motion trajectory of each drill arm end from the laser point cloud data and / or image information collected at the construction site; uses sensors to sense the working state of the drill arm in real time; constructs a multi-drill arm deep reinforcement learning model, defines the state and action based on the kinematic model of the drill arm of the multi-arm rock drilling rig, and establishes a reward function; uses the multi-drill arm deep reinforcement learning model, and plans the joint motion parameters of each drill arm at the next sampling moment based on the joint state of each drill arm and the posture of the drill arm end at the current sampling moment, so that each drill arm end is in a better posture at the next sampling moment.
[0008] Furthermore, the method comprises the following steps:
[0009] Step 1, establishing a kinematic model of the drill arm of a multi-arm drilling rig with a multi-joint chain based on the DH method;
[0010] Step 2: construct a deep spatiotemporal graph convolutional network, collect laser point cloud data and image information of the drilling site, compile a training data set A, and train the deep spatiotemporal graph convolutional network with the training data set A; obtain the position and motion trajectory of each drill arm end with the trained deep spatiotemporal graph convolutional network, and obtain the spatial distribution of each borehole and the spatiotemporal relationship between the motion trajectories of each drill arm end;
[0011] Step 3: construct a multi-drill arm deep reinforcement learning model, define actions, state space and state transfer functions, and construct a reward function based on the success of drilling positioning, the risk of drill arm self-collision, the working time of the drill arm and the movement of the drill arm joints; compile a training data set B based on the spatial distribution of completed drilling holes and the spatiotemporal relationship data between the motion trajectories of the end of each drill arm; train the multi-drill arm deep reinforcement learning model based on the training data set B;
[0012] Step 4: Collect laser point cloud data and image information from the drilling construction site in real time, and input them into the trained deep spatiotemporal graph convolutional network. The deep spatiotemporal graph convolutional network outputs the spatial distribution of the current drilling and the spatiotemporal relationship data between the motion trajectories of each drill arm to the multi-drill arm deep reinforcement learning model. The multi-drill arm deep reinforcement learning model outputs the joint motion parameters of each drill arm and the planning strategy for the motion trajectory of each drill arm end.
[0013] Furthermore, in step 2, the deep spatiotemporal graph convolutional network includes a graph convolutional network and a spatiotemporal convolutional network; the graph convolutional network is used to extract the spatial characteristics of the arrangement of drilling holes, and the spatiotemporal convolutional network is used to extract the spatiotemporal characteristics between the motion trajectories of the drill arm end.
[0014] Furthermore, in step 2, laser radar is used to collect laser point cloud data of the tunnel face; a camera is used to collect on-site construction pictures of the drilling; and the collected data are synchronized in time to compile training data set A.
[0015] Furthermore, in step 3, the state space is defined as the spatial position of the drill arm and the wear state of the machine; the spatial position is the position and posture information of the drill arm in the current workspace; the wear state is the wear condition of the drill arm joints; its influencing factors include the number of rotations of the drill arm, the rotation angle, and the applied force; multiple interrelated drill arms are regarded as multi-agents; the actions of the multi-agents are defined as the position control, posture adjustment, moving speed, drilling process, and multi-drill arm cooperation strategy of the drill arm; the state transition of the multi-agent is defined as the multi-agent state at the next sampling moment under given state and given action conditions; the multi-agent state includes the position, posture, speed, and acceleration of the drill arm; the state transition satisfies the Malv decision process.
[0016] Furthermore, in step 3, the method for constructing a reward function based on the success of drilling positioning, the risk of drill arm self-collision, the drill arm working time and the drill arm joint movement amount includes the following method steps:
[0017] Every time the end of the drill arm reaches a drilling position, there will be a fixed positive reward r a ;
[0018] Considering the self-collision risk of multiple drill arms, when the minimum distance between two arms is lower than the safety threshold, a penalty term r inversely proportional to the distance is set. d , the penalty term r d The absolute value of increases rapidly as the distance decreases, and its expression is as follows:
[0019]
[0020] Among them, d min is the minimum distance calculated after collision detection between the two drill arms using the bounding box algorithm, h 1 and h 2 is a set of penalty coefficients;
[0021] The working time T of the drill arm is defined as: the moving time T from the initial position to the drilling position move and the actual drilling time T drill ; Reward for working hours T The setup is as follows:
[0022]
[0023] Among them, g is the calculation coefficient, and the time reward r T Negative exponential relationship with working hours;
[0024] The joint wear of the drill arm is related to the joint motion Q, and the joint motion is related to the number of joint rotations n, the rotation angle θ, and the applied force F. The following function is used to represent the joint motion Q:
[0025] Q=w n ·n+w θ ·θ+w F ·F;
[0026] Among them, w n 、w θ 、w F They correspond to the weight coefficients of the number of rotations n, the rotation angle θ, and the applied force F, respectively. The size of the coefficient is determined according to the actual wear model and engineering experience;
[0027] Joint wear is used as a penalty item to encourage the drill arm to minimize joint movement when taking action to reduce energy consumption and loss. θ The calculation formula is as follows:
[0028] r θ = -γ(θ+x) (5)
[0029] Among them, γ is a positive weight coefficient used to adjust the influence of joint motion in the reward function, and x is other penalty terms related to the actual environmental conditions of the project.
[0030] Furthermore, the axis-aligned bounding box algorithm is used to detect the collision of the drill arm working space. For each drill arm, an AABB is defined, whose size and position are based on the current state of the drill arm. Then, simulation is performed using MATLAB Robotic Toolbox as the platform to obtain the shortest distance between the drill arms, so as to detect whether there is a collision between them.
[0031] Furthermore, in step 2, a multi-drill arm deep reinforcement learning model is constructed based on the DDPG algorithm.
[0032] Furthermore, in step 1, according to the structural and kinematic characteristics of the multi-joint chain drill arm, the drill arm kinematic model is solved based on the DH method to determine the spatial range of the coordinated work of each drill arm; the following kinematic model of the drill arm of the multi-arm drilling rig is constructed:
[0033]
[0034] Where:
[0035] i represents the joint number;
[0036] represents the homogeneous transformation matrix from the coordinate system of the i-1th joint to the coordinate system of the i-th joint;
[0037] i represents the joint number;
[0038] x represents the x-axis direction in the coordinate system;
[0039] z represents the z-axis direction in the coordinate system;
[0040] α i-1 represents the torsion angle of the i-1th joint;
[0041] a i-1 Represents the translation distance of the i-1th joint along the x-axis;
[0042] a i Represents the translation distance of the i-th joint along the x-axis;
[0043] θ i represents the rotation angle of the i-th joint;
[0044] d i is the offset of the i-th joint axis;
[0045] Rot(x,α i-1 ) describes the change in the rotation state around the x-axis, indicating that the rotation angle from the coordinate system of the i-1th joint to the intermediate coordinate system is α i ;
[0046] Trans(x,a i-1 ) describes the change in the translation state along the x-axis, indicating that the intermediate coordinate system is translated along the x-axis by a i Distance, reaching the origin position of the u-th joint;
[0047] Rot(z,θ i ) describes the change in rotation state around the z-axis, indicating the rotation from the intermediate coordinate system to the i-th coordinate system. The rotation angle is the joint variable θ i ;
[0048] Trans(z,d i ) describes the change in the translation state along the z-axis, translating the coordinate system of the i-th joint along the z-axis by d i distance;
[0049] The posture of the entire drill arm is obtained by multiplying the DH transformation matrices of all joints; given the rotation angle of each joint, the position and posture of each drill arm end are calculated.
[0050] The present invention also provides a device for a multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning, comprising a memory and a processor, wherein the memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of the multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning as described above when executing the computer program.
[0051] The advantages and positive effects of the present invention are:
[0052] By constructing an improved deep spatiotemporal graph convolutional network, using the graph convolutional network to extract the spatial characteristics of the arrangement of drilling holes, using the spatiotemporal convolutional network to consider the spatiotemporal influence between the motion trajectories of the drill arm, and based on the established drill arm operation space collision detection axis alignment bounding box algorithm, the dynamic optimization decision of the operation hole sequence and joint action of each drill arm is realized. Based on the established deep reinforcement learning model, the present invention is able to achieve dynamic collaborative optimization of the operation task by real-time perception of the current drilling task execution status, drill arm working status and other dynamic environmental states under the condition of considering potential spatial conflicts and motion smoothness. Dynamic collaborative optimization is achieved by learning to adjust the mutual influence between multiple drill arms and the environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a workflow diagram of a multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning of the present invention.
[0054] Figure 2 This is a working principle diagram of a multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning of the present invention. DETAILED DESCRIPTION
[0055] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0056] The Chinese meanings of the following English words, phrases and abbreviations are as follows:
[0057] AABB: Axis-aligned bounding box. It can be used as an algorithm for collision detection. It detects whether objects collide by defining an axis-aligned bounding box (a minimum rectangular box) for each object and calculating the distance between these boxes. If the distance between two bounding boxes is less than a certain safety threshold, it is considered that a collision may occur.
[0058] MATLAB Robotic Toolbox: MATLAB Robotic Toolbox, a software toolbox for robot modeling, simulation, and analysis, provides rich functionality to support a variety of tasks in robotics research.
[0059] DH method: A modeling method used to describe the geometric relationship between robot joints and links. By defining a coordinate system for each joint and using a homogeneous transformation matrix to describe the position and posture relationship between adjacent joints, the position and posture of the robot end effector can be calculated.
[0060] DDPG algorithm: A deep reinforcement learning algorithm suitable for reinforcement learning problems in continuous action space. It combines the advantages of deep learning and reinforcement learning by using deep neural networks to approximate policy and value functions and updating network parameters through policy gradient methods.
[0061] GCN: A neural network for processing graph-structured data.
[0062] TCN: A neural network for processing time series data, especially suitable for capturing the spatiotemporal characteristics in time series.
[0063] See also Figure 1 to Figure 2 A multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning, the method establishes a kinematic model of the drill arm of a multi-arm rock drilling rig based on the DH method, calculates and solves the position of the drill arm end according to the motion parameters of the drill arm joint, so as to determine the spatial range of the collaborative operation of each drill arm; constructs a deep spatiotemporal graph convolution network, which is used to extract the posture and motion trajectory of each drill arm end from the laser point cloud data and / or image information collected at the construction site; uses sensors to sense the working status of the drill arm in real time; constructs a multi-drill arm deep reinforcement learning model, defines the state and action based on the kinematic model of the drill arm of the multi-arm rock drilling rig, and establishes a reward function; uses the multi-drill arm deep reinforcement learning model, and plans the joint motion parameters of each drill arm at the next sampling moment according to the joint state of each drill arm and the posture of the drill arm end at the current sampling moment, so that each drill arm end is in a better posture at the next sampling moment.
[0064] Preferably, the method comprises the following steps:
[0065] Step 1: Based on the DH method, a kinematic model of the drill arm of a multi-arm drilling rig with a multi-joint chain is established.
[0066] Step 2: construct a deep spatiotemporal graph convolutional network, collect laser point cloud data and image information from the drilling site, compile a training data set A, and use the training data set A to train the deep spatiotemporal graph convolutional network; obtain the terminal position and terminal motion trajectory of each drill arm from the trained deep spatiotemporal graph convolutional network, and obtain the spatial distribution of each borehole and the spatiotemporal relationship between the motion trajectories of each drill arm terminal.
[0067] Step 3: Build a multi-drill arm deep reinforcement learning model to process and analyze sensor data to achieve comprehensive real-time perception of the progress of drilling task execution and the working status of the drill arm. Define actions, state space, and state transition functions, and build a reward function based on the success of drilling positioning, the risk of drill arm self-collision, the working time of the drill arm, and the movement of the drill arm joints; compile a training data set B based on the spatial distribution of completed drilling and the spatiotemporal relationship data between the motion trajectories of each drill arm end; and train the multi-drill arm deep reinforcement learning model based on the training data set B. The reward function of the multi-drill arm deep reinforcement learning model is used to quantify the effect of the agent taking a certain action, which is a key factor in guiding the agent's learning.
[0068] Step 4: Collect laser point cloud data and image information from the drilling construction site in real time, and input them into the trained deep spatiotemporal graph convolutional network. The deep spatiotemporal graph convolutional network outputs the spatial distribution of the current drilling and the spatiotemporal relationship data between the motion trajectories of each drill arm to the multi-drill arm deep reinforcement learning model. The multi-drill arm deep reinforcement learning model outputs the joint motion parameters of each drill arm and the planning strategy for the motion trajectory of each drill arm end.
[0069] Preferably, in step 2, the deep spatiotemporal graph convolutional network may include a graph convolutional network and a spatiotemporal convolutional network; the graph convolutional network can be used to extract the spatial characteristics of the arrangement of drilling holes, and the spatiotemporal convolutional network can be used to extract the spatiotemporal characteristics between the motion trajectories of the drill arm end.
[0070] Preferably, in step 2, a laser radar may be used to collect laser point cloud data of the tunnel face; a camera may be used to collect on-site construction pictures of the drilling; and the collected data may be synchronized in time to compile a training data set A.
[0071] Use preprocessing methods such as filtering, denoising, and normalization to remove noise and irrelevant information that may be contained in the collected raw data.
[0072] Time domain statistical features, frequency domain features, time-frequency features, etc. can be extracted from the preprocessed data to help understand the progress of the drilling task and the working status of the drill arm. The extracted features are used to estimate the real-time progress of the drilling task and the current working status of the drill arm through the Kalman filter.
[0073] Preferably, in step 3, the state space can be defined as the spatial position of the drill arm and the wear state of the machine; the spatial position is the position and posture information of the drill arm in the current workspace; the wear state is the wear of the drill arm joint; its influencing factors include the number of rotations of the drill arm, the rotation angle, and the applied force; multiple interrelated drill arms can be regarded as multi-agents; the actions of the multi-agents are defined as the position control, posture adjustment, movement speed, drilling process, and multi-drill arm cooperation strategy of the drill arm; the state transition of the multi-agent can be defined as the multi-agent state at the next sampling moment under given state and given action conditions; the multi-agent state includes the position posture, speed, and acceleration of the drill arm; the state transition satisfies the Malv decision process. The state of the deep reinforcement learning model is a description of the environment, which contains all the information the agent needs to make a decision.
[0074] Preferably, in step 3, the method of constructing a reward function based on the success of drilling positioning, the risk of drill arm self-collision, the drill arm working time and the drill arm joint movement amount may include the following method steps.
[0075] Every time the end of the drill arm reaches a drilling position, there will be a fixed positive reward r a .
[0076] Considering the self-collision risk of multiple drill arms, when the minimum distance between two arms is lower than the safety threshold, a penalty term r inversely proportional to the distance can be set. d , the penalty term r d The absolute value of will increase rapidly as the distance decreases, and its expression can be as follows:
[0077]
[0078] Among them, d min is the minimum distance calculated after collision detection between the two drill arms using the bounding box algorithm, h 1 and h 2 is a set of penalty coefficients.
[0079] The working time T of the drill arm is defined as: the moving time T from the initial position to the drilling position move and the actual drilling time T drill ; Reward for working hours T The setup is as follows:
[0080]
[0081] Among them, g is the calculation coefficient, and the time reward r T It has a negative exponential relationship with working hours.
[0082] The joint wear of the drill arm is related to the joint motion Q, and the joint motion is related to the number of joint rotations n, the rotation angle θ, and the applied force F. The joint motion Q can be expressed by the following function.
[0083] Q=w n ·n+w θ ·θ+w F ·F;
[0084] Among them, w n 、w θ 、w F They correspond to the weight coefficients of the number of rotations n, the rotation angle θ, and the applied force F, respectively. The size of the coefficient is determined according to the actual wear model and engineering experience.
[0085] Joint wear is used as a penalty item to encourage the drill arm to minimize joint movement when taking action to reduce energy consumption and loss. θ The calculation formula can be as follows:
[0086] r θ = -γ(θ+x) (5)
[0087] Among them, γ is a positive weight coefficient used to adjust the influence of joint motion in the reward function, and x is other penalty terms related to the actual environmental conditions of the project.
[0088] Preferably, the axis-aligned bounding box algorithm can be used to detect the collision of the drill arm working space; for each drill arm, an AABB is defined, whose size and position are based on the current state of the drill arm; then, simulation is performed using MATLAB Robotic Toolbox as a platform to obtain the shortest distance between the drill arms, so as to detect whether there is a collision between them.
[0089] Preferably, in step 2, a multi-drill arm deep reinforcement learning model is constructed based on the DDPG algorithm. The state space is a continuous action space, so the deep deterministic policy gradient (DDPG) algorithm is selected, and a deep neural network is used for learning and storage. The DDPG algorithm uses a deep neural network to approximate the policy and value functions, and updates the neural network parameters through a policy gradient method.
[0090] Preferably, in step 1, according to the structural and kinematic characteristics of the multi-joint chain drill arm, the drill arm kinematic model is solved based on the DH method to determine the spatial range of the coordinated work of each drill arm; the following multi-arm drilling rig drill arm kinematic model is constructed:
[0091]
[0092] Where:
[0093] i represents the joint number;
[0094] Represents the homogeneous transformation matrix from the coordinate system of the i-1th joint to the coordinate system of the i-th joint; describes the position and posture relationship between two adjacent joint coordinate systems;
[0095] i represents the joint number;
[0096] x represents the x-axis direction in the coordinate system;
[0097] z represents the z-axis direction in the coordinate system;
[0098] α i-1 represents the torsion angle of the i-1th joint;
[0099] a i-1 Represents the translation distance of the i-1th joint along the x-axis;
[0100] a i Represents the translation distance of the i-th joint along the x-axis;
[0101] θ i represents the rotation angle of the i-th joint;
[0102] d i is the offset of the i-th joint axis;
[0103] Rot(x,α i-1 ) describes the change in the rotation state around the x-axis, indicating that the rotation angle from the coordinate system of the i-1th joint to the intermediate coordinate system is α i ;
[0104] Trans(x,a i-1 ) describes the change in the translation state along the x-axis, indicating that the intermediate coordinate system is translated along the x-axis by a i Distance, reaching the origin position of the i-th joint;
[0105] Rot(z,θ i ) describes the change in rotation state around the z-axis, indicating the rotation from the intermediate coordinate system to the i-th coordinate system. The rotation angle is the joint variable θ i ;
[0106] Trans(z,d i ) describes the change in the translation state along the z-axis, translating the coordinate system of the i-th joint along the z-axis by d i distance;
[0107] The posture of the entire drill arm is obtained by multiplying the DH transformation matrices of all joints; given the rotation angle of each joint, the position and posture of each drill arm end are calculated.
[0108] The present invention also provides a device for a multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning, comprising a memory and a processor, wherein the memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of the multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning as described above when executing the computer program.
[0109] The following is a preferred embodiment of the present invention to further illustrate the working process and working principle of the present invention:
[0110] A multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning. This method establishes a kinematic model of the drill arm of a multi-arm rock drilling rig based on the DH method, clarifies the spatial range of collaborative operation of each drill arm, and uses multi-source intelligent sensors to perceive the progress of drilling task execution and the working status of the drill arm in real time. By building a multi-agent deep reinforcement learning model, defining states and actions, and establishing a reward function model, dynamic collaborative optimization of multi-drill arm tasks is achieved.
[0111] Step A: According to the structural and kinematic characteristics of the multi-joint chain drill arm, the drill arm kinematic model is solved based on the DH method to determine the spatial range in which each drill arm can work together. The established drill arm posture kinematic model expression is:
[0112]
[0113] Where:
[0114] i represents the joint number;
[0115] represents the homogeneous transformation matrix from the coordinate system of the i-1th joint to the coordinate system of the i-th joint;
[0116] i represents the joint number;
[0117] x represents the x-axis direction in the coordinate system;
[0118] z represents the z-axis direction in the coordinate system;
[0119] α i-1 represents the torsion angle of the i-1th joint;
[0120] a i-1 Represents the translation distance of the i-1th joint along the x-axis;
[0121] a i Represents the translation distance of the i-th joint along the x-axis;
[0122] θ i represents the rotation angle of the i-th joint;
[0123] d iis the offset of the i-th joint axis;
[0124] Rot(x,α i-1 ) describes the change in the rotation state around the x-axis, indicating that the rotation angle from the coordinate system of the i-1th joint to the intermediate coordinate system is α i ;
[0125] Trans(x,a i-1 ) describes the change in the translation state along the x-axis, indicating that the intermediate coordinate system is translated along the x-axis by a i Distance, reaching the origin position of the i-th joint;
[0126] Rot(z,θ i ) describes the change in rotation state around the z-axis, indicating the rotation from the intermediate coordinate system to the i-th coordinate system. The rotation angle is the joint variable θ i ;
[0127] Trans(z,d i ) describes the change in the translation state along the z-axis, translating the coordinate system of the i-th joint along the z-axis by d i distance;
[0128] The posture of the entire drill arm is obtained by multiplying the DH transformation matrices of all joints; given the rotation angle of each joint, the position and posture of each drill arm end are calculated.
[0129] The posture of the entire robot arm can be obtained by multiplying the DH transformation matrices of all joints. In this way, given the rotation angle of each joint, the position and posture of the end effector can be calculated.
[0130] Step B: Based on laser radar, high-definition camera and other devices, high-precision laser point cloud data and high-definition image information of the face are obtained, and then a deep neural network model of laser point cloud and high-definition image fusion is constructed to obtain the spatial distribution of completed drilling holes, and then the execution progress of the drilling task is obtained. Based on multi-source intelligent sensors such as posture, vibration, sound, mechanics, and kinematics (laser radar, high-definition camera, posture sensor, etc.), the spatial posture of each drill arm and the working parameters of the drill arm such as propulsion speed, rotation speed, propulsion pressure, and impact pressure are perceived in real time, and cross-modal fusion diagnosis of abnormal working conditions of the machinery is performed. By monitoring the changes and statistical characteristics of the eigenvalues, anomaly detection algorithms are used to identify abnormal states of the machinery, and the estimated states and detected anomalies are fed back to the corresponding control system in real time.
[0131] Step C: Based on steps 1 and 2, a multi-agent deep reinforcement learning model is established to define actions, state spaces, and state transfer functions. Meanwhile, a reward function is constructed by comprehensively considering the rewards for drilling positioning, the risk of self-collision of the robotic arm, the working time of the drilling arm, and the movement of the drilling arm joints.
[0132] Among them, this patent defines the state space as the spatial position of the drill arm and the wear state of the machine. The spatial position refers to the position and posture information of the robot arm in the current workspace, and the wear state refers to the wear of the drill arm joints, and its influencing factors include the number of rotations of the robot arm, the rotation angle, the applied force, etc. The actions of the multi-agent are defined as the position control, posture adjustment, movement speed, drilling process, multi-drill arm cooperation strategy, etc. of the drill arm. The state transition of the multi-agent is defined as the state of the multi-agent at the next sampling moment under a given state and given action conditions, including the position posture, speed, acceleration, etc. of the drill arm, and the state transition satisfies the Markov decision process. The reward function of the multi-agent is determined by the reward for successful drilling positioning, the penalty for collision, and the working time T of the drill arm and the joint movement θ of the drill arm during drilling.
[0133] (1) Every time the drill arm reaches a drilling position, there will be a fixed positive reward r a .
[0134] (2) Considering the self-collision risk of multiple robotic arms, when the minimum distance d between the two arms min When it is below the safety threshold, a penalty term r is set that is inversely proportional to the distance d , the penalty term r d The absolute value of will increase rapidly as the distance decreases, and its expression is:
[0135]
[0136] Among them, d min is the minimum distance calculated after collision detection between two manipulators using bounding box algorithm, h 1 and h 2 is a set of penalty coefficients.
[0137] (3) The working time T of the drill arm is defined as follows: The moving time T from the initial position to the drilling position move and the actual drilling time T drill The reward setting for working hours can be expressed as:
[0138]
[0139] Among them, g is the calculation coefficient, and the time reward r T It has a negative exponential relationship with working hours.
[0140] (4) The joint wear of the drill arm is usually related to the joint motion θ, which can include factors such as the number of joint rotations n, the rotation angle θ, and the applied force F, which can be expressed by the following function:
[0141] θ=wn ·n+w θ ·θ+w F ·F (4)
[0142] Among them, w n 、w θ 、w F They correspond to the weight coefficients of the number of rotations n, the rotation angle θ, and the applied force F, respectively. The size of the coefficient is determined according to the actual wear model and engineering experience.
[0143] Joint wear is used as a penalty term to encourage the drill arm to minimize joint movement when taking action, ensuring that energy consumption and loss are not too large. The solution is shown in the following formula:
[0144] r θ = -γ(θ+x) (5)
[0145] Among them, γ is a positive weight coefficient used to adjust the influence of joint motion in the reward function, and x is other penalty terms related to the actual environmental conditions of the project.
[0146] Step D: Use deep neural networks for learning and storage, and select the Deep Deterministic Policy Gradient (DDPG) algorithm, which is a method that combines deep learning and reinforcement learning, and is particularly suitable for reinforcement learning problems in continuous action spaces. The DDPG algorithm uses deep neural networks to approximate policy and value functions, and updates these networks through policy gradient methods.
[0147] DDPG uses a deterministic policy function μ(s; θ μ ), where s is the state, θ μ are the parameters of the policy function, which outputs the best action in a given state. In addition, DDPG also uses a value function Q(s,a;θ Q ), which estimates the expected cumulative reward of taking action a in state s and following policy μ. Q are the parameters of the value function, a is the action, Q is the value, and μ is the strategy.
[0148] In order to stabilize the training process, DDPG introduces the strategy μ′ and value Q′ of the target network, and their corresponding parameters θ μ′ and θ Q′ , corresponding to the main network parameter θ μ and θ Q A time-delayed copy of .
[0149] DDPG uses an experience replay buffer to store samples of recent states, actions, rewards, and next states. These samples are then used for training to reduce the variance of the data and improve sample efficiency. At the same time, the deterministic policy gradient theorem is used to update the policy network. The policy gradient theorem shows that the gradient of the policy network can be calculated by the value function. The parameter θ of the value function Q Q Update by minimizing the following loss function:
[0150] L(θ Q )=E[(Q(s,a;θ Q )-Y) 2 ] (6)
[0151] Where Y is the target q-value, usually calculated using the Bellman equation.
[0152] Finally, the soft update of the parameters, the target network parameters θ μ′ and θ Q′ By soft updating the corresponding main network parameters θ μ and θ Q Get closer.
[0153] Step E: Based on steps A to D, in actual application, by constructing an improved deep spatiotemporal graph convolutional network, the graph convolutional network (GCN) is used to extract the spatial characteristics of the drilling hole arrangement, and the spatiotemporal convolutional network (TCN) is used to consider the spatiotemporal influence between the motion trajectories of the drill arm. Based on the established axis-aligned bounding box algorithm for collision detection in the drilling arm operation space, dynamic optimization decisions of the hole sequence and joint movements of each drilling arm are realized.
[0154] (1) A graph convolutional network (GCN) is used for spatial feature extraction. The input is the hole arrangement, which can be represented as a graph G = (V, E), where V is the set of hole nodes and E is the set of edges representing the spatial relationship between the hole positions. i ∈V has a eigenvector z i .
[0155] (2) A spatiotemporal convolutional network (TCN) is used to consider the spatiotemporal influence of the motion trajectory of the drill arm. The input is the motion trajectory of the drill arm, which can be expressed as a series of states S = {s 1 ,s 2 ,...,s T}, where s t It is the state at time step t, which can include position, velocity, acceleration, etc.
[0156] (3) In order to realize the collision detection of the working space of the drill arm, the axis-aligned bounding box (AABB) algorithm can be used. For each drill arm, an AABB is defined, and its size and position are based on the current state of the drill arm. Then, the simulation is carried out using MATLAB Robotic Toolbox as the platform to obtain the shortest distance between the drill arms, so as to detect whether there is a collision between them.
[0157] (4) Combining the spatial features extracted by GCN and the spatiotemporal features captured by TCN, as well as the collision detection results, and the objective function of minimizing the operation time T and the joint motion Q, an optimization model can be constructed to determine the operation hole sequence and joint motion of each drill arm while satisfying a series of constraints such as collision avoidance, operation efficiency, and safety.
[0158] The above-mentioned deep spatiotemporal graph convolutional network, sensor, lidar, camera, axis-aligned bounding box algorithm, DH method and other components, systems, and algorithms adopt applicable components, systems, and algorithms in the prior art, or adopt components, systems, and algorithms in the prior art and are constructed using conventional technical means.
[0159] The embodiments described above are only used to illustrate the technical ideas and features of the present invention, and their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. The patent scope of the present invention cannot be limited only by this embodiment, that is, any equivalent changes or modifications made to the spirit disclosed by the present invention still fall within the patent scope of the present invention.
Claims
1. A multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning, characterized in that: The method establishes a kinematic model of the drill arm of a multi-arm rock drilling rig based on the DH method, calculates and solves the position of the drill arm end according to the motion parameters of the drill arm joint, so as to determine the spatial range of the collaborative operation of each drill arm; constructs a deep spatiotemporal graph convolution network to extract the posture and motion trajectory of each drill arm end from the laser point cloud data and / or image information collected at the construction site; uses sensors to sense the working state of the drill arm in real time; constructs a multi-drill arm deep reinforcement learning model, defines the state and action based on the kinematic model of the drill arm of the multi-arm rock drilling rig, and establishes a reward function; uses the multi-drill arm deep reinforcement learning model to plan the joint motion parameters of each drill arm at the next sampling moment based on the joint state of each drill arm and the posture of the drill arm end at the current sampling moment, so that each drill arm end is in a better posture at the next sampling moment.
2. The multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning according to claim 1 is characterized in that: The method comprises the following steps: Step 1, establishing a kinematic model of the drill arm of a multi-arm drilling rig with a multi-joint chain based on the DH method; Step 2: construct a deep spatiotemporal graph convolutional network, collect laser point cloud data and image information of the drilling site, compile a training data set A, and train the deep spatiotemporal graph convolutional network with the training data set A; obtain the position and motion trajectory of each drill arm end with the trained deep spatiotemporal graph convolutional network, and obtain the spatial distribution of each borehole and the spatiotemporal relationship between the motion trajectories of each drill arm end; Step 3: Build a multi-drill arm deep reinforcement learning model, define actions, state space, and state transfer functions, and build a reward function based on the success of drilling positioning, the risk of drill arm self-collision, the working time of the drill arm, and the movement of the drill arm joints; A training data set B is compiled based on the spatial distribution of completed drilling holes and the spatiotemporal relationship data between the motion trajectories of the end of each drill arm; a multi-drill arm deep reinforcement learning model is trained based on the training data set B; Step 4: Collect laser point cloud data and image information from the drilling construction site in real time, and input them into the trained deep spatiotemporal graph convolutional network. The deep spatiotemporal graph convolutional network outputs the spatial distribution of the current drilling and the spatiotemporal relationship data between the motion trajectories of each drill arm to the multi-drill arm deep reinforcement learning model. The multi-drill arm deep reinforcement learning model outputs the joint motion parameters of each drill arm and the planning strategy for the motion trajectory of each drill arm end.
3. The multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning according to claim 2 is characterized in that: In step 2, the deep spatiotemporal graph convolutional network includes a graph convolutional network and a spatiotemporal convolutional network; the graph convolutional network is used to extract the spatial characteristics of the arrangement of drilling holes, and the spatiotemporal convolutional network is used to extract the spatiotemporal characteristics between the motion trajectories of the drill arm end.
4. The multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning according to claim 2 is characterized in that: In step 2, laser radar is used to collect laser point cloud data of the tunnel face; a camera is used to collect on-site construction pictures of the drilling; and the collected data are synchronized in time to compile training data set A.
5. The multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning according to claim 2 is characterized in that: In step 3, the state space is defined as the spatial position of the drill arm and the wear state of the machine; the spatial position is the position and posture information of the drill arm in the current workspace; the wear state is the wear condition of the drill arm joint; The influencing factors include the number of drill arm rotations, rotation angles, and applied force; multiple interrelated drill arms are regarded as multi-agents; the actions of the multi-agents are defined as the position control, posture adjustment, movement speed, drilling process, and multi-drill arm cooperation strategy of the drill arm; The state transition of the multi-agent is defined as the multi-agent state at the next sampling moment under given state and given action conditions; the multi-agent state includes the position, posture, speed and acceleration of the drill arm; the state transition satisfies the Malv decision process.
6. The multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning according to claim 2 is characterized in that: In step 3, the method for constructing a reward function based on the success of drilling positioning, the risk of drill arm self-collision, the drill arm working time and the drill arm joint movement amount includes the following method steps: Every time the end of the drill arm reaches a drilling position, there will be a fixed positive reward r a ; Considering the self-collision risk of multiple drill arms, when the minimum distance between two arms is lower than the safety threshold, a penalty term r inversely proportional to the distance is set. d , the penalty term r d The absolute value of increases rapidly as the distance decreases, and its expression is as follows: Among them, d min It is the minimum distance calculated after collision detection between two drill arms using bounding box algorithm, h1 and h2 are a set of penalty coefficients; The working time T of the drill arm is defined as: the moving time T from the initial position to the drilling position move and the actual drilling time T drill ; Reward for working hours T The setup is as follows: Among them, g is the calculation coefficient, and the time reward r T Negative exponential relationship with working hours; The joint wear of the drill arm is related to the joint motion Q, and the joint motion is related to the number of joint rotations n, the rotation angle θ, and the applied force F. The following function is used to represent the joint motion Q: Q=w n ·n+w θ ·θ+w F ·F; Among them, w n 、w θ 、w F They correspond to the weight coefficients of the number of rotations n, the rotation angle θ, and the applied force F, respectively. The size of the coefficient is determined according to the actual wear model and engineering experience; Joint wear is used as a penalty item to encourage the drill arm to minimize joint movement when taking action to reduce energy consumption and loss. θ The calculation formula is as follows: r θ =-γ(θ+x) (5) Among them, γ is a positive weight coefficient used to adjust the influence of joint motion in the reward function, and x is other penalty terms related to the actual environmental conditions of the project.
7. The multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning according to claim 6 is characterized in that: The axis-aligned bounding box algorithm is used to detect the collision of the drill arm working space. For each drill arm, an AABB is defined, and its size and position are based on the current state of the drill arm. Then, simulation is performed using MATLAB Robotic Toolbox as the platform to obtain the shortest distance between the drill arms, so as to detect whether there is a collision between them.
8. The multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning according to claim 2 is characterized in that: In step 2, a multi-drill arm deep reinforcement learning model is constructed based on the DDPG algorithm.
9. The multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning according to claim 2 is characterized in that: In step 1, according to the structural and kinematic characteristics of the multi-joint chain drill arm, the drill arm kinematic model is solved based on the DH method to determine the spatial range of the coordinated work of each drill arm; the following kinematic model of the drill arm of the multi-arm drilling rig is constructed: Where: i represents the joint number; represents the homogeneous transformation matrix from the coordinate system of the i-1th joint to the coordinate system of the i-th joint; i represents the joint number; x represents the x-axis direction in the coordinate system; z represents the z-axis direction in the coordinate system; α i-1 represents the torsion angle of the i-1th joint; a i-1 Represents the translation distance of the i-1th joint along the x-axis; a i Represents the translation distance of the i-th joint along the x-axis; θ i represents the rotation angle of the i-th joint; d i is the offset of the i-th joint axis; Rot(x,α i-1 ) describes the change in the rotation state around the x-axis, indicating that the rotation angle from the coordinate system of the i-1th joint to the intermediate coordinate system is α i ; Trans(x,a i-1 ) describes the change in the translation state along the x-axis, indicating that the intermediate coordinate system is translated along the x-axis by a i Distance, reaching the origin position of the i-th joint; Rot(z,θ i ) describes the change in rotation state around the z-axis, indicating the rotation from the intermediate coordinate system to the i-th coordinate system. The rotation angle is the joint variable θ i ; Trans(z,d i ) describes the change in the translation state along the z-axis, translating the coordinate system of the i-th joint along the z-axis by d i distance; The posture of the entire drill arm is obtained by multiplying the DH transformation matrices of all joints; given the rotation angle of each joint, the position and posture of each drill arm end are calculated.
10. A device for a multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning, comprising a memory and a processor, characterized in that: The memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of the multi-drill arm dynamic task collaborative optimization method based on deep reinforcement learning as described in any one of claims 1 to 9 when executing the computer program.
Citation Information
Cited By
Hole forming position optimization processing method and device, equipment and storage medium
CN120844999A