Drill pipe multi-mechanism cooperative adaptive feeding and unloading integrated equipment and control method thereof
Patent Information
- Application Number
- CN202610930211.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-26
AI Technical Summary
[0005]本申请目的是提供钻杆多机构协同式自适应上下料一体化设备及其控制方法,以解决现有技术中钻杆上下料过程自适应能力不足的问题
[0011] The multi-mechanism collaborative adaptive loading and unloading integrated method for drill pipe provided in this application has the following beneficial effects: First, by collecting various data, this application can lay an information foundation for state perception in complex operating environments; then, by mapping heterogeneous data into a dynamic graph structure and using a graph attention network for neighborhood aggregation, it can mine the spatial correlation features between various mechanisms, thereby forming an accurate spatiotemporal state representation; subsequently, by calculating the coaxiality deviation and confidence level based on the three-dimensional point cloud and identifying the drive wheel resistance based on the six-dimensional force, it can quantify key working condition indicators; next, by inputting the link identifier, spatiotemporal state characteristics, deviation and resistance data into the proximal strategy optimization network for decision-making, it can autonomously generate adaptive control commands for different links; then, based on the commands, it solves the collaborative control sequence in the future time domain, which can pre-simulate the collaborative trajectory of multiple mechanisms; finally, it performs adaptive compensation on the control sequence based on the compensation coefficient and generates drive commands, thus realizing an adaptive and precise control closed loop from perception to execution.
Smart Images

Figure CN122488520B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of drilling machinery control, and in particular to a multi-mechanism collaborative adaptive loading and unloading integrated equipment for drill pipes and its control method. Background Technology
[0002] In the field of drilling construction, realizing multi-mechanism coordination and adaptive control in the drill pipe loading and unloading process is a key technology direction for improving operational efficiency and ensuring construction safety, and it also has important application value for promoting the intelligent development of deep hole drilling equipment.
[0003] Currently, common methods for controlling drill pipe loading and unloading mainly rely on preset program logic or manual assistance. For example, programmable logic controllers can be used to sequentially control hydraulic actuators, or operators can rely on their experience to grab and connect drill pipes. Although some automated equipment can achieve basic loading and unloading functions, their control strategies are mostly open-loop or simple closed-loop regulation.
[0004] However, existing control methods often struggle to adjust control parameters in a timely manner when faced with complex operating conditions such as changes in drill pipe specifications, disturbances in the working environment, and fluctuations in equipment status. This results in insufficient coordination among multiple actuators and an inability to cope with the impact of dynamic factors such as changes in screw resistance and spatial position deviations on the operation process. Summary of the Invention
[0005] The purpose of this application is to provide a multi-mechanism collaborative adaptive loading and unloading integrated equipment for drill pipe and its control method, so as to solve the problem of insufficient adaptive capability in the drill pipe loading and unloading process in the prior art.
[0006] To address the aforementioned technical problems, in a first aspect, this application provides a multi-mechanism collaborative adaptive loading and unloading control method for drill pipes, comprising: Collect six-dimensional force data of the drive wheel, as well as position detection data of hydraulic rods, hydraulic cylinders, lateral hydraulic swing cylinders and longitudinal hydraulic swing cylinders, and obtain three-dimensional point cloud data of the docking station of the upper and lower drill rods and acoustic emission waveform data of the gripper. The six-dimensional force data, the position detection data, the three-dimensional point cloud data, and the acoustic emission waveform data are mapped into a dynamic graph structure whose node attributes are updated over time. The dynamic graph structure is then processed by a graph attention network to aggregate neighborhood node features, thereby obtaining spatiotemporal state features. The coaxiality deviation between the gripper and the lower drill rod and the corresponding confidence coefficient are calculated based on the three-dimensional point cloud data, and the resistance data of the drive wheel is identified based on the six-dimensional force data. The process identifier of the current operation, the spatiotemporal state characteristics, the coaxiality deviation, the confidence coefficient, and the resistance data are input into the near-end strategy optimization network for strategy decision processing to obtain the corresponding process control command. Based on the aforementioned control instructions, the coordinated control sequence of the hydraulic rod, the hydraulic cylinder, the lateral hydraulic swing cylinder, the longitudinal hydraulic swing cylinder, and the hydraulic motor within a future preset time domain is solved. Based on the compensation coefficient, the coordinated control sequence is adaptively compensated to obtain the drive command. Based on the drive command, the hydraulic rod, the hydraulic cylinder, the lateral hydraulic swing cylinder, the longitudinal hydraulic swing cylinder, and the hydraulic motor are controlled.
[0007] Optionally, the step of inputting the current operation stage's stage identifier, the spatiotemporal state characteristics, the coaxiality deviation, the confidence coefficient, and the resistance data into the proximal strategy optimization network for strategy decision processing to obtain the corresponding stage's stage control instructions includes: The mapping unit of the network is optimized by using a near-end strategy to map the process identifier of the current operation process to the corresponding identifier embedding vector. The spatiotemporal state features are input into the first encoding unit of the near-end policy optimization network. The spatiotemporal state features are compressed to a first intermediate dimension through the compression layer of the first encoding unit. A first linear rectified function is applied to the first intermediate dimension through the suppression layer to obtain a first intermediate vector. The first intermediate vector is mapped to a state embedding vector through the mapping layer. The coaxiality deviation, the confidence coefficient, and the resistance data are input into the second encoding unit of the near-end policy optimization network. The coaxiality deviation, the confidence coefficient, and the resistance data are compressed to a second intermediate dimension through the compression layer of the second encoding unit. A second linear rectification function is applied to the second intermediate dimension through the suppression layer to obtain a second intermediate vector. The second intermediate vector is mapped to a deviation embedding vector through the mapping layer. The fusion unit of the network is optimized by a near-end strategy to fuse the identifier embedding vector, the state embedding vector, and the bias embedding vector to obtain a fused feature vector; The policy unit of the near-end policy optimization network is used to suppress and map the fused feature vector to obtain the control command under the current operation. Based on the control commands, the first position command of the hydraulic rod, the second position command of the hydraulic cylinder, the rotation angle command of the transverse hydraulic swing cylinder and the longitudinal hydraulic swing cylinder, and the speed command of the hydraulic motor are generated.
[0008] Optionally, the policy unit of the network optimized by the near-end policy performs suppression and mapping processing on the fused feature vector to obtain the control instructions under the current operation, including: The fused feature vector is input into the policy unit of the near-end policy optimization network, and the corresponding action head in the policy unit is activated according to the stage identifier of the current operation stage. The fused feature vector is mapped through the fully connected layer of the action head to obtain the original feature vector; The original feature vector is split into mean sub-vectors and variance sub-vectors in dimensional order. A Gaussian distribution is constructed based on the mean sub-vectors and variance sub-vectors. The original action vector with the same dimension as the mean sub-vector is obtained by sampling from the Gaussian distribution through the reparameterization method. Based on the preset allowable range of the current operation, the original motion vector is scaled to obtain the position increment command value, clamping force command value, rotation speed command value and angle command value; The position increment command value, the clamping force command value, the rotation speed command value, and the rotation angle command value are combined into a control command for the current operation.
[0009] Optionally, mapping the six-dimensional force data, the position detection data, the three-dimensional point cloud data, and the acoustic emission waveform data into a dynamic graph structure whose node attributes are updated over time includes: The drive wheel, the gripper, the hydraulic rod, the hydraulic cylinder, the transverse hydraulic swing cylinder, and the longitudinal hydraulic swing cylinder are defined as corresponding mechanism nodes, the upper drill rod and the lower drill rod are defined as corresponding workpiece nodes, and the docking station is defined as an environmental node. The values of each node at the corresponding time in the six-dimensional force data, the position detection data, the three-dimensional point cloud data, and the acoustic emission waveform data are used as the load characteristics of the corresponding node. The clamping relationship between the mechanism node and the workpiece node, the motion constraint relationship between the mechanism node and the environment node, and the spatial positioning relationship between the workpiece node and the environment node are taken as an edge set. Each edge in the edge set is assigned the six-dimensional force data, the position detection data, the force value and relative distance value corresponding to the start and end times in the three-dimensional point cloud data. By combining the load characteristics, the edge set, the force values, the relative distance values, and the three-dimensional spatial coordinates of each node, a dynamic graph structure in which node attributes are updated over time is constructed.
[0010] In a second aspect, this application provides a drill pipe multi-mechanism collaborative adaptive loading and unloading integrated device, including: a frame, a drive mechanism, a clamping mechanism and an industrial AI controller, wherein the industrial AI controller is disposed on the frame and is used to execute the drill pipe multi-mechanism collaborative adaptive loading and unloading control method as described in any one of the first aspects; The telescopic end of the hydraulic cylinder is fixed to the frame. The drive mechanism is installed on one side of the frame. The drive mechanism includes a swing assembly. The swing assembly includes a transverse hydraulic swing cylinder fixedly installed on the top of the mounting frame. The drive shaft of the transverse hydraulic swing cylinder is fixed with a rotating shaft. A swing arm rotatably connected to the mounting base is fixed on the rotating shaft. A longitudinal hydraulic swing cylinder with the drive shaft fixed to the mounting base is fixedly installed at one end of the swing arm. The clamping mechanism includes a mounting base, on which two grippers are rotatably connected, and a hydraulic rod is rotatably connected between the two grippers. A connecting component is provided on the grippers. The connecting assembly includes a hydraulic motor fixedly installed on one side of the gripper, and two drive wheels rotatably connected to the other side of the gripper, one of which is fixed to the drive shaft of the hydraulic motor. The hydraulic motor is activated to screw the upper drill rod into the lower drill rod through the drive wheel.
[0011] The multi-mechanism collaborative adaptive loading and unloading integrated method for drill pipe provided in this application has the following beneficial effects: First, by collecting various data, this application can lay an information foundation for state perception in complex operating environments; then, by mapping heterogeneous data into a dynamic graph structure and using a graph attention network for neighborhood aggregation, it can mine the spatial correlation features between various mechanisms, thereby forming an accurate spatiotemporal state representation; subsequently, by calculating the coaxiality deviation and confidence level based on the three-dimensional point cloud and identifying the drive wheel resistance based on the six-dimensional force, it can quantify key working condition indicators; next, by inputting the link identifier, spatiotemporal state characteristics, deviation and resistance data into the proximal strategy optimization network for decision-making, it can autonomously generate adaptive control commands for different links; then, based on the commands, it solves the collaborative control sequence in the future time domain, which can pre-simulate the collaborative trajectory of multiple mechanisms; finally, it performs adaptive compensation on the control sequence based on the compensation coefficient and generates drive commands, thus realizing an adaptive and precise control closed loop from perception to execution.
[0012] Furthermore, this application optimizes the mapping unit, encoding unit, and fusion unit within the network through a near-end strategy to map the link identifier, spatiotemporal state characteristics, and deviation resistance data into embedding vectors and fuse them, thereby achieving the organic integration of multi-source information at the decision-making level. Then, the strategy unit suppresses and maps the fused features to generate control commands, which are then decomposed into specific commands for each hydraulic component. This improves the targeting and response speed of strategy decisions in complex scenarios, thereby ensuring that the control commands for each link are accurately adapted to the current working conditions. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating a multi-mechanism collaborative adaptive loading and unloading control method for drill pipe provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating a specific implementation of a multi-mechanism collaborative adaptive loading and unloading control method for drill pipes, provided in an embodiment of this application. Figure 3 A schematic diagram of the integrated multi-mechanism collaborative adaptive loading and unloading device for drill pipe provided in an embodiment of this application; Figure 4 A schematic diagram of the swing component of the drill pipe multi-mechanism collaborative adaptive loading and unloading integrated equipment provided in the embodiments of this application; Figure 5 A schematic diagram of the hydraulic cylinder of the drill pipe multi-mechanism collaborative adaptive loading and unloading integrated equipment provided in the embodiments of this application; Figure 6 A schematic diagram of the gripper of the drill pipe multi-mechanism collaborative adaptive loading and unloading integrated equipment provided in the embodiments of this application; Figure 7 A schematic diagram of the hydraulic rod of the drill pipe multi-mechanism collaborative adaptive loading and unloading integrated equipment provided in the embodiments of this application.
[0015] Legend: 10. Frame; 20. Drive mechanism; 21. Positioning rod; 22. Mounting frame; 23. Hydraulic cylinder; 24. Swing assembly; 241. Lateral hydraulic swing cylinder; 242. Rotary shaft; 243. Swing arm; 244. Limit adjustment bolt; 30. Clamping mechanism; 31. Mounting base; 32. Support rod; 33. Gripper; 34. Hydraulic rod; 35. Positioning unit; 351. Sleeve; 352. Cage; 36. Connecting assembly; 361. Drive wheel; 362. Hydraulic motor; 37. Longitudinal hydraulic swing cylinder; 40. Industrial AI Controller. Detailed Implementation
[0016] In drill pipe loading and unloading control, existing methods mostly rely on preset program logic or manual operation, and adopt open-loop or simple closed-loop control strategies. When faced with complex working conditions such as changes in drill pipe specifications, disturbances in the working environment, and fluctuations in equipment status, these methods often cannot adjust control parameters in a timely manner. This results in insufficient coordination among multiple actuators and an inability to effectively cope with the continuous impact of dynamic factors such as changes in screw thread resistance and spatial position deviations on the operation process.
[0017] To address this, this application proposes a multi-mechanism collaborative adaptive loading and unloading control method for drill pipes. The core of this method lies in: collecting multi-source data and constructing a dynamic graph structure; then using a graph attention network to extract the spatiotemporal collaborative features between the mechanisms; furthermore, calculating coaxiality deviation based on 3D point clouds and identifying resistance data based on six-dimensional force, along with link identifiers, and inputting these into a proximal strategy optimization network for intelligent decision-making; finally, solving for the collaborative control sequence in the future time domain and generating drive commands. This method achieves end-to-end adaptive control from state perception and feature extraction to strategy decision-making, effectively overcoming the problems of insufficient collaboration and poor adaptability in existing technologies.
[0018] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] The core of this application is to provide a multi-mechanism collaborative adaptive loading and unloading control method for drill pipes, and a flowchart of one specific implementation is shown below. Figure 1 As shown, the method includes: S101. Collect six-dimensional force data of the drive wheel, as well as position detection data of the hydraulic rod, hydraulic cylinder, transverse hydraulic swing cylinder and longitudinal hydraulic swing cylinder, and obtain three-dimensional point cloud data of the docking position of the upper drill rod and the lower drill rod, as well as acoustic emission waveform data of the gripper.
[0020] Among them, six-dimensional force data refers to the force and torque information of the drive wheel and drill pipe in three spatial axes when they are in contact, which is used to sense the resistance changes during the screwing process; three-dimensional point cloud data refers to the spatial position point set of the upper and lower drill pipe docking station obtained by the vision sensor, which is used to analyze the relative pose between the two drill pipes; acoustic emission waveform data refers to the elastic wave signal generated by friction or deformation on the material surface when the gripper holds the drill pipe, which is used to monitor the stability of the gripping state.
[0021] In step S101, six-dimensional force data is first collected by a six-dimensional force sensor installed on the drive wheel, and position detection data of the hydraulic rod, hydraulic cylinder, lateral hydraulic swing cylinder and longitudinal hydraulic swing cylinder are collected by displacement sensor and angle sensor respectively; at the same time, three-dimensional point cloud data of the docking position of the upper drill rod and the lower drill rod are obtained by a three-dimensional vision sensor, and acoustic emission waveform data are collected by an acoustic emission sensor installed on the gripper.
[0022] S102. The six-dimensional force data, the position detection data, the three-dimensional point cloud data, and the acoustic emission waveform data are mapped into a dynamic graph structure whose node attributes are updated over time. The dynamic graph structure is then processed by a graph attention network to aggregate neighborhood node features, thereby obtaining spatiotemporal state features.
[0023] In one specific implementation, step S102 includes: Step 1021: Define the drive wheel, the gripper, the hydraulic rod, the hydraulic cylinder, the transverse hydraulic swing cylinder, and the longitudinal hydraulic swing cylinder as corresponding mechanism nodes, define the upper drill rod and the lower drill rod as corresponding workpiece nodes, and define the docking station as an environmental node.
[0024] In step 1021, based on the type of physical entity, each object participating in the operation is first classified and mapped to three types of nodes in the graph structure. Among them, the drive wheel, gripper, hydraulic rod, hydraulic cylinder, transverse hydraulic swing cylinder, and longitudinal hydraulic swing cylinder are defined as mechanism nodes, the upper drill rod and lower drill rod are defined as workpiece nodes, and the docking station is defined as environment nodes.
[0025] Step 1022: Use the values of each node at the corresponding time in the six-dimensional force data, the position detection data, the three-dimensional point cloud data, and the acoustic emission waveform data as the load characteristics of the corresponding node.
[0026] In step 1022, the sensor values corresponding to each node at the current moment are extracted from the collected six-dimensional force data, position detection data, three-dimensional point cloud data, and acoustic emission waveform data to serve as the load characteristics of that node. For example, the load characteristics of the drive wheel node include three axial forces and three axial moments measured by the six-dimensional force sensor; the load characteristics of the hydraulic rod node include the extension length measured by the displacement sensor; and the load characteristics of the gripper node include the amplitude of the elastic wave signal measured by the acoustic emission sensor.
[0027] Step 1023: Take the clamping relationship between the mechanism node and the workpiece node, the motion constraint relationship between the mechanism node and the environment node, and the spatial positioning relationship between the workpiece node and the environment node as an edge set, and assign each edge in the edge set the six-dimensional force data, the position detection data, the force value and relative distance value corresponding to the start and end time in the three-dimensional point cloud data.
[0028] In step 1023, based on physical connections and operational logic, an edge set between nodes is established, namely: first, three types of relationships are defined: clamping relationships between mechanism nodes and workpiece nodes, such as the contact between the gripper and the upper drill rod, and between the drive wheel and the drill rod; motion constraint relationships between mechanism nodes and environmental nodes, such as the extension and retraction direction constraint of the hydraulic rod and the rotation range constraint of the swing cylinder; and spatial positioning relationships between workpiece nodes and environmental nodes, such as the alignment relationship between the upper drill rod and the docking station. Subsequently, each edge is assigned a corresponding attribute value: force values are extracted from six-dimensional force data, such as the clamping force magnitude, and relative distance values are extracted from position detection data and three-dimensional point cloud data, such as the spatial distance between two nodes.
[0029] Step 1024: Combining the load characteristics, the edge set, the force values, the relative distance values, and the three-dimensional spatial coordinates corresponding to each node, construct a dynamic graph structure in which node attributes are updated over time.
[0030] Among them, three-dimensional spatial coordinates refer to the position information of each node in three-dimensional space: the coordinates of the mechanism nodes are obtained by forward kinematics calculation based on the geometric parameters of the equipment's mechanical structure and the current displacement and angle feedback values of each hydraulic component; the coordinates of the workpiece nodes and the environment nodes are extracted from the collected three-dimensional point cloud data through point cloud segmentation and feature point recognition algorithms.
[0031] In step 1024, firstly, based on the load characteristics and edge set and their attributes of each node, and combined with the three-dimensional spatial coordinates of each node at the current moment, a complete attribute set of each node is formed. These attributes are continuously updated with time steps, thus forming a dynamic graph structure that can dynamically reflect changes in the work scenario.
[0032] Step 1025: Through the linear transformation layer of the graph attention network, perform nonlinear transformation on all features in the dynamic graph structure to obtain the initial embedding vector of all nodes at the current time.
[0033] In step 1025, the original features of all nodes in the constructed dynamic graph structure are input into the linear transformation layer of the graph attention network. This layer linearly combines the original features through a trainable weight matrix and adds a bias term to obtain the initial embedding vector of each node.
[0034] Step 1026: Through the first graph attention layer of the graph attention network, calculate the spatial attention coefficients of each node and its corresponding neighboring nodes on the edge set in the dynamic graph structure, and perform a weighted summation of the spatial attention coefficients and the initial embedding vectors of the corresponding neighboring nodes to obtain the spatial aggregation features of each node.
[0035] In step 1026, spatial information aggregation is performed using the first graph attention layer of the graph attention network. Specifically, for each node, the spatial attention coefficient between it and all its neighboring nodes is calculated. This coefficient is obtained through a learnable attention mechanism based on the initial embedding vectors of the current node and its neighboring nodes, and is dimensionless. The specific formula is as follows: ,in, This represents the spatial attention coefficient of node i to its neighbor node j, with a value ranging from [0,1]. Let be a linear rectified activation function with leakage, and k be any node in the neighbor set of node i. Let i represent the set of neighboring nodes. and Let be the initial embedding vectors for nodes i and j, respectively. To share the linear transformation matrix, || denotes vector concatenation, a T The attention parameter vector is transposed; then, the attention coefficients are used as weights to perform a weighted summation of the initial embedding vectors of the corresponding neighboring nodes, thereby obtaining the spatial aggregation feature of the current node, the dimension of which is the same as that of the initial embedding vector.
[0036] Step 1027: By using the fusion layer of the graph attention network, the spatial aggregation features are superimposed with the initial embedding vector to obtain the residual enhancement features of each node.
[0037] In step 1027, the spatial aggregation features are added element-wise to the initial embedding vector through the fusion layer to obtain the residual enhancement features of each node. This residual connection preserves the original node information, enabling the network to be effectively trained as it deepens, while also enhancing the stability of the features.
[0038] Step 1028: Through the second graph attention layer of the graph attention network, multiple neighbor paths are defined for each node based on clamping relationships, motion constraint relationships, and spatial positioning relationships. For all the neighbor paths corresponding to each node, the path attention coefficients of each neighbor path are calculated based on the residual enhancement features. Based on all the path attention coefficients, the path collaboration features of each node are obtained.
[0039] In step 1028, the collaborative information in multi-hop relationships is further mined through the second graph attention layer of the graph attention network. Specifically, firstly, based on the clamping relationship, motion constraint relationship, and spatial positioning relationship defined in step 1023, multiple neighbor paths are defined for each node. For example, a gripper node can obtain long-distance association information through the path "gripper to upper drill pipe to lower drill pipe to docking station". Then, for each path, the path attention coefficient is calculated based on the residual enhancement features. Specifically, the residual enhancement features of all nodes on the path are concatenated and input into a fully connected layer to obtain the path feature representation. Then, the path feature representation is multiplied by the learnable attention weight vector and normalized by the LeakyReLU activation function to obtain the path attention coefficient of the path, which reflects the importance of the path to the current node. Finally, the path attention coefficients of all paths are weighted and aggregated with the node features on the corresponding path to obtain the path collaborative feature of each node. This feature integrates the collaborative information transmitted by the multi-hop path.
[0040] Step 1029: Using the path collaboration features of all historical moments before the preset window length as the query sequence, calculate the temporal attention coefficients between different time steps in the query sequence through the temporal self-attention layer of the graph attention network according to the node type of each node, and perform weighted summation of the path collaboration features based on the temporal attention coefficients to obtain the spatiotemporal state features at the current moment.
[0041] In step 1029, the dynamic evolution pattern in the time dimension is captured through the temporal self-attention layer of the graph attention network. Specifically: First, the path collaboration features of all historical moments within a preset window length are used as the query sequence; then, for each node, the temporal attention coefficient between different time steps is calculated according to its node type. This coefficient is based on the similarity between the feature vector at the current moment and the feature vector at historical moments, and the specific formula is as follows: ,in, This represents the temporal attention coefficient of node i at historical time t. This represents the set of time steps within a preset window length. This is the query vector at the current moment. Let i be the transpose of the query vector at the current time step. Let be the key vector at historical time t. and d is a learnable linear transformation matrix, and d is the feature dimension. After obtaining the attention coefficients, the spatiotemporal state features of the current moment are obtained by weighted summation based on the path collaboration features of the historical moments.
[0042] This application achieves a deep fusion representation of multi-entity relationships and dynamic processes in complex operational scenarios, thereby providing high-quality state features that include spatial collaboration and temporal evolution for subsequent strategy decisions.
[0043] S103. Calculate the coaxiality deviation between the gripper and the lower drill rod and the corresponding confidence coefficient based on the three-dimensional point cloud data, and identify the resistance data of the drive wheel based on the six-dimensional force data.
[0044] In one specific implementation, step S103 includes: Step 1031: Segment the three-dimensional point cloud data into point clouds at the upper drill rod end and point clouds at the lower drill rod end, and identify the overlapping area between the point clouds at the upper drill rod end and the point clouds at the lower drill rod end. Based on the overlapping area, obtain the overlap metric value.
[0045] In step 1031, the acquired 3D point cloud data of the docking station is first segmented. This is achieved by using a geometric feature-based clustering algorithm, and based on the spatial distribution and local shape features of the point cloud, dividing it into two parts: the point cloud at the upper drill rod end and the point cloud at the lower drill rod end. Then, spatial matching is performed on the two segmented point clouds, and the overlapping area between them is identified. This overlapping area is identified by calculating the nearest neighbor distance between the two point cloud sets: if the distance between a point in the lower drill rod point cloud and a point in the upper drill rod point cloud is less than a preset threshold, then that point is considered to belong to the overlapping area. The preset threshold is set based on the accuracy of the point cloud acquisition equipment and the geometric dimensions of the drill rod end, and its value ranges from 1mm to 5mm. Based on the identified overlapping area, the ratio of the number of point clouds contained within that area to the total number of point clouds at the upper drill rod end is calculated, or the ratio of the spatial volume of the overlapping area to the total spatial volume of the upper drill rod end is calculated, thus obtaining an overlap metric value. This value ranges from 0.0 to 1.0, and a larger value indicates that the two drill rod ends are closer to alignment.
[0046] Step 1032: Based on the overlapping area, calculate the rotational deviation and translational deviation of the gripper relative to the lower drill rod, and combine the rotational deviation and the translational deviation into a coaxiality deviation.
[0047] In step 1032, spatial registration calculation is performed based on the identified overlapping area. Specifically, feature points of the point clouds at the ends of the upper and lower drill rods in the overlapping area are extracted. For example, the principal axis direction of each point cloud is calculated using principal component analysis as the axis direction of the drill rod, and the coordinates of the end center point are obtained through centroid calculation. Then, the rigid body transformation parameters from the lower drill rod coordinate system to the upper drill rod coordinate system are solved. The Euler angles corresponding to the rotation matrix are the rotation deviation, and the translation vector is the translation deviation. These two deviations are then combined to obtain the coaxiality deviation characterizing the relative pose of the two drill rods. The smaller the deviation value, the higher the alignment accuracy of the upper and lower drill rods carried by the gripper.
[0048] Step 1033: Combining the overlap measurement value, the coaxiality deviation, and the point cloud density value at the end of the upper drill pipe, obtain the first fluctuation range of the rotational deviation and the second fluctuation range of the translational deviation.
[0049] In step 1033, an error analysis is performed on the coaxiality deviation. Specifically, considering that the acquisition quality of the point cloud data and the size of the overlapping area affect the accuracy of the deviation calculation, this step combines the overlap metric value obtained in step 1031 and the point cloud density value of the current point cloud to estimate the confidence range of the rotation deviation and translation deviation. Generally, the larger the overlap metric value and the higher the point cloud density value, the higher the accuracy of the deviation calculation and the smaller the fluctuation range. Then, the first fluctuation range of the rotation deviation and the second fluctuation range of the translation deviation are obtained through a pre-constructed mapping table. Specifically, the mapping table is formed by collecting the overlap metric value and point cloud density value under different working conditions during the equipment debugging phase and comparing them with the deviation fluctuation range measured by the high-precision measurement equipment. The overlap metric value and point cloud density value are divided into several intervals, and the deviation fluctuation range under each interval combination is statistically analyzed, thus forming the mapping table, as shown in Table 1.
[0050] Table 1: Mapping Table
[0051] Step 1034: Calculate the confidence coefficient based on the first fluctuation range and the second fluctuation range.
[0052] In step 1034, based on the first fluctuation range and the second fluctuation range, the confidence coefficient of the coaxiality deviation is calculated using the following formula: Where C represents the confidence coefficient. This indicates the width of the first fluctuation range. This indicates the width of the second fluctuation range. The formula, where is the normalization constant, means that the smaller the fluctuation range, the closer the confidence coefficient is to 1, indicating that the calculation result of the coaxiality deviation is more reliable; the larger the fluctuation range, the smaller the confidence coefficient, indicating that the uncertainty of the calculation result is higher.
[0053] Step 1035: Extract the axial force and tangential force values of the drive wheel acting on the drill pipe from the six-dimensional force data, calculate the ratio of the axial force value to the tangential force value, and obtain the instantaneous friction coefficient.
[0054] In step 1035, firstly, based on the geometric relationship between the contact points of the drive wheel and the drill pipe, the total force vector is decomposed from the collected six-dimensional force data into axial force values along the drill pipe axis and tangential force values perpendicular to the axis; then, the ratio of these two components, i.e., the instantaneous friction coefficient, is calculated. ,in, Indicates the instantaneous coefficient of friction. Indicates the tangential force value. This indicates the value of the axial force.
[0055] Step 1036: The vector synthesis result of the axial force value and the tangential force value is taken as the total contact force between the drive wheel and the drill pipe contact surface.
[0056] In step 1036, the axial force value and the tangential force value are vector-synthesized by vector summation to obtain the total contact force between the drive wheel and the drill pipe contact surface.
[0057] Step 1037: When the current operation is a screw-in connection, mark the projection component of the total contact force along the drill pipe axis as the screw-in resistance; or, when the current operation is pipe unloading and disassembly, mark the projection component of the total contact force along the drill pipe tangential direction as the screw-out resistance.
[0058] In step 1037, the current operation stage identifier of the system is first obtained, and it is determined whether it is a screw connection or a pipe disassembly. When it is in the screw connection stage, the upper drill pipe needs to move axially to screw into the lower drill pipe. At this time, the projection component of the total contact force calculated in step 1036 along the drill pipe axis is extracted and marked as the screwing resistance. When it is in the pipe disassembly stage, the upper drill pipe needs to rotate to unscrew from the lower drill pipe. At this time, the projection component of the total contact force along the drill pipe tangential direction is extracted and marked as the unscrewing resistance.
[0059] Step 1038: Associate and store the inward or outward resistance with the instantaneous friction coefficient to generate the resistance data of the drive wheel.
[0060] In step 1038, the identified screwing-in resistance or screwing-out resistance is associated with the instantaneous friction coefficient and stored to form complete resistance data of the drive wheel. This data includes the force and friction state of the drive wheel during the screwing process at the current moment, which can provide key working condition feedback information for subsequent intelligent decision-making. For example, when the screwing-in resistance is too large and the friction coefficient is abnormal, the system can determine that there may be thread jamming and the screwing strategy needs to be adjusted.
[0061] This application accurately calculates the coaxiality deviation and its confidence coefficient between the gripper and the lower drill pipe by processing three-dimensional point cloud data, quantifies the alignment accuracy and evaluates the reliability of the results; by analyzing six-dimensional force data, it accurately identifies the resistance data in the screwing process, which can provide key working condition feedback information for subsequent strategy decisions.
[0062] S104. Input the current operation stage identifier, the spatiotemporal state characteristics, the coaxiality deviation, the confidence coefficient, and the resistance data into the near-end strategy optimization network for strategy decision processing to obtain the corresponding stage control command.
[0063] Among them, the near-end policy optimization network adopts a reinforcement learning framework for end-to-end training. During the training process, historical job data is used as samples, and a reward function is constructed with indicators such as job success rate, completion time, and energy consumption. The network parameters are updated through the policy gradient method. The parameters of the mapping unit, encoding unit, fusion unit, and policy unit are optimized simultaneously through joint training, so that the network can autonomously learn the mapping relationship from multi-source state inputs to the optimal control command. In addition, the design of the action head enables the network to learn differentiated decision-making strategies for different job stages, thereby improving the overall control performance.
[0064] In one specific implementation, such as Figure 2 As shown, step S104 includes: Step 1041: Optimize the mapping unit of the network through the near-end strategy, and map the process identifier of the current operation process to the corresponding identifier embedding vector.
[0065] Among them, the link identifier refers to the number or code used to distinguish different operation stages in the drill pipe loading and unloading process. In this method, it specifically includes links such as grabbing, lifting, centering, twisting connection, pipe unloading and disassembly, and unloading and returning to the warehouse. Each link corresponds to a unique identifier, usually using one-hot encoding or integer encoding. The mapping unit is a module in the near-end policy optimization network, implemented by the embedding layer. Its function is to convert discrete link identifiers into continuous dense vector representations.
[0066] In step 1041, the current operation stage of the system is first obtained, such as the grasping stage or the screw connection stage. This information exists in the form of a stage identifier. Then, the stage identifier is input into the mapping unit of the near-end policy optimization network. The unit maintains a learnable embedding matrix. For the input discrete identifier, the mapping unit retrieves the corresponding row vector from the embedding matrix through a lookup table operation, thereby obtaining the identifier embedding vector.
[0067] Step 1042: Input the spatiotemporal state features into the first encoding unit of the near-end policy optimization network, compress the spatiotemporal state features to a first intermediate dimension through the compression layer of the first encoding unit, apply a first linear rectified function to the first intermediate dimension through the suppression layer to obtain a first intermediate vector, and map the first intermediate vector to a state embedding vector through the mapping layer.
[0068] The first encoding unit is a module in the near-end policy optimization network, consisting of a compression layer, a suppression layer, and a mapping layer. It is used to process high-dimensional spatiotemporal state features. Its compression layer is a fully connected layer used to reduce the dimension of the input features to a preset first intermediate dimension. The suppression layer uses a linear rectified function as the activation function to introduce nonlinearity and suppress negative responses. The mapping layer is another fully connected layer used to further map the intermediate vector into the final state embedding vector.
[0069] In step 1042, the spatiotemporal state features are first input into the compression layer of the first encoding unit, and the high-dimensional spatiotemporal state features are compressed to a preset first intermediate dimension through fully connected computation to reduce the computational complexity of subsequent operations. Then, the compressed vector is input into the suppression layer and a linear rectification function is applied to retain positive information and set negative values to zero, thereby enhancing the sparsity and nonlinear expressive power of the features and obtaining the first intermediate vector. Finally, the intermediate vector is input into the mapping layer and mapped to a fixed-dimensional state embedding vector through another fully connected layer. This vector condenses the comprehensive state information of the current operation scenario.
[0070] Step 1043: Input the coaxiality deviation, the confidence coefficient, and the drag data into the second encoding unit of the near-end policy optimization network. Compress the coaxiality deviation, the confidence coefficient, and the drag data to the second intermediate dimension through the compression layer of the second encoding unit. Apply the second linear rectification function to the second intermediate dimension through the suppression layer to obtain the second intermediate vector. Map the second intermediate vector to the deviation embedding vector through the mapping layer.
[0071] The second coding unit is another coding module in the near-end policy optimization network. It also consists of a compression layer, a suppression layer, and a mapping layer, and is used to process low-dimensional operating parameters composed of coaxiality deviation, confidence coefficient, and drag data.
[0072] In step 1043, the coaxiality deviation, confidence coefficient, and resistance data are first concatenated into a one-dimensional vector and input into the compression layer of the second encoding unit to compress it to a preset second intermediate dimension through fully connected computation. Then, the compressed vector is input into the suppression layer and a linear rectification function is applied to obtain the second intermediate vector. Finally, the intermediate vector is input into the mapping layer to be mapped into a fixed-dimensional deviation embedding vector through the fully connected layer. This vector condenses the key information of the current alignment accuracy and buckle resistance.
[0073] Step 1044: Optimize the fusion unit of the near-end strategy network to fuse the identifier embedding vector, the state embedding vector and the bias embedding vector to obtain a fused feature vector.
[0074] In step 1044, the identifier embedding vector, state embedding vector, and bias embedding vector are input into the fusion unit of the near-end policy optimization network. This unit uses vector concatenation to connect the three vectors along the feature dimension to form a new fusion feature vector.
[0075] Step 1045: The policy unit of the near-end policy optimization network is used to suppress and map the fused feature vector to obtain the control command under the current operation.
[0076] Step 1045 may specifically include the following steps: Step a1: Input the fused feature vector into the policy unit of the near-end policy optimization network, and activate the corresponding action head in the policy unit according to the stage identifier of the current operation stage.
[0077] Among them, the action head refers to the output branch in the strategy unit that is specifically used to process a specific task. Each action head contains a fully connected layer, a parameter splitting module, and a sampling module, which is responsible for mapping the fusion features of the task to specific action parameters.
[0078] In step a1, the fused feature vector is input into the policy unit, which maintains multiple parallel action heads, each corresponding to a specific task, such as a grasping action head or a rotating action head. Subsequently, based on the currently input task identifier, the policy unit activates the corresponding action head, while other action heads remain silent. This design enables the network to learn specialized decision-making strategies for the characteristics of different tasks, thereby improving control accuracy.
[0079] Step a2: Map the fused feature vector through the fully connected layer of the action head to obtain the original feature vector.
[0080] In step a2, the fully connected layer inside the activated action head performs a linear mapping on the input fused feature vector. This fully connected layer transforms the dimension of the fused feature vector to a preset output dimension to obtain the original feature vector, which contains all the information needed to generate specific action parameters.
[0081] Step a3: Split the original feature vector into mean sub-vectors and variance sub-vectors according to the dimensional order. Construct a Gaussian distribution based on the mean sub-vectors and variance sub-vectors. Sample from the Gaussian distribution using the reparameterization method to obtain the original action vector with the same dimension as the mean sub-vector.
[0082] In step a3, the original feature vector is processed. First, according to the preset dimensionality partitioning rules, the vector is split into two parts: a mean sub-vector and a variance sub-vector, both with the same dimension. Then, a multidimensional Gaussian distribution is constructed using the mean sub-vector as the mean and the variance sub-vector as the variance. In order to maintain the propagation of gradients during training, a reparameterization method is used to sample from this multidimensional Gaussian distribution: first, a random noise vector is sampled from the standard normal distribution, and then the original action vector is obtained by adding the product of noise and variance to the mean sub-vector. This original action vector is the action representation in the probability space, reflecting the exploratory nature of the strategy.
[0083] Step a4: Based on the preset allowable range of the current operation, scale the original motion vector to obtain the position increment command value, clamping force command value, rotation speed command value and angle command value.
[0084] The allowable range refers to the minimum and maximum values of each execution parameter preset for different work stages, such as the maximum displacement increment of the hydraulic rod and the maximum speed of the hydraulic motor.
[0085] In step a4, the original motion vector in the probability space is mapped to the actual physical execution range. That is, firstly, the allowable range of each parameter preset in the current operation is obtained. For example, for the screw connection link, the allowable range of rotation speed is 0 to the maximum value. Then, a linear scaling method is used to map each element in the original motion vector from the value range of the probability space to the corresponding physical allowable range, thereby obtaining the position increment command value, clamping force command value, rotation speed command value and angle command value respectively. These command values have actual physical meaning and can be directly used for actuator control.
[0086] Step a5: Combine the position increment command value, the clamping force command value, the rotation speed command value, and the rotation angle command value into a control command for the current operation.
[0087] In step a5, the various instruction values are combined and packaged to form a complete control instruction. This instruction is a structured data that contains all the execution parameters required for the current operation.
[0088] Step 1046: Based on the control command, generate the first position command of the hydraulic rod, the second position command of the hydraulic cylinder, the rotation angle command of the transverse hydraulic swing cylinder and the longitudinal hydraulic swing cylinder, and the speed command of the hydraulic motor.
[0089] In step 1046, the parts related to the hydraulic rod are extracted from the control commands to generate a first position command; the parts related to the hydraulic cylinder are extracted to generate a second position command; the parts related to the swing cylinder are extracted to generate rotation angle commands for the lateral and longitudinal hydraulic swing cylinders; and the parts related to the hydraulic motor are extracted to generate a speed command for the hydraulic motor.
[0090] This application realizes the generation of adaptive control instructions for different work stages, and accurately decomposes the instructions into specific action targets of each actuator, providing a clear control benchmark for subsequent coordinated control.
[0091] S105. Based on the control instructions of the link, solve the coordinated control sequence of the hydraulic rod, the hydraulic cylinder, the transverse hydraulic swing cylinder, the longitudinal hydraulic swing cylinder and the hydraulic motor in the future preset time domain.
[0092] In one specific implementation, step S105 includes: Step 1051: Using the displacement of the hydraulic rod, the displacement of the hydraulic cylinder, the swing angle of the lateral hydraulic swing cylinder, the swing angle of the longitudinal hydraulic swing cylinder, and the rotational speed of the hydraulic motor as state variables, and the increment corresponding to each displacement as control variables, establish a linear prediction model between the state variables and the control variables at each time point in a future preset time domain.
[0093] In step 1051, the system's state variables are first determined as the displacement of the hydraulic rod, the displacement of the hydraulic cylinder, the swing angle of the lateral hydraulic swing cylinder, the swing angle of the longitudinal hydraulic swing cylinder, and the rotational speed of the hydraulic motor, denoted as the system state vector x(k). Then, the increments of each state variable are used as control variables, i.e., the control input at the current moment, denoted as the control variable vector u(k). Subsequently, a discrete-time linear prediction model is established, in the following form: ,in, Let H be the system state vector at time k+1, where k represents the discrete time step, H be the state transition matrix describing the dynamic characteristics of the system itself, and L be the control input matrix describing the influence of the control variables on the state variables. This model is predetermined by a system identification method, such as least squares fitting based on historical operating data. Finally, by iterating over this model, the state values at each time point within the preset time domain N can be calculated.
[0094] Step 1052: Convert the target position increment, target rotation angle, and target rotation speed in the link control command into the expected sequence of each moment in the future preset time domain.
[0095] In step 1052, information such as target position increment, target rotation angle, and target rotation speed are extracted from the link control command. These target values are usually the final endpoint values to be reached, rather than the trajectory of the entire time domain. At this time, they need to be converted into the expected sequence r(k) of the next N time moments. The conversion method can adopt a smooth transition curve, such as an S-curve or polynomial interpolation, so that the state variables can smoothly transition from the current value to the target value, avoiding the impact of sudden changes on the actuator.
[0096] Step 1053: The sum of squared deviations between the expected sequence and the predicted sequence output by the linear prediction model is used as the tracking cost, the sum of squared control variables is used as the energy cost, and the sum of squared differences between control variables at adjacent time points is used as the smoothing cost. The weighted summation formula of the tracking cost, the energy cost, and the smoothing cost is used as the first objective function.
[0097] In step 1053, the first objective function of the optimization problem is constructed, and its mathematical expression is: In this context, the first term represents the tracking cost, Q is the tracking cost weight matrix used to adjust the tracking accuracy for different state variables, and N is the total number of time steps in the future preset time domain, which is a positive integer. The first term is the expected state sequence vector at time k; the second term is the energy cost, R is the energy cost weight matrix, used to suppress the magnitude of the control quantity; the third term is the smoothing cost, S is the smoothing cost weight matrix, used to penalize abrupt changes in the control quantity.
[0098] Step 1054: Unify the preset set of link constraints and the preset obstacle avoidance constraints into a set of linear inequalities, and use a sequential quadratic programming solver to iteratively solve the problem until the first objective function reaches the preset convergence condition.
[0099] In step 1054, all physical constraints are uniformly expressed as a system of linear inequalities: ,in, Here is the state constraint matrix. For the state constraint threshold vector, To control the constraint matrix, To control the constraint threshold vector, This is the joint state and control constraint matrix. The joint constraint threshold vector comprises inequalities that cover the motion range, speed limits, acceleration limits, and spatial constraints required for obstacle avoidance of each actuator. Subsequently, a sequential quadratic programming solver is used to iteratively solve the constrained optimization problem. The solution process is as follows: First, initialize: given an initial control sequence. The value is usually taken as zero or the value from the previous time step; then, in each iteration, the original nonlinear programming problem is approximated as a quadratic programming subproblem at the current point; subsequently, this quadratic programming subproblem is solved to obtain the search direction. Then, the step size α is determined through line search, and the control sequence is updated: Finally, check the convergence conditions, such as the change in the objective function and the gradient norm being less than the threshold. If these conditions are met, stop; otherwise, return to the step of approximating the original nonlinear programming problem as a quadratic programming subproblem at the current point and continue iterating.
[0100] Through the above iterations, the optimal control sequence that satisfies the constraints and minimizes the first objective function is finally obtained.
[0101] Step 1055: Combine the first displacement increment of the hydraulic rod, the second displacement increment of the hydraulic cylinder, the swing angle increment of the lateral hydraulic swing cylinder, the swing angle increment of the longitudinal hydraulic swing cylinder, and the speed increment of the hydraulic motor when the convergence condition is met into a cooperative control sequence for the current moment.
[0102] In step 1055, the optimal control sequence output by the solver is decomposed according to the actuator: the part corresponding to the hydraulic rod is extracted as the first displacement increment of the hydraulic rod, the part corresponding to the hydraulic cylinder is extracted as the second displacement increment of the hydraulic cylinder, the part corresponding to the lateral hydraulic swing cylinder is extracted as the swing angle increment of the lateral hydraulic swing cylinder, the part corresponding to the longitudinal hydraulic swing cylinder is extracted as the swing angle increment of the longitudinal hydraulic swing cylinder, and the part corresponding to the hydraulic motor is extracted as the speed increment of the hydraulic motor. These increments are then combined in time sequence to form a complete cooperative control sequence.
[0103] This application enables coordinated motion planning of multiple actuators, thereby ensuring the accuracy, safety, and smoothness of the operation process.
[0104] S106. Based on the compensation coefficient, adaptive compensation is performed on the cooperative control sequence to obtain a drive command. Based on the drive command, the hydraulic rod, the hydraulic cylinder, the lateral hydraulic swing cylinder, the longitudinal hydraulic swing cylinder, and the hydraulic motor are controlled.
[0105] In one specific implementation, step S106 includes: Step 1061: Collect the actual displacement and actual rotation speed values of the hydraulic rod, the hydraulic cylinder, the transverse hydraulic swing cylinder, the longitudinal hydraulic swing cylinder, and the hydraulic motor.
[0106] In step 1061, for hydraulic rods and hydraulic cylinders, their current actual displacement values are read by displacement sensors; for lateral hydraulic swing cylinders and longitudinal hydraulic swing cylinders, their current actual swing angle values are read by angle sensors; and for hydraulic motors, their current actual speed values are read by speed sensors.
[0107] Step 1062: Calculate the deviation sequence between the cooperative control sequence and the actual displacement value and the actual rotational speed value respectively, and concatenate the deviation sequence with the cooperative control sequence to form an adaptive state vector.
[0108] In step 1062, the expected value at the corresponding moment in the cooperative control sequence is first compared with the actual displacement value and actual rotation speed value collected, and the difference between the two is calculated to obtain the deviation sequence. This sequence reflects the degree of deviation between the theoretical control target and the actual execution result. Then, this deviation sequence is concatenated with the original cooperative control sequence to form a dimension-extended adaptive state vector.
[0109] Step 1063: Input the adaptive state vector into the adaptation layer of the meta-reinforcement learning network. The adaptation layer compresses the adaptive state vector to a third intermediate dimension through a shared coding unit. Based on the third intermediate dimension, compensation is performed through a compensation unit to obtain compensation coefficients.
[0110] Among them, the meta-reinforcement learning network is a reinforcement learning network architecture that can quickly adjust its own parameters when facing new tasks or environmental changes. It consists of a basic learner and a meta-learner. The meta-reinforcement learning network adopts a meta-training method, first training on a large number of historical tasks, and each task contains multiple execution processes. The network quickly adapts to specific working conditions in the inner loop and optimizes the initial parameters of the adaptation layer in the outer loop. The adaptation layer is an online adjustment module in the meta-reinforcement learning network, responsible for quickly generating compensation based on the current state deviation. It typically contains a shared coding unit and a compensation unit. The shared coding unit is a fully connected network used to compress the input adaptive state vector into a unified feature space. The compensation unit is another fully connected network that outputs specific compensation coefficients based on the compressed features.
[0111] In step 1063, the adaptive state vector is input into the pre-trained meta-reinforcement learning network, and its adaptation layer is called for processing. The adaptation layer first performs nonlinear compression on the input adaptive state vector through a shared coding unit and reduces the dimensionality to a preset third intermediate dimension, and then extracts the core features related to compensation. Then, the compressed feature vector is sent to the compensation unit, which calculates the compensation coefficient with the same dimension as the cooperative control sequence through a fully connected layer. This compensation coefficient reflects the proportion of adjustment required to the original control command under the current deviation state.
[0112] Step 1064: Construct a support set using the adaptive state vectors of all historical moments before the preset window length and the corresponding compensation coefficients. Construct a second objective function using the support set. Perform a single gradient step on the parameters of the adaptive layer using the second objective function to obtain the updated compensation coefficients.
[0113] The support set refers to a set of data samples used for rapid adaptation during the meta-learning process. In this step, the support set consists of the adaptive state vectors and corresponding compensation coefficients from multiple recent historical moments. A single gradient step refers to calculating the gradient based on the second objective function and updating the parameters of the adaptation layer with gradient descent to enable the network to adapt to the current working conditions.
[0114] In step 1064, firstly, all adaptive state vectors and their corresponding compensation coefficients within the most recent preset window length are retrieved from the historical database and used to form a support set; then, based on this support set, a second objective function is constructed, specifically in the form of the mean square error between the predicted compensation effect and the ideal compensation effect: AdaptLayer ,in, The loss function value for adaptive compensation, where W represents the support set. For each historical moment, the adaptive state vector is... These are the compensation coefficients output by the network at the corresponding time points. For the ideal compensation target, AdaptLayer is the adaptation layer mapping function of the meta-reinforcement learning network. This function measures the difference between the compensation coefficients output by the network and the ideal compensation effect under the current conditions. Then, the gradient of the objective function with respect to the adaptation layer parameters is calculated, and a single gradient step is performed to fine-tune the internal parameters of the adaptation layer. Finally, after this fast update, the updated adaptation layer is used again to process the current adaptive state vector to obtain the updated compensation coefficients.
[0115] Step 1065: Multiply the updated compensation coefficients element-wise with the corresponding displacement increment, swing angle increment, and speed increment in the cooperative control sequence to generate the drive command at the current moment.
[0116] In step 1065, the updated compensation coefficients are multiplied element-wise with the obtained cooperative control sequence. Term greater than 1 in the compensation coefficients amplifies the corresponding control increment, while term less than 1 reduces the corresponding control increment, thereby achieving adaptive correction of the original control sequence. The result of the multiplication is the final drive command, which includes the compensated hydraulic rod displacement command, hydraulic cylinder displacement command, lateral hydraulic swing cylinder rotation angle command, longitudinal hydraulic swing cylinder rotation angle command, and hydraulic motor speed command. These commands are then sent to the corresponding actuators to achieve precise control of the entire loading and unloading process.
[0117] This application can effectively eliminate the influence of model errors and environmental disturbances on control accuracy, thereby achieving high-precision and robust control of the actuator.
[0118] Figure 3 This is a schematic diagram of a specific embodiment of a drill pipe multi-mechanism collaborative adaptive loading and unloading integrated device provided in this application, with reference to... Figure 3 The device may include: a frame 10, a drive mechanism 20, a clamping mechanism 30, and an industrial AI controller 40, wherein the industrial AI controller 40 is disposed on the frame 10; The telescopic end of the hydraulic cylinder 23 is fixed to the frame 10. A drive mechanism 20 is installed on one side of the frame 10. The drive mechanism 20 includes a positioning rod 21 fixed to one side of the frame 10. A mounting frame 22 is slidably connected to the positioning rod 21. The hydraulic cylinder 23 is fixedly installed on the top of the mounting frame 22.
[0119] Reference Figure 4 The drive mechanism 20 includes a swing assembly 24, which includes a transverse hydraulic swing cylinder 241 fixedly mounted on the top of the mounting frame 22. The drive shaft of the transverse hydraulic swing cylinder 241 is fixed with a rotating shaft 242. A swing arm 243 rotatably connected to the mounting base 31 is fixed on the rotating shaft 242. A longitudinal hydraulic swing cylinder 37, whose drive shaft is fixed to the mounting base 31, is fixedly mounted at one end of the swing arm 243.
[0120] Reference Figure 3 The bottom of the rotating shaft 242 is threaded with a limit adjustment bolt 244. The rotation angle of the rotating shaft 242 is controlled by the limit adjustment bolt 244, and the swing position of the drill rod is adjusted so as to achieve the alignment of drill rods of different diameters.
[0121] Reference Figure 3 The system also includes a clamping mechanism 30, as shown in the reference. Figure 5 , Figure 6 and Figure 7The clamping mechanism 30 includes a mounting base 31, on which two grippers 33 are rotatably connected, and a hydraulic rod 34 is rotatably connected between the two grippers 33. A connecting component 36 is provided on the grippers 33.
[0122] Reference Figure 3 and Figure 5 A support rod 32 is fixed on the mounting base 31. In order to improve the alignment accuracy of the upper and lower drill rods and keep the drill rods horizontally stable when feeding, a positioning unit 35 is provided on the support rod 32. The positioning unit 35 includes two sleeves 351 sleeved on the support rod 32. Both ends of the sleeves 351 are fixedly connected to retainers 352 for supporting the drill rods and playing a positioning role. The sleeves 351 are slidably connected to the support rod 32 and fixed by fixing pins. By adjusting the distance between the two sleeves 351, it is easy to support and position drill rods of different lengths. When connecting the drill rods, the connection points of the upper and lower drill rods are positioned by the support of the two retainers 352 to improve the coaxiality of the drill rods and ensure that the two drill rods are connected smoothly.
[0123] The connecting assembly 36 includes a hydraulic motor 362 fixedly installed on one side of the gripper 33. Two drive wheels 361 are rotatably connected to the other side of the gripper 33, and one of the drive wheels 361 is fixed to the drive shaft of the hydraulic motor 362. The hydraulic motor 362 is started to screw the upper drill rod into the lower drill rod through the drive wheel 361.
[0124] This application provides an embodiment of a drill pipe multi-mechanism collaborative adaptive loading and unloading integrated device to implement the aforementioned drill pipe multi-mechanism collaborative adaptive loading and unloading integrated method. Therefore, the specific implementation of the drill pipe multi-mechanism collaborative adaptive loading and unloading integrated device can be found in the embodiment section of the drill pipe multi-mechanism collaborative adaptive loading and unloading integrated method described above. The specific implementation can be referred to the description of the corresponding embodiments, which will not be repeated here.
[0125] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0126] The above provides a detailed description of the integrated multi-mechanism cooperative adaptive loading and unloading device for drill pipe and its control method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A multi-mechanism collaborative adaptive loading and unloading control method for drill pipes, characterized in that, include: Collect six-dimensional force data of the drive wheel, as well as position detection data of hydraulic rods, hydraulic cylinders, lateral hydraulic swing cylinders and longitudinal hydraulic swing cylinders, and obtain three-dimensional point cloud data of the docking station of the upper and lower drill rods and acoustic emission waveform data of the gripper. The six-dimensional force data, the position detection data, the three-dimensional point cloud data, and the acoustic emission waveform data are mapped into a dynamic graph structure whose node attributes are updated over time. The dynamic graph structure is then processed by a graph attention network to aggregate neighborhood node features, thereby obtaining spatiotemporal state features. The coaxiality deviation between the gripper and the lower drill rod and the corresponding confidence coefficient are calculated based on the three-dimensional point cloud data, and the resistance data of the drive wheel is identified based on the six-dimensional force data. The process identifier of the current operation, the spatiotemporal state characteristics, the coaxiality deviation, the confidence coefficient, and the resistance data are input into the near-end strategy optimization network for strategy decision processing to obtain the corresponding process control command. Based on the aforementioned control instructions, the coordinated control sequence of the hydraulic rod, the hydraulic cylinder, the lateral hydraulic swing cylinder, the longitudinal hydraulic swing cylinder, and the hydraulic motor within a future preset time domain is solved. Based on the compensation coefficient, the coordinated control sequence is adaptively compensated to obtain the drive command. Based on the drive command, the hydraulic rod, the hydraulic cylinder, the lateral hydraulic swing cylinder, the longitudinal hydraulic swing cylinder, and the hydraulic motor are controlled. The adaptive compensation of the coordinated control sequence based on the compensation coefficient to obtain the driving command includes: The actual displacement and actual rotational speed values of the hydraulic rod, the hydraulic cylinder, the transverse hydraulic swing cylinder, the longitudinal hydraulic swing cylinder, and the hydraulic motor are collected. Calculate the deviation sequences between the cooperative control sequence and the actual displacement value and the actual rotation speed value respectively, and concatenate the deviation sequences with the cooperative control sequence to form an adaptive state vector; The adaptive state vector is input into the adaptation layer of the meta-reinforcement learning network. The adaptation layer compresses the adaptive state vector to a third intermediate dimension through a shared coding unit. Based on the third intermediate dimension, a compensation unit is used to perform compensation to obtain compensation coefficients. A support set is constructed using the adaptive state vectors of all historical moments before the preset window length and the corresponding compensation coefficients. A second objective function is constructed using the support set. A single gradient step is performed on the parameters of the adaptive layer using the second objective function to obtain the updated compensation coefficients. The updated compensation coefficients are multiplied element-wise with the corresponding displacement increment, swing angle increment, and speed increment in the cooperative control sequence to generate the drive command for the current moment.
2. The method according to claim 1, characterized in that, The process of inputting the current operation stage's stage identifier, the spatiotemporal state characteristics, the coaxiality deviation, the confidence coefficient, and the resistance data into the near-end strategy optimization network for strategy decision processing to obtain the corresponding stage's stage control instructions includes: The mapping unit of the network is optimized by using a near-end strategy to map the process identifier of the current operation process to the corresponding identifier embedding vector. The spatiotemporal state features are input into the first encoding unit of the near-end policy optimization network. The spatiotemporal state features are compressed to a first intermediate dimension through the compression layer of the first encoding unit. A first linear rectified function is applied to the first intermediate dimension through the suppression layer to obtain a first intermediate vector. The first intermediate vector is mapped to a state embedding vector through the mapping layer. The coaxiality deviation, the confidence coefficient, and the resistance data are input into the second encoding unit of the near-end policy optimization network. The coaxiality deviation, the confidence coefficient, and the resistance data are compressed to a second intermediate dimension through the compression layer of the second encoding unit. A second linear rectification function is applied to the second intermediate dimension through the suppression layer to obtain a second intermediate vector. The second intermediate vector is mapped to a deviation embedding vector through the mapping layer. The fusion unit of the network is optimized by a near-end strategy to fuse the identifier embedding vector, the state embedding vector, and the bias embedding vector to obtain a fused feature vector; The policy unit of the near-end policy optimization network is used to suppress and map the fused feature vector to obtain the control command under the current operation. Based on the control commands, the first position command of the hydraulic rod, the second position command of the hydraulic cylinder, the rotation angle command of the transverse hydraulic swing cylinder and the longitudinal hydraulic swing cylinder, and the speed command of the hydraulic motor are generated.
3. The method according to claim 2, characterized in that, The policy unit that optimizes the network through a near-end policy performs suppression and mapping processing on the fused feature vector to obtain control instructions for the current operation, including: The fused feature vector is input into the policy unit of the near-end policy optimization network, and the corresponding action head in the policy unit is activated according to the stage identifier of the current operation stage. The fused feature vector is mapped through the fully connected layer of the action head to obtain the original feature vector; The original feature vector is split into mean sub-vectors and variance sub-vectors in dimensional order. A Gaussian distribution is constructed based on the mean sub-vectors and variance sub-vectors. The original action vector with the same dimension as the mean sub-vector is obtained by sampling from the Gaussian distribution through the reparameterization method. Based on the preset allowable range of the current operation, the original motion vector is scaled to obtain the position increment command value, clamping force command value, rotation speed command value and angle command value; The position increment command value, the clamping force command value, the rotation speed command value, and the rotation angle command value are combined into a control command for the current operation.
4. The method according to claim 1, characterized in that, The process of mapping the six-dimensional force data, the position detection data, the three-dimensional point cloud data, and the acoustic emission waveform data into a dynamic graph structure whose node attributes are updated over time includes: The drive wheel, the gripper, the hydraulic rod, the hydraulic cylinder, the transverse hydraulic swing cylinder, and the longitudinal hydraulic swing cylinder are defined as corresponding mechanism nodes, the upper drill rod and the lower drill rod are defined as corresponding workpiece nodes, and the docking station is defined as an environmental node. The values of each node at the corresponding time in the six-dimensional force data, the position detection data, the three-dimensional point cloud data, and the acoustic emission waveform data are used as the load characteristics of the corresponding node. The clamping relationship between the mechanism node and the workpiece node, the motion constraint relationship between the mechanism node and the environment node, and the spatial positioning relationship between the workpiece node and the environment node are taken as an edge set. Each edge in the edge set is assigned the six-dimensional force data, the position detection data, the force value and relative distance value corresponding to the start and end times in the three-dimensional point cloud data. By combining the load characteristics, the edge set, the force values, the relative distance values, and the three-dimensional spatial coordinates of each node, a dynamic graph structure in which node attributes are updated over time is constructed.
5. The method according to claim 1, characterized in that, The step of aggregating neighborhood node features of the dynamic graph structure using a graph attention network to obtain spatiotemporal state features includes: By using the linear transformation layer of the graph attention network, all features in the dynamic graph structure are processed by nonlinear transformation to obtain the initial embedding vector of all nodes at the current time. The spatial attention coefficients of each node and its corresponding neighboring nodes on the edge set are calculated through the first graph attention layer of the graph attention network. The spatial attention coefficients are then weighted and summed with the initial embedding vector of the corresponding neighboring node to obtain the spatial aggregation features of each node. By using the fusion layer of the graph attention network, the spatial aggregation features are superimposed with the initial embedding vector to obtain the residual enhancement features of each node; Through the second graph attention layer of the graph attention network, multiple neighbor paths are defined for each node based on clamping relationships, motion constraint relationships, and spatial positioning relationships. For all the neighbor paths corresponding to each node, the path attention coefficients of each neighbor path are calculated based on the residual enhancement features. Based on all the path attention coefficients, the path collaboration features of each node are obtained. Using the path collaboration features of all historical moments before the preset window length as the query sequence, the temporal attention coefficients between different time steps in the query sequence are calculated through the temporal self-attention layer of the graph attention network according to the node type of each node, and the path collaboration features are weighted and summed based on the temporal attention coefficients to obtain the spatiotemporal state features at the current moment.
6. The method according to claim 1, characterized in that, The step of solving the coordinated control sequence of the hydraulic rod, hydraulic cylinder, lateral hydraulic swing cylinder, longitudinal hydraulic swing cylinder, and hydraulic motor within a future preset time domain based on the link control command includes: Using the displacement of the hydraulic rod, the displacement of the hydraulic cylinder, the swing angle of the lateral hydraulic swing cylinder, the swing angle of the longitudinal hydraulic swing cylinder, and the rotational speed of the hydraulic motor as state variables, and the increment corresponding to each displacement as control variables, a linear prediction model between the state variables and the control variables is established at each time in a future preset time domain. The target position increment, target rotation angle, and target rotation speed in the control command are converted into the expected sequence of each moment in the future preset time domain; The tracking cost is the sum of squared deviations between the expected sequence and the predicted sequence output by the linear prediction model; the energy cost is the sum of squared control variables; and the smoothing cost is the sum of squared differences between control variables at adjacent time points. The weighted summation formula of the tracking cost, the energy cost, and the smoothing cost serves as the first objective function. The preset set of process constraints and the preset obstacle avoidance constraints are uniformly represented as a set of linear inequalities, and iteratively solved using a sequential quadratic programming solver until the first objective function reaches the preset convergence condition. The first displacement increment of the hydraulic rod, the second displacement increment of the hydraulic cylinder, the swing angle increment of the lateral hydraulic swing cylinder, the swing angle increment of the longitudinal hydraulic swing cylinder, and the speed increment of the hydraulic motor when the convergence condition is met are combined into a cooperative control sequence for the current moment.
7. The method according to claim 1, characterized in that, The step of calculating the coaxiality deviation between the gripper and the lower drill pipe and the corresponding confidence coefficient based on the three-dimensional point cloud data includes: The three-dimensional point cloud data is segmented into point clouds at the upper drill rod end and point clouds at the lower drill rod end, and the overlapping area of the point clouds at the upper drill rod end and the lower drill rod end is identified. Based on the overlapping area, an overlap metric value is obtained. Based on the overlapping area, the rotational deviation and translational deviation of the gripper relative to the lower drill pipe are calculated, and the rotational deviation and translational deviation are combined into a coaxiality deviation. By combining the overlap metric, the coaxiality deviation, and the point cloud density value at the end of the upper drill pipe, the first fluctuation range of the rotational deviation and the second fluctuation range of the translational deviation are obtained. Calculate the confidence coefficient based on the first fluctuation range and the second fluctuation range.
8. The method according to claim 1, characterized in that, The step of identifying the resistance data of the drive wheel based on the six-dimensional force data includes: Extract the axial force and tangential force values of the drive wheel acting on the drill pipe from the six-dimensional force data, calculate the ratio of the axial force value to the tangential force value, and obtain the instantaneous friction coefficient. The vector synthesis result of the axial force value and the tangential force value is taken as the total contact force between the drive wheel and the drill pipe contact surface; When the current operation is a screw-in connection, the projection component of the total contact force along the drill pipe axis is marked as the screw-in resistance; or, when the current operation is pipe unloading and disassembly, the projection component of the total contact force along the drill pipe tangential direction is marked as the screw-out resistance. The inward or outward resistance is associated with the instantaneous friction coefficient and stored to generate resistance data for the drive wheel.
9. A multi-mechanism collaborative adaptive loading and unloading integrated device for drill pipes, characterized in that, include: The frame (10), drive mechanism (20), clamping mechanism (30) and industrial AI controller (40) are provided on the frame (10) for executing the drill pipe multi-mechanism cooperative adaptive loading and unloading control method as described in any one of claims 1 to 8; The telescopic end of the hydraulic cylinder (23) is fixed to the frame (10). The drive mechanism (20) is installed on one side of the frame (10). The drive mechanism (20) includes a swing assembly (24). The swing assembly (24) includes a transverse hydraulic swing cylinder (241) fixedly installed on the top of the mounting frame (22). The drive shaft of the transverse hydraulic swing cylinder (241) is fixed with a rotating shaft (242). A swing arm (243) rotatably connected to the mounting base (31) is fixed on the rotating shaft (242). A longitudinal hydraulic swing cylinder (37) with the drive shaft fixed to the mounting base (31) is fixedly installed at one end of the swing arm (243). The clamping mechanism (30) includes a mounting base (31), on which two grippers (33) are rotatably connected, and a hydraulic rod (34) is rotatably connected between the two grippers (33). A connecting component (36) is provided on the grippers (33). The connecting assembly (36) includes a hydraulic motor (362) fixedly installed on one side of the gripper (33). Two drive wheels (361) are rotatably connected to the other side of the gripper (33), and one of the drive wheels (361) is fixed to the drive shaft of the hydraulic motor (362). The hydraulic motor (362) is started to screw the upper drill rod into the lower drill rod through the drive wheel (361).
Citation Information
Patent Citations
Resistor disc defect online detection system and grading method based on machine vision
CN120765532A
Sea area cross-medium unmanned system task allocation method based on graph attention network and deep reinforcement learning, and electronic equipment
CN121168911A