Well repair robot control algorithm fusion method

By using an improved YOLOv7 and Res-Net-50 multi-branch structure, combined with a dynamic weight fusion network and the A* algorithm, an interactive 3D model is generated, which solves the accuracy and stability issues of detection and grasping of well repair robots in underground environments and achieves efficient path planning and grasping control.

CN120606386AActive Publication Date: 2025-09-09QINGDAO BEIHAI JUNHUI ELECTRONIC INSTR CO LTD

Patent Information

Application Number
CN202510699155.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-09
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing well repair robots have difficulty accurately detecting target objects and extracting multi-dimensional features, have difficulty efficiently fusing multimodal data to generate interactive three-dimensional models, lack the ability to combine multiple algorithms to generate safe paths and dynamically avoid obstacles, and lack the ability to quantify multiple indicators and perform comprehensive scoring to determine whether the target object needs to be recaptured.

Method used

The improved YOLOv7 and Res-Net-50 multi-branch structure, combined with a dynamic weight fusion network and the A* algorithm, uses high-resolution cameras, laser rangefinders, and IMU sensors to generate interactive 3D models, plan safe paths, and perform grasping control. Multi-sensor data is combined to generate a working environment map, define the wellhead and robot coordinate systems, adjust the trajectory in real time, and judge the success of the grasping through comprehensive indicator scoring.

Benefits of technology

It improves the accuracy of downhole detection, enhances the accuracy of downhole tool classification and operational efficiency, achieves high-precision path planning and grasping control, solves the problem of inaccurate single-modal data modeling in complex downhole scenarios, and ensures the stability of grasping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120606386A_ABST
    Figure CN120606386A_ABST
Patent Text Reader

Abstract

The invention discloses a well repair robot control algorithm fusion method, and relates to the field of oil and gas field well repair. Comprising the following steps: collecting a video stream through a high-resolution camera, carrying out target detection and multi-dimensional feature extraction by using an improved YOLOv7 and Res-Net-50 network, and generating an interactive three-dimensional model in combination with a classifier result; fusing the point cloud density and the image texture features, and training a lightweight semantic segmentation model to realize a target object category; selecting an optimal grabbing point and strategy according to categories and features, constructing a working environment map by combining an A * algorithm and multi-sensor data, and planning a safe path; dynamic track adjustment and pose control are realized through coordinate system alignment and real-time pose feedback; and finally, according to the grabbing strategy, all finger joints of the mechanical arm are controlled to grab the target object, the comprehensive index influence score is calculated, whether grabbing needs to be conducted again or not is judged, accurate grabbing of the robot is controlled, and the stability and reliability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of oil and gas field well repair, and in particular relates to a control algorithm fusion method for a well repair robot. Background Art

[0002] In the field of oil and gas field well repair, the traditional model adopts a workover rig + wellhead operator model. The wellhead operators need to push or support or receive pipe rods and elevators, operate the opening and closing of elevators, fasten pipe rods, operate hydraulic pliers to make or break out the buckles, operate slips, etc. Usually 2-3 workers are needed to cooperate, the labor intensity is high, and the wellhead equipment is numerous and bulky, which can easily lead to safety accidents and casualties.

[0003] In recent years, although wellhead automation equipment consisting of robotic arms, hydraulic flip elevators, automatic hydraulic clamps, etc. has emerged to replace the labor of wellhead workers, these devices have defects such as complexity and dispersion, high cost, frequent failures, and low efficiency in mutual coordination. It is difficult to accurately detect target objects and extract multi-dimensional features, and it is also difficult to efficiently fuse multimodal data and generate interactive three-dimensional models. There is a lack of combining multiple algorithms to generate safe paths and dynamically avoid obstacles, and there is a lack of comprehensive scoring after quantifying multiple indicators to determine whether the target object needs to be re-grabbed.

[0004] Therefore, a control algorithm fusion method for well workover robots came into being. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a control algorithm fusion method for a well repair robot to solve the following technical problems:

[0006] It is difficult to accurately detect target objects and extract multi-dimensional features, and it is also difficult to efficiently fuse multimodal data and generate interactive three-dimensional models. There is a lack of combining multiple algorithms to generate safe paths and dynamically avoid obstacles. There is a lack of comprehensive scoring after quantifying multiple indicators to determine whether the target object needs to be recaptured.

[0007] To solve the above problems, the present invention provides a control algorithm fusion method for a well repair robot, comprising the following steps:

[0008] S1: The video stream is collected by a high-resolution camera for preprocessing, the target area is detected using the improved YOLOv7, and a multi-branch structure is designed based on the Res-Net-50 network to extract the multi-dimensional features of the target object;

[0009] S2: Fuse the YOLOv7 object detection results, Res-Net-50 image features, and classifier recognition results to generate an interactive 3D model.

[0010] S3: Design a dynamic weight fusion network to assign modal weights based on point cloud density and image texture confidence, generate a joint feature representation, and train a lightweight semantic segmentation model based on the joint features to distinguish different target object categories;

[0011] S4: Select the optimal grasping point and grasping strategy based on the target object category and characteristics, plan a safe path based on the A* algorithm and obstacle expansion processing, fuse multi-sensor data to generate a work environment map, and mark the obstacle type;

[0012] S5: Define the wellhead center coordinate system and the robot coordinate system, calculate the pose and joint angles, obtain the real-time pose through the IMU and camera, calculate the error and dynamically adjust the trajectory, and perform pose control and trajectory correction based on sensor feedback;

[0013] S6: Based on the grasping planning results, control the finger joints of the manipulator to grasp the target object. By normalizing the contact force, posture change, visual inspection results and stability index, and then adding the weighted comprehensive index influence scores, it is determined whether re-grasping is needed.

[0014] Preferably, the step S1 includes the following steps:

[0015] Use a high-resolution camera to capture real-time video streams of the workspace and perform preprocessing on the captured images, including denoising, color correction, and white balancing.

[0016] YOLOv7 was selected as the basic object detection model. Deformable convolutions were introduced into its E-ELAN backbone network. A coordinate attention mechanism was added to the detection head. Gaussian heatmap regression was used instead of traditional bounding box regression. Feature maps of different scales were dynamically fused using the BiFPN weighted bidirectional feature pyramid.

[0017] Res-Net-50 is selected as the basic network, CBAM is embedded in the residual block, and a multi-branch network is designed to extract the shape topology structure, RGB-HSV color space features and LBP texture features of the object respectively. The features extracted by different branches are fused to form a multi-dimensional semantic information representation.

[0018] Preferably, the step S2 comprises the following steps:

[0019] The target area detected by YOLOv7 is input into the Res-Net-50 network enhanced by the attention mechanism to extract the multi-dimensional features of the target object;

[0020] Based on the extracted multi-dimensional features, a classifier is used to classify and identify the target object;

[0021] The sub-pixel bounding box output by YOLOv7 is modified twice, and the multi-dimensional semantic information extracted by ResNet-50 is integrated with the target detection results to form a complete target description, including geometric attributes: bounding box coordinates and shape outline; appearance attributes: color histogram and LBP texture features; semantic attributes: category label and multi-dimensional semantic vector;

[0022] The integrated target description is output in the form of an interactive three-dimensional model.

[0023] Preferably, the step S3 comprises the following steps:

[0024] In the working space, the target object is scanned by a laser rangefinder to obtain the original point cloud data and pre-process the point cloud data;

[0025] Align the point cloud with the high-resolution image through a time synchronization mechanism and unify the coordinate system based on the IMU attitude data;

[0026] Extract features from point cloud data, including normal vectors, curvature, and local surface features, and perform adaptive downsampling on high-density point clouds;

[0027] Combined with the Res-Net-50 network to extract the multi-dimensional features of the target object;

[0028] Design a transmembrane feature projection module: map the multi-dimensional features extracted by the Res-Net-50 network to a point cloud coordinate system and perform spatial alignment through feature matching;

[0029] Design a dynamic weight fusion network to dynamically assign modal weights based on point cloud density and image texture confidence to generate a joint feature representation;

[0030] A lightweight semantic segmentation model is trained based on joint features. The model outputs multi-category confidence for each voxel and assigns a semantic label to each voxel based on the confidence output by the model.

[0031] The radiation field of the voxel is fused with the spatial position characteristics and image characteristics of the voxel to generate the radiation intensity or color of the voxel.

[0032] Preferably, the pre-processing of the point cloud data comprises the following steps:

[0033] Use PCL's voxelGrid to calculate the point density of each voxel, set high-density thresholds and low-density thresholds; recursively subdivide octree nodes for high-density areas, merge adjacent voxels for low-density areas; and construct a KD-tree inside the octree node for fast neighborhood search;

[0034] Huffman encoding of voxel semantic labels;

[0035] Adjust the filter strength according to the voxel semantic label and use bilateral filtering to preserve the sharpness of the voxel boundary. Specifically:

[0036]

[0037] Among them, W(p) is the weight value of point p, p is the position vector of the current point, and p c is the position vector of the center point, I(p) and I(p c ) are point p and point p respectively c The intensity value, σ d is the scale parameter of the spatial domain, σ r is the scale parameter of the intensity domain;

[0038] An axial attention module is introduced to calculate the position of the point cloud projected onto the image plane and align it with the image segmentation mask.

[0039] Preferably, the step of fusing the spatial position characteristics and image characteristics of voxels through the neural network radiation field to generate the radiation intensity or color of the voxels comprises the following steps:

[0040] The spatial position features of the voxels and the color and texture features extracted by the ResNet-50 network are input into the neural network radiation field to generate the radiation intensity or color of the voxels, specifically:

[0041] c=σ(ω1·P_f+ω2·I_f)

[0042] Where c is the radiation intensity or color of the voxel, P_f is the spatial position feature, I_f is the image feature, including color and texture features, σ is the activation function, ω1 and ω2 are the corresponding weights;

[0043] Match the material library according to voxel semantic labels and use UV mapping to map the material to the 3D model surface.

[0044] Preferably, the step S4 comprises the following steps:

[0045] Based on the 3D model and voxel semantic labels of the target object, the optimal grasping point is selected and a grasping strategy is formulated: for regular large smooth objects, a multi-point grasping strategy is selected with evenly distributed grasping points; for irregular small rough objects, a single-point grasping strategy is selected with the geometric center or center of gravity as the grasping point;

[0046] Plan the robot's path from its current position to the target grasping point using the A* algorithm;

[0047] Use cameras and lidar to obtain information about the robot's working environment, align point cloud data with high-resolution images, generate a working environment map, and combine semantic segmentation results to mark the types of obstacles;

[0048] During the A* algorithm path planning process, obstacles are regarded as impassable nodes, and the obstacle area is expanded using the expansion operation.

[0049] Preferably, the step S5 comprises the following steps:

[0050] Define the wellhead center coordinate system and the robot coordinate system, and align the origin of the robot coordinate system with the wellhead center;

[0051] Use DH parameters to describe the joint and link structure of the robot, calculate the pose of the end effector based on the joint angle, and calculate the joint angle based on the target pose;

[0052] The joint angle is obtained by numerical iteration method, and the optimal solution is selected according to actual needs. i , i = 1, 2, ..., n, n is the number of joints;

[0053] Discretize the joint angle sequence, assign it to the time axis, use the interpolation method to generate a smooth trajectory, map the joint angle sequence to the time axis, and generate a trajectory with continuous velocity and acceleration;

[0054] The real-time pose of the end effector is obtained through the IMU sensor, the error between the real-time pose and the target pose is calculated, and the joint angle is adjusted according to the error; at the same time, the image of the grasping point is obtained using a camera, the real-time pose of the grasping point is estimated through a visual algorithm, and the trajectory is adjusted according to the visual feedback.

[0055] Preferably, the steps of defining the wellhead center coordinate system and the robot coordinate system and aligning the origin of the robot coordinate system with the wellhead center include the following steps:

[0056] Define the wellhead center coordinate system: the origin is located at the wellhead center, the Z axis is perpendicular to the wellhead plane and points upward, and the X axis and Y axis are in the horizontal direction respectively;

[0057] Define the robot coordinate system: the origin is located at the robot base, the Z axis is along the robot's main axis, and the X and Y axes point to the front and back and left and right directions of the robot respectively;

[0058] According to the current position and posture of the robot, the transformation matrix between the robot coordinate system and the wellhead center coordinate system is calculated, and the position and posture of the grasping point are transformed from the robot coordinate system to the wellhead center coordinate system. Specifically:

[0059] P w =T r_w ·P r

[0060] Among them, P w is the wellhead center coordinate system, T r_w is the coordinate system transformation matrix, P ris the robot coordinate system.

[0061] Preferably, the step S6 comprises the following steps:

[0062] According to the results of the grasping plan, the finger joints of the manipulator are controlled to grasp the target object;

[0063] Force sensors are installed at the ends of the manipulator's finger joints to measure contact force in real time. An IMU sensor is installed on the target object to measure the object's acceleration and angular velocity in real time. The displacement change is calculated through integration, and a camera is used to capture an image of the target object. The pose change is calculated through feature point matching. A high-resolution camera is installed above the manipulator and a deep learning algorithm is used to detect whether the target object is in the manipulator. If the target object is in the manipulator, the binary output is 1, otherwise it is 0. The stability index is calculated based on the contact force distribution of each finger joint and the center of gravity of the object.

[0064] Normalize the value of each indicator and assign a weight to each indicator. The total weight is 1. The weighted sum is added to obtain the comprehensive indicator impact score, which is as follows:

[0065]

[0066] Among them, E is the comprehensive indicator impact score, F ′ , ΔP ′ 、l ′ 、S ′ and M′ are the normalized contact force, posture change, visual inspection results and stability index, ω F 、ω p 、ω l and ω are the corresponding weights respectively;

[0067] If E ≥ threshold, the crawling is successful and no re-crawl is required; if E < threshold, the crawling fails, and the crawling strategy is replanned and the crawling is attempted again.

[0068] Beneficial effects of the present invention:

[0069] By combining the improved YOLOv7 with the Res-Net-50 multi-branch structure, this paper can effectively extract the multi-dimensional features of the target object, improving the detection accuracy in low-light and high-noise environments underground. At the same time, by fusing the YOLOv7 detection results, Res-Net-50 image features and classifier results, an interactive 3D model is generated, solving the problem of inaccurate modeling of single-modal data in complex underground scenes and supporting subsequent path planning and grasping control.

[0070] The present invention dynamically assigns modal weights through point cloud density and image texture confidence, generates joint feature representation, trains a lightweight semantic segmentation model, and improves the accuracy of downhole tool classification.

[0071] This invention generates a working environment map by fusing multi-sensor data based on the A* algorithm, marks obstacle types, and plans collision-free paths, thereby improving the operating efficiency of the well repair robot. At the same time, by defining the wellhead center coordinate system and the robot coordinate system, combining the IMU and camera to obtain the position and posture in real time, calculate errors, and dynamically adjust the trajectory, high-precision positioning and control are achieved.

[0072] The present invention generates a comprehensive score by weighted addition of normalized contact force, posture change, visual detection results and stability index to determine whether re-grasping is needed, thereby solving the problem of unstable grasping of downhole tools caused by irregular shape or heavy weight. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 Schematic diagram of the method flow of the present invention;

[0074] Figure 2 A schematic diagram of the interactive three-dimensional modeling process of the present invention;

[0075] Figure 3 This is a schematic diagram of the multimodal perception weighted decision-making and evaluation process of the present invention. DETAILED DESCRIPTION

[0076] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0077] See also Figure 1 As shown, the present invention is a control algorithm fusion method for a well repair robot, comprising the following steps:

[0078] S1: The video stream is collected by a high-resolution camera for preprocessing, the target area is detected using the improved YOLOv7, and a multi-branch structure is designed based on the Res-Net-50 network to extract the multi-dimensional features of the target object;

[0079] S2: Fuse the YOLOv7 object detection results, Res-Net-50 image features, and classifier recognition results to generate an interactive 3D model.

[0080] S3: Design a dynamic weight fusion network to assign modal weights based on point cloud density and image texture confidence, generate a joint feature representation, and train a lightweight semantic segmentation model based on the joint features to distinguish different target object categories;

[0081] S4: Select the optimal grasping point and grasping strategy based on the target object category and characteristics, plan a safe path based on the A* algorithm and obstacle expansion processing, fuse multi-sensor data to generate a work environment map, and mark the obstacle type;

[0082] S5: Define the wellhead center coordinate system and the robot coordinate system, calculate the pose and joint angles, obtain the real-time pose through the IMU and camera, calculate the error and dynamically adjust the trajectory, and perform pose control and trajectory correction based on sensor feedback;

[0083] S6: Based on the grasping planning results, control the finger joints of the manipulator to grasp the target object. By normalizing the contact force, posture change, visual inspection results and stability index, and then adding the weighted comprehensive index influence scores, it is determined whether re-grasping is needed.

[0084] In one embodiment of the present invention, step S1 includes the following steps:

[0085] Use a high-resolution camera to capture real-time video streams of the workspace and perform preprocessing on the captured images, including denoising, color correction, and white balancing.

[0086] YOLOv7 was selected as the basic object detection model. Deformable convolutions were introduced into its E-ELAN backbone network. A coordinate attention mechanism was added to the detection head. Gaussian heatmap regression was used instead of traditional bounding box regression. Feature maps of different scales were dynamically fused using the BiFPN weighted bidirectional feature pyramid.

[0087] Res-Net-50 is selected as the basic network, CBAM is embedded in the residual block, and a multi-branch network is designed to extract the shape topology structure, RGB-HSV color space features and LBP texture features of the object respectively. The features extracted by different branches are fused to form a multi-dimensional semantic information representation.

[0088] Specifically, Gaussian filtering or bilateral filtering is used to remove image noise, color deviation is corrected through the grayscale world hypothesis or white balance algorithm, and the image color temperature is adjusted using the adaptive white balance algorithm; deformable convolution is inserted into the convolution layer of the E-ELAN in the YOLOv7-based target detection model; a coordinate attention mechanism is added to the detection head, and the CA module is used to weight the feature map; the traditional bounding box regression is replaced by Gaussian heat map regression to generate a Gaussian heat map of the target center point and predict the center coordinates and size; a weighted bidirectional feature pyramid is used to dynamically fuse feature maps of different scales, and high-level features and low-level features are weightedly fused to improve multi-scale target detection capabilities; Res-Net-50 is used as the backbone network to extract preliminary features. The multi-branch design includes using edge detection or skeleton extraction algorithms to extract shape features, converting RGB images to HSV color space, extracting color histogram features, and using the local binary pattern algorithm to extract texture features; the features extracted by each branch, including shape, color and texture, are spliced ​​or weightedly fused to form a multi-dimensional semantic representation, and channel and spatial attention modules are embedded in the residual block to improve feature selection capabilities.

[0089] In one embodiment of the present invention, step S2 includes the following steps:

[0090] The target area detected by YOLOv7 is input into the Res-Net-50 network enhanced by the attention mechanism to extract the multi-dimensional features of the target object;

[0091] Based on the extracted multi-dimensional features, a classifier is used to classify and identify the target object;

[0092] The sub-pixel bounding box output by YOLOv7 is modified twice, and the multi-dimensional semantic information extracted by ResNet-50 is integrated with the target detection results to form a complete target description, including geometric attributes: bounding box coordinates and shape outline; appearance attributes: color histogram and LBP texture features; semantic attributes: category label and multi-dimensional semantic vector;

[0093] The integrated target description is output in the form of an interactive three-dimensional model.

[0094] Specifically, according to the bounding box coordinates output by YOLOv7, the target area is cropped from the original image and input into the Res-Net-50 network enhanced by the attention mechanism. Multi-dimensional features are extracted from the output of Res-Net-50, and the output features of Res-Net-50 are used as input for classification through the fully connected layer. The multi-dimensional features are input into the classifier and the category label is output. The edge features or shape contours extracted by Res-Net-50 are used to perform a secondary correction on the bounding box of YOLOv7. The bounding box coordinates of YOLOv7 and the multi-dimensional features extracted by Res-Net-50 are fused to form a complete target description. A three-dimensional model of the target object is generated according to the geometric properties.

[0095] In one embodiment of the present invention, step S3 includes the following steps:

[0096] In the working space, the target object is scanned by a laser rangefinder to obtain the original point cloud data and pre-process the point cloud data;

[0097] Align the point cloud with the high-resolution image through a time synchronization mechanism and unify the coordinate system based on the IMU attitude data;

[0098] Extract features from point cloud data, including normal vectors, curvature, and local surface features, and perform adaptive downsampling on high-density point clouds;

[0099] Combined with the Res-Net-50 network to extract the multi-dimensional features of the target object;

[0100] Design a transmembrane feature projection module: map the multi-dimensional features extracted by the Res-Net-50 network to a point cloud coordinate system and perform spatial alignment through feature matching;

[0101] Design a dynamic weight fusion network to dynamically assign modal weights based on point cloud density and image texture confidence to generate a joint feature representation;

[0102] A lightweight semantic segmentation model is trained based on joint features. The model outputs multi-category confidence for each voxel and assigns a semantic label to each voxel based on the confidence output by the model.

[0103] The radiation field of the voxel is fused with the spatial position characteristics and image characteristics of the voxel to generate the radiation intensity or color of the voxel.

[0104] Specifically, the point cloud frames and image frames are aligned by timestamps, and the point cloud and image are converted to the global coordinate system using the posture number of the IMU; the normal vector is calculated using PCA or neighborhood fitting plane, the curvature is estimated by the neighborhood point density, FPFH features are extracted as local surface features, and Res-Net-50 is used to extract the multi-dimensional features of the image; a cross-modal feature projection module is designed: the image features are projected to the point cloud coordinate system through bilinear interpolation or spatial transformation, and the point cloud and image features are aligned through nearest neighbor search or feature similarity, such as cosine similarity; the confidence is calculated based on the local density of the point cloud, and the confidence is calculated through the feature response of Res-Net-50, and the weights of the point cloud and image are dynamically assigned according to the confidence; the point cloud is converted into a voxel grid, each voxel contains joint features, and the voxel features are processed using a 3D convolutional network. According to the confidence output by the model, a semantic label is assigned to each voxel, and the model is optimized using cross-entropy loss.

[0105] In one embodiment of the present invention, the pre-processing of the point cloud data includes the following steps:

[0106] Use PCL's voxelGrid to calculate the point density of each voxel, set high-density thresholds and low-density thresholds; recursively subdivide octree nodes for high-density areas, merge adjacent voxels for low-density areas; and construct a KD-tree inside the octree node for fast neighborhood search;

[0107] Huffman encoding of voxel semantic labels;

[0108] Adjust the filter strength according to the voxel semantic label and use bilateral filtering to preserve the sharpness of the voxel boundary. Specifically:

[0109]

[0110] Among them, W(p) is the weight value of point p, p is the position vector of the current point, and p c is the position vector of the center point, I(p) and I(p c ) are point p and point p respectively c The intensity value, σ d is the scale parameter of the spatial domain, σ r is the scale parameter of the intensity domain;

[0111] An axial attention module is introduced to calculate the position of the point cloud projected onto the image plane and align it with the image segmentation mask.

[0112] Specifically, PCL's VoxelGrid is used to calculate the point density of each voxel, high-density threshold and low-density threshold are set, octree nodes are recursively subdivided for high-density voxels, adjacent voxels are merged for low-density voxels, and a KD-tree is constructed inside the octree node for fast neighborhood search; the frequency of occurrence of each semantic label is counted, a Huffman tree is constructed based on the label frequency, the Huffman tree is traversed, the Huffman code of each label is generated, and the voxel label is replaced with the Huffman code; the bilateral filtering formula is used to adjust the weight value of the point, each point is traversed, its weight value is calculated and the position is updated; the point cloud is projected onto the image plane, the pixel coordinates are obtained, the axial attention module is designed, the alignment weight of the point cloud and the image segmentation mask is calculated, and the point cloud features are aligned with the image segmentation mask.

[0113] In one embodiment of the present invention, the method of fusing the spatial position characteristics and image characteristics of voxels through a neural network radiation field to generate the radiation intensity or color of the voxels includes the following steps:

[0114] The spatial position features of the voxels and the color and texture features extracted by the ResNet-50 network are input into the neural network radiation field to generate the radiation intensity or color of the voxels, specifically:

[0115] c=σ(ω1·P_f+ω2·I_f)

[0116] Where c is the radiation intensity or color of the voxel, P_f is the spatial position feature, I_f is the image feature, including color and texture features, σ is the activation function, ω1 and ω2 are the corresponding weights;

[0117] Match the material library according to voxel semantic labels and use UV mapping to map the material to the 3D model surface.

[0118] Specifically, PCL's VoxelGrid is used to divide the point cloud into a voxel grid, and the center coordinates of each voxel are recorded as the spatial position feature x. The color and texture features of the image are extracted using the ResNet-50 model. A simple neural network is designed to combine the extracted feature f with the spatial position feature x of the voxel as input. The radiation intensity or color is output, and the activation function uses ReLU with weights ω1 and ω2 of 0.4 and 0.6 respectively. The true radiation intensity or color of the voxel is used as a supervisory signal to train the network. According to the semantic label of the voxel, the corresponding material map is selected from the material library, and the material map is loaded using OpenCV or PIL. The UV coordinates are calculated for each vertex of the 3D model, and the material map is mapped to the 3D model surface according to the UV coordinates.

[0119] In one embodiment of the present invention, step S4 includes the following steps:

[0120] Based on the 3D model and voxel semantic labels of the target object, the optimal grasping point is selected and a grasping strategy is formulated: for regular large smooth objects, a multi-point grasping strategy is selected with evenly distributed grasping points; for irregular small rough objects, a single-point grasping strategy is selected with the geometric center or center of gravity as the grasping point;

[0121] Plan the robot's path from its current position to the target grasping point using the A* algorithm;

[0122] Use cameras and lidar to obtain information about the robot's working environment, align point cloud data with high-resolution images, generate a working environment map, and combine semantic segmentation results to mark the types of obstacles;

[0123] During the A* algorithm path planning process, obstacles are regarded as impassable nodes, and the obstacle area is expanded using the expansion operation.

[0124] Specifically, based on the three-dimensional model and voxel semantic labels of the target object, the object is classified as "regular large smooth object" or "irregular small rough object". For regular large smooth objects: a multi-point grasping strategy is used to evenly distribute the grasping points, calculate the surface normal vector of the object, and select the point whose normal vector direction is suitable for grasping; for irregular small rough objects, a single-point grasping strategy is used, with the geometric center or center of gravity as the grasping point; the optimal grasping point and grasping strategy are output based on the classification results; cameras and lidars are used to obtain working environment information, point cloud data is aligned with high-resolution images, and a working environment map is generated. Combined with the semantic segmentation results, the type of obstacles is marked. In the A* algorithm, obstacles are regarded as impassable nodes, and the expansion operation is used to expand the obstacle area and increase the safety distance; the robot is controlled to move along the planned path to the target grasping point, and the grasping action is performed according to the grasping strategy.

[0125] In one embodiment of the present invention, step S5 includes the following steps:

[0126] Define the wellhead center coordinate system and the robot coordinate system, and align the origin of the robot coordinate system with the wellhead center;

[0127] Use DH parameters to describe the joint and link structure of the robot, calculate the pose of the end effector based on the joint angle, and calculate the joint angle based on the target pose;

[0128] The joint angle is obtained by numerical iteration method, and the optimal solution is selected according to actual needs. i , i = 1, 2, ..., n, n is the number of joints;

[0129] Discretize the joint angle sequence, assign it to the time axis, use the interpolation method to generate a smooth trajectory, map the joint angle sequence to the time axis, and generate a trajectory with continuous velocity and acceleration;

[0130] The real-time pose of the end effector is obtained through the IMU sensor, the error between the real-time pose and the target pose is calculated, and the joint angle is adjusted according to the error; at the same time, the image of the grasping point is obtained using a camera, the real-time pose of the grasping point is estimated through a visual algorithm, and the trajectory is adjusted according to the visual feedback.

[0131] Specifically, with the wellhead center as the origin, a wellhead center coordinate system is established, and the Z axis is defined to be perpendicular to the wellhead plane and upward, and the X axis and Y axis are respectively in the horizontal direction; with the robot base as the origin, a robot coordinate system is established, and the Z axis is defined to be along the robot main axis direction, and the X axis and Y axis are respectively pointing to the front and back and left and right directions of the robot. According to the current position and posture of the robot, the transformation matrix between the robot coordinate system and the wellhead center coordinate system is calculated, and the posture of the grasping point is transformed from the robot coordinate system to the wellhead center coordinate system; define the DH parameter table, fill in the robot's connecting rod length ai, connecting rod torsion angle αi, joint offset di and joint angle θi, and use the DH parameter table to fill in the robot's connecting rod length ai, connecting rod torsion angle αi, joint offset di and joint angle θi. The transformation matrix of each joint is calculated step by step by numerical methods, and the posture of the end effector is finally obtained. The joint angle is solved using inverse kinematics; the joint angle sequence is discretized and assigned to the time axis, and interpolation methods such as polynomial interpolation or spline interpolation are used to generate a smooth trajectory. The joint angle trajectory is differentiated to obtain the velocity and acceleration; the real-time posture of the end effector is obtained through the IMU, the error between the real-time posture and the target posture is calculated, and the joint angle is adjusted according to the error; the image of the grasping point is obtained using a camera, the real-time posture of the grasping point is estimated through a visual algorithm, and the trajectory is adjusted according to the visual feedback; the joint angle trajectory is adjusted in real time by combining the IMU and visual feedback.

[0132] In one embodiment of the present invention, defining the wellhead center coordinate system and the robot coordinate system, and aligning the origin of the robot coordinate system with the wellhead center, includes the following steps:

[0133] Define the wellhead center coordinate system: the origin is located at the wellhead center, the Z axis is perpendicular to the wellhead plane and points upward, and the X axis and Y axis are in the horizontal direction respectively;

[0134] Define the robot coordinate system: the origin is located at the robot base, the Z axis is along the robot's main axis, and the X and Y axes point to the front and back and left and right directions of the robot respectively;

[0135] According to the current position and posture of the robot, the transformation matrix between the robot coordinate system and the wellhead center coordinate system is calculated, and the position and posture of the grasping point are transformed from the robot coordinate system to the wellhead center coordinate system. Specifically:

[0136] P w =T r_w ·P r

[0137] Among them, Pw is the wellhead center coordinate system, T r_w is the coordinate system transformation matrix, P r is the robot coordinate system.

[0138] In one embodiment of the present invention, step S6 includes the following steps:

[0139] According to the results of the grasping plan, the finger joints of the manipulator are controlled to grasp the target object;

[0140] Force sensors are installed at the ends of the manipulator's finger joints to measure contact force in real time. An IMU sensor is installed on the target object to measure the object's acceleration and angular velocity in real time. The displacement change is calculated through integration, and a camera is used to capture an image of the target object. The pose change is calculated through feature point matching. A high-resolution camera is installed above the manipulator and a deep learning algorithm is used to detect whether the target object is in the manipulator. If the target object is in the manipulator, the binary output is 1, otherwise it is 0. The stability index is calculated based on the contact force distribution of each finger joint and the center of gravity of the object.

[0141] Normalize the value of each indicator and assign a weight to each indicator. The total weight is 1. The weighted sum is added to obtain the comprehensive indicator impact score, which is as follows:

[0142]

[0143] Among them, E is the comprehensive indicator impact score, F ′ , ΔP ′ 、l ′ 、S ′ and M′ are the normalized contact force, posture change, visual inspection results and stability index, ω F 、ω p 、ω l and ω are the corresponding weights respectively;

[0144] If E ≥ threshold, the crawling is successful and no re-crawl is required; if E < threshold, the crawling fails, and the crawling strategy is replanned and the crawling is attempted again.

[0145] Specifically, a grasping strategy is generated based on the shape, size, weight, and posture of the target object, including the optimal grasping point and grasping strategy. The robot kinematics and dynamics models are used to plan the motion trajectory of the manipulator. The grasping planning results are sent to the manipulator controller, which controls the movement of each finger joint according to the planned trajectory to complete the grasping of the target object. Force sensors are installed at the ends of the finger joints of the manipulator to collect contact force data of each finger joint in real time. IMU sensors are installed on the target object to collect acceleration and angular velocity data of the target object in real time. The displacement change is calculated by integration. A high-resolution camera is installed above the manipulator to capture the image of the target object in real time. A deep learning algorithm is used to detect whether the target object is in the manipulator and the binary output is: if the target object is in the manipulator, the output is L=1; otherwise, the output is L=0. The contact force distribution of each finger joint and the center of gravity position of the target object are obtained, and the uniformity of the contact force distribution is calculated. Specifically, Among them, S is the force uniformity, F i is the contact force of the i-th finger joint, is the average contact force, N is the number of finger joints; obtain the mass distribution of the target object, calculate its center of gravity and center of gravity coordinates by geometric methods, project the center of gravity coordinates onto the palm plane to obtain the projection point, if the palm plane is tilted, it is necessary to project the center of gravity onto the palm plane through coordinate transformation, the stable area is usually a polygonal area of ​​the palm of the manipulator (such as a rectangle, circle or custom shape), which is determined by the structure of the manipulator. For example, for a parallel gripping manipulator, the stable area may be the center area between two fingers; for a multi-finger manipulator, the stable area may be a polygon surrounded by each finger joint, and different judgment methods are used to determine whether the projection is within the stable area, such as geometric judgment, defining the stable area as a set of polygon vertices, and using the ray intersection method or the point in polygon algorithm to determine whether the projection point is within the polygon; if the projection point is within the stable area, it is assigned a value of 1, otherwise it is assigned a value of 0; the above indicators are normalized, and each indicator is given a weighted sum to obtain a comprehensive indicator impact score, ω F 、ω p 、ω l The values ​​of and ω are 0.2, 0.2, 0, 3, and 0.3, respectively. Based on the relationship between the comprehensive index impact score and the threshold, determine whether re-grasping is necessary. In multiple grasping experiments, the comprehensive index impact score and grasping result of each grasp are recorded, and the corresponding distribution graph is drawn. The distribution of success and failure values ​​is observed, and the dividing point between success and failure is found as the threshold.

[0146] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A control algorithm fusion method for a well repair robot, characterized in that: The following steps are involved: S1: The video stream is collected by a high-resolution camera for preprocessing, the target area is detected using the improved YOLOv7, and a multi-branch structure is designed based on the Res-Net-50 network to extract the multi-dimensional features of the target object; S2: Fuse the YOLOv7 object detection results, Res-Net-50 image features, and classifier recognition results to generate an interactive 3D model. S3: Design a dynamic weight fusion network to assign modal weights based on point cloud density and image texture confidence, generate a joint feature representation, and train a lightweight semantic segmentation model based on the joint features to distinguish different target object categories; S4: Select the optimal grasping point and grasping strategy based on the target object category and characteristics. Based on the A* algorithm, it fuses multi-sensor data to generate a work environment map, mark obstacle types, and plan a safe path. S5: Define the wellhead center coordinate system and the robot coordinate system, calculate the pose and joint angles, obtain the real-time pose through the IMU and camera, calculate the error and dynamically adjust the trajectory, and perform pose control and trajectory correction based on sensor feedback; S6: Based on the grasping planning results, control the finger joints of the manipulator to grasp the target object. By normalizing the contact force, posture change, visual inspection results and stability index, and then adding the weighted comprehensive index influence scores, it is determined whether re-grasping is needed.

2. The control algorithm fusion method for a well repair robot according to claim 1, characterized in that: The step S1 comprises the following steps: Use a high-resolution camera to capture real-time video streams of the workspace and perform preprocessing operations on the captured images, including denoising, color correction, and white balancing. YOLOv7 was selected as the basic object detection model. Deformable convolutions were introduced into its E-ELAN backbone network. A coordinate attention mechanism was added to the detection head. Gaussian heatmap regression was used instead of traditional bounding box regression. Feature maps of different scales were dynamically fused using the BiFPN weighted bidirectional feature pyramid. Res-Net-50 is selected as the basic network, CBAM is embedded in the residual block, and a multi-branch network is designed to extract the shape topology structure, RGB-HSV color space features and LBP texture features of the object respectively. The features extracted by different branches are fused to form a multi-dimensional semantic information representation.

3. The control algorithm fusion method for a well repair robot according to claim 1, characterized in that: The step S2 comprises the following steps: The target area detected by YOLOv7 is input into the Res-Net-50 network enhanced by the attention mechanism to extract the multi-dimensional features of the target object; Based on the extracted multi-dimensional features, a classifier is used to classify and identify the target object; The sub-pixel bounding box output by YOLOv7 is modified twice, and the multi-dimensional semantic information extracted by ResNet-50 is integrated with the target detection results to form a complete target description, including geometric attributes: bounding box coordinates and shape outline; appearance attributes: color histogram and LBP texture features; semantic attributes: category label and multi-dimensional semantic vector; The integrated target description is output in the form of an interactive three-dimensional model.

4. The control algorithm fusion method for a well repair robot according to claim 1, characterized in that: The step S3 comprises the following steps: In the working space, the target object is scanned by a laser rangefinder to obtain the original point cloud data and pre-process the point cloud data; Align the point cloud with the high-resolution image through a time synchronization mechanism and unify the coordinate system based on the IMU attitude data; Extract features from point cloud data, including normal vectors, curvature, and local surface features, and perform adaptive downsampling on high-density point clouds; Combined with the Res-Net-50 network to extract the multi-dimensional features of the target object; Design a transmembrane feature projection module: map the multi-dimensional features extracted by the Res-Net-50 network to a point cloud coordinate system and perform spatial alignment through feature matching; Design a dynamic weight fusion network to dynamically assign modal weights based on point cloud density and image texture confidence to generate a joint feature representation; A lightweight semantic segmentation model is trained based on joint features. The model outputs multi-category confidence for each voxel and assigns a semantic label to each voxel based on the confidence output by the model. The radiation field of the voxel is fused with the spatial position characteristics and image characteristics of the voxel to generate the radiation intensity or color of the voxel.

5. The control algorithm fusion method for a well repair robot according to claim 4, characterized in that: The point cloud data is preprocessed, comprising the following steps: Use PCL's voxelGrid to calculate the point density of each voxel, set high-density thresholds and low-density thresholds; recursively subdivide octree nodes for high-density areas, merge adjacent voxels for low-density areas; and construct a KD-tree inside the octree node for fast neighborhood search; Huffman encoding of voxel semantic labels; Adjust the filter strength according to the voxel semantic label and use bilateral filtering to preserve the sharpness of the voxel boundary. Specifically: Among them, W(p) is the weight value of point p, p is the position vector of the current point, and p c is the position vector of the center point, I(p) and I(p c ) are point p and point p respectively c The intensity value, σ d is the scale parameter of the spatial domain, σ r is the scale parameter of the intensity domain; An axial attention module is introduced to calculate the position of the point cloud projected onto the image plane and align it with the image segmentation mask.

6. The control algorithm fusion method for a well repair robot according to claim 4, characterized in that: The method of fusing the spatial position characteristics and image characteristics of voxels through the neural network radiation field to generate the radiation intensity or color of the voxels includes the following steps: The spatial position features of the voxels and the color and texture features extracted by the ResNet-50 network are input into the neural network radiation field to generate the radiation intensity or color of the voxels, specifically: c=σ(ω1·P_f+ω2·I_f) Where c is the radiation intensity or color of the voxel, P_f is the spatial position feature, I_f is the image feature, including color and texture features, σ is the activation function, ω1 and ω2 are the corresponding weights; Match the material library according to voxel semantic labels and use UV mapping to map the material to the 3D model surface.

7. The control algorithm fusion method for a well repair robot according to claim 1, characterized in that: The step S4 comprises the following steps: Based on the 3D model and voxel semantic labels of the target object, the optimal grasping point is selected and a grasping strategy is formulated: for regular large smooth objects, a multi-point grasping strategy is selected with evenly distributed grasping points; for irregular small rough objects, a single-point grasping strategy is selected with the geometric center or center of gravity as the grasping point; Plan the robot's path from its current position to the target grasping point using the A* algorithm; Use cameras and lidar to obtain information about the robot's working environment, align point cloud data with high-resolution images, generate a working environment map, and combine semantic segmentation results to mark the types of obstacles; During the A* algorithm path planning process, obstacles are regarded as impassable nodes, and the expansion operation is used to expand the obstacle area.

8. The control algorithm fusion method for a well repair robot according to claim 1, characterized in that: The step S5 comprises the following steps: Define the wellhead center coordinate system and the robot coordinate system, and align the origin of the robot coordinate system with the wellhead center; Use DH parameters to describe the joint and link structure of the robot, calculate the pose of the end effector based on the joint angle, and calculate the joint angle based on the target pose; The joint angle is obtained by numerical iteration method, and the optimal solution is selected according to actual needs. i , i = 1, 2, ..., n, n is the number of joints; Discretize the joint angle sequence, assign it to the time axis, use the interpolation method to generate a smooth trajectory, map the joint angle sequence to the time axis, and generate a trajectory with continuous velocity and acceleration; The real-time pose of the end effector is obtained through the IMU sensor, the error between the real-time pose and the target pose is calculated, and the joint angle is adjusted according to the error; at the same time, the image of the grasping point is obtained using a camera, the real-time pose of the grasping point is estimated through a visual algorithm, and the trajectory is adjusted according to the visual feedback.

9. The control algorithm fusion method for a well repair robot according to claim 8, characterized in that: Defining the wellhead center coordinate system and the robot coordinate system and aligning the origin of the robot coordinate system with the wellhead center includes the following steps: Define the wellhead center coordinate system: the origin is located at the wellhead center, the Z axis is perpendicular to the wellhead plane and points upward, and the X axis and Y axis are in the horizontal direction respectively; Define the robot coordinate system: the origin is located at the robot base, the Z axis is along the robot's main axis, and the X and Y axes point to the front and back and left and right directions of the robot respectively; According to the current position and posture of the robot, the transformation matrix between the robot coordinate system and the wellhead center coordinate system is calculated, and the position and posture of the grasping point are transformed from the robot coordinate system to the wellhead center coordinate system. Specifically: P w =T r_w ·P r Among them, P w is the wellhead center coordinate system, T r_w is the coordinate system transformation matrix, P r is the robot coordinate system.

10. The control algorithm fusion method for a well repair robot according to claim 1, characterized in that: The step S6 comprises the following steps: According to the results of the grasping plan, the finger joints of the manipulator are controlled to grasp the target object; Force sensors are installed at the ends of the manipulator's finger joints to measure contact force in real time. An IMU sensor is installed on the target object to measure the object's acceleration and angular velocity in real time. The displacement change is calculated through integration, and a camera is used to capture an image of the target object. The pose change is calculated through feature point matching. A high-resolution camera is installed above the manipulator and a deep learning algorithm is used to detect whether the target object is in the manipulator. If the target object is in the manipulator, the binary output is 1, otherwise it is 0. The stability index is calculated based on the contact force distribution of each finger joint and the center of gravity of the object. Normalize the value of each indicator and assign a weight to each indicator. The total weight is 1. The weighted sum is added to obtain the comprehensive indicator impact score, which is as follows: Among them, E is the comprehensive indicator impact score, F ′ , ΔP ′ 、l ′ 、S ′ and M′ are the normalized contact force, posture change, visual inspection results and stability index, ω F 、ω p 、ω l and ω are the corresponding weights respectively; If E ≥ threshold, the crawling is successful and no re-crawl is required; if E < threshold, the crawling fails, and the crawling strategy is replanned and the crawling is attempted again.

Citation Information

Patent Citations

  • Path planning method for mechanical arm of well repair equipment

    CN114986508A

  • Agricultural robot control method and device and agricultural robot

    CN116021526A

  • Visual guidance grabbing method of robot

    CN116787432A

  • Device and method for detecting targets based on guide feature maps

    KR102655237B1

  • Device and Method to Position an End Effector in a Well

    US20200055196A1

Cited By

  • Robot grabbing pose planning method based on three-dimensional point cloud recognition

    CN121223809A