A multi-modal pipe column intelligent identification and positioning method based on joint intelligent screwing and unscrewing

By employing a multimodal intelligent identification and positioning method for drill strings, utilizing the synchronous triggering and joint calibration of industrial cameras and lidar, and combining neural networks for feature fusion, the positioning error and identification robustness issues in complex drilling environments are resolved. Stable identification and real-time positioning are achieved under harsh conditions, assisting drillers in intelligent unloading.

CN122115574APending Publication Date: 2026-05-29JILIN UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2026-04-09
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In complex drilling environments, existing technologies suffer from problems such as time synchronization errors, spatial inconsistencies across sensors, insufficient robustness in identification, and difficulty in simultaneously achieving real-time performance. These issues lead to large string positioning errors, a high probability of misjudging equipment status, and increased operational risks.

Method used

A multimodal intelligent identification and positioning method for pipe columns is adopted. Through joint calibration and synchronous triggering of industrial cameras and LiDAR, combined with convolutional neural networks and point cloud neural networks, the temporal and spatial consistency of image and point cloud data is achieved. Multi-layer fusion feature extraction and recognition are performed to output the spatial position and boundary feature lines of the joint.

Benefits of technology

The system's robustness and adaptability have been improved, enabling stable identification under harsh conditions such as vibration, reflection, obstruction, rain, fog, and mud pollution. It meets real-time processing requirements and provides three-dimensional rotating frame parameters and boundary feature lines to assist iron drillers in intelligent unloading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115574A_ABST
    Figure CN122115574A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal pipe column intelligent identification and positioning method based on joint intelligence screwing, comprising: constructing pipe column identification system, including industrial camera, laser radar, synchronous trigger controller, host processing unit, pipe column includes several pipe bodies connected by joint, joint is the connecting position between two adjacent pipe columns for realizing mechanical connection and torque transmission;Utilize synchronous trigger controller to generate multi-way synchronous pulse signal to drive industrial camera and laser radar synchronous acquisition;Joint calibration is carried out to industrial camera and laser radar;Preprocess synchronized image and point cloud data;Image features and point cloud geometric features are extracted and fused;Based on the fusion result, the pipe column target is identified and positioned, and the position, attitude, size information or boundary feature line position is output;The identification result is sent to terminal.The application reduces the influence of space-time mismatch by hardware synchronization time service and joint calibration, and improves the recognition stability in complex environment by multi-modal fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of drilling equipment automation, industrial vision perception and multimodal intelligent recognition technology, specifically involving a multimodal intelligent identification and positioning method for tubing based on intelligent joint unscrewing. Background Technology

[0002] In industrial settings such as ultra-deep scientific drilling and oil and gas drilling, the connection, disassembly, and relocation of drilling rigs are frequently performed. These work areas typically involve harsh conditions including high-frequency vibrations, highly reflective surfaces, drastic changes in day and night lighting, dust, rain, fog, and mud contamination. These environmental factors can significantly affect the output of a single sensor due to nonlinear noise, leading to increased drilling rig positioning errors, a higher probability of misjudging equipment status, and consequently, increased operational risks.

[0003] Existing technologies relying on inertial navigation, encoders, or electromagnetic sensors are prone to cumulative errors and are susceptible to electromagnetic interference in metallic environments. Methods based solely on traditional image processing lack robustness under complex conditions and struggle to simultaneously achieve real-time performance, accuracy, and generalization capability. Single-lidar solutions suffer from degraded point cloud quality on highly reflective surfaces and in rain or fog, and the lack of texture information limits recognition accuracy. With the increasing demands for automated wellhead guidance, drill pipe movement monitoring, and multi-robot collaborative operations, there is an urgent need for an intelligent pipe string identification and positioning solution that integrates spatiotemporal synchronization and multi-source fusion while also considering engineering deployment capabilities, to assist drillers in intelligent joint tightening and loosening. Summary of the Invention

[0004] To address the problems of time synchronization errors, cross-sensor spatial inconsistencies, insufficient recognition robustness, and difficulty in achieving real-time performance in existing technologies in complex drilling environments, this invention proposes a multimodal intelligent identification and positioning method for tubing strings based on intelligent joint unscrewing.

[0005] This invention provides a method for intelligent identification and positioning of multimodal tubing strings based on intelligent joint unscrewing, comprising the following steps:

[0006] S1. Construct a pipe column recognition system, configuring an industrial camera, LiDAR, synchronous trigger controller, and host processing unit; the pipe column comprises several sections of pipe connected by joints, the joints being the connection points between adjacent pipe sections for mechanical connection and torque transmission; the industrial camera is used to acquire image data of the drilling operation area, and the LiDAR is used to acquire 3D point cloud data of the operation area; the synchronous trigger controller is used for unified triggering and timing control of the industrial camera and LiDAR; the host processing unit is used for data reception, spatiotemporal alignment, feature fusion, and recognition and positioning calculations; the host processing unit is configured with an image branch and a point cloud branch; the image branch is used for feature extraction, coordinate transformation, and joint boundary feature line positioning processing of the image data; the point cloud branch is used for geometric feature extraction and spatial parameter estimation of the point cloud data.

[0007] S2. The synchronous trigger controller generates multiple synchronous pulse signals to drive the industrial camera and lidar to acquire data synchronously, and establishes a unified time reference to establish the time correspondence between image frames and point cloud frames.

[0008] S3. Perform joint calibration on the industrial camera and LiDAR to obtain joint calibration parameters, i.e., cross-sensor spatial transformation parameters, and align the image space and point cloud space based on the spatial transformation parameters. The joint calibration includes: establishing the correspondence between image pixels and point cloud 3D points; solving the initial value of the extrinsic parameter matrix based on the correspondence; projecting the point cloud data collected by the LiDAR onto the image plane; extracting the image grayscale information of the corresponding pixel region on the image plane; simultaneously extracting the laser reflection intensity information of the corresponding point cloud points; optimizing the initial value of the extrinsic parameter matrix by minimizing the cross-modal similarity error between the image grayscale information and the point cloud laser reflection intensity information; constructing a similarity error function to obtain the extrinsic parameter calibration result. The similarity error function is:

[0009] ;

[0010] in The grayscale value of the image pixels. The laser reflection intensity corresponds to the point cloud point. Let the external parameters to be optimized be the rotation matrix and the translation vector. This is the intensity normalization coefficient;

[0011] S4. Preprocess the synchronized image data and point cloud data to obtain multimodal input data for recognition;

[0012] S5. Extract image features and point cloud geometric features through the image branch and point cloud branch respectively, and fuse the extracted image features and point cloud geometric features using at least one of the following fusion methods: data-level fusion, feature-level fusion, or decision-level fusion; wherein, the image branch includes an image feature extraction network, which uses at least one of the following methods to encode the image features and generate an image feature map; the point cloud branch includes a point cloud geometric feature extraction network, which uses at least one of the following methods to extract point cloud geometric features: PointNet, PointNet++, or sparse convolutional network.

[0013] S6. Based on the fusion results, identify and locate the joint area of ​​the tubing, and output at least one of the following: spatial position of the joint, axial direction, joint feature line, end face boundary position, and size parameters. Determine the relative position deviation, axial deviation, angular deviation, or spacing deviation between the tubing to be connected based on the joint parameters.

[0014] S7. The joint identification and positioning results are sent to the iron drill terminal to control the iron drill clamping mechanism, screwing mechanism or moving mechanism to perform joint centering, guiding and approaching, thread engagement start position adjustment, screwing and tightening or reverse unscrewing actions, and to correct the execution trajectory in real time according to the joint position change during the operation.

[0015] The core innovation of this invention lies in achieving stable perception through a system-level technology chain of "time synchronization + spatial synchronization + multimodal fusion + recognition output linkage". The solution reduces the impact of multi-sensor time drift and spatial mismatch on recognition results through hardware-synchronized time synchronization and joint calibration; it improves recognition stability in scenarios with vibration, reflection, occlusion, rain, fog, and mud pollution through complementary perception of images and point clouds and a multi-layer fusion strategy; the recognition module can output 3D rotating frame parameters to assist drillers in intelligent joint unloading, and can also output key boundary feature line positions as auxiliary constraints for depth estimation and action judgment, exhibiting good engineering scalability.

[0016] Compared with the prior art, the present invention has the following beneficial effects:

[0017] (1) Improved system-level robustness: By using hardware synchronization and joint calibration, the impact of time drift and spatial mismatch of multiple sensors on the recognition results is reduced, ensuring the temporal and spatial consistency of images and point cloud data;

[0018] (2) Strong adaptability to complex environments: By employing complementary perception of images and point clouds and a multi-layer fusion strategy, the recognition stability is improved in harsh scenarios such as vibration, reflection, occlusion, rain, fog, and mud pollution. Images provide rich texture and color information, while point clouds provide accurate geometric depth information. The fusion of the two compensates for the limitations of a single sensor.

[0019] (3) Good engineering scalability: The recognition module can output three-dimensional rotating box parameters to assist the iron driller's work, and can also output the position of key boundary feature lines as auxiliary constraints for depth estimation and action judgment, meeting the needs of different application scenarios;

[0020] (4) Real-time deployment friendly: By adopting a lightweight network structure, parameter smoothing and timing filtering strategy, it meets the real-time processing requirements on site and can realize low-latency inference on edge computing devices. Attached Figure Description

[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the present invention and are not intended to limit the scope of this disclosure. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0022] Figure 1 This is a flowchart illustrating the multimodal tubing intelligent identification system disclosed herein.

[0023] Figure 2 This is a schematic diagram of the multi-sensor synchronization triggering and timestamp synchronization process disclosed in this publication;

[0024] Figure 3 This is a schematic diagram of the joint calibration and spatial alignment process of industrial cameras and lidar disclosed in this paper.

[0025] Figure 4 This is a schematic diagram of the multimodal data preprocessing, feature extraction and fusion process disclosed in this publication. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. However, the scope of protection of this invention is not limited to the following embodiments.

[0027] like Figure 1 As shown, this embodiment provides a method for intelligent identification and positioning of multimodal tubing based on intelligent joint unscrewing, including the following steps:

[0028] Step S1: Construct a tubular column identification system.

[0029] A multimodal tubing column identification system based on intelligent joint unscrewing was constructed. The system includes an industrial camera, LiDAR, synchronous trigger controller, and host processing unit.

[0030] The tubing string comprises several sections of tubing connected by joints. Each joint is a connection point between adjacent tubing sections used for mechanical connection and torque transmission, typically a threaded joint formed by the engagement of a male and female threaded end. The joint area includes the joint body, joint boundary, end face edge, thread initiation position, and adjacent transition area of ​​the tubing. The main feature of this invention is the boundary characteristic line of the joint area, used to assist steelworkers in achieving intelligent unloading.

[0031] An industrial camera is fixedly mounted in front of the drilling work area to acquire two-dimensional image data of the pipe joint area. A lidar is rigidly connected to the industrial camera and calibrated to map them to the same coordinate system, used to acquire three-dimensional point cloud data of the work area.

[0032] The synchronous trigger controller is used to perform unified triggering and timing control of industrial cameras and lidar. It can be implemented using a microcontroller (such as STM32), which synchronously drives each sensor to sample by outputting hardware trigger signals and generates a unified timestamp.

[0033] The host processing unit is used to perform multimodal data processing and recognition / localization tasks, and it is internally configured with an image branch and a point cloud branch. The image branch is used to extract image features, and the point cloud branch is used to extract three-dimensional geometric features.

[0034] Step S2: Multi-sensor time synchronization acquisition.

[0035] like Figure 2 The diagram illustrates a multi-sensor synchronization triggering and timestamp synchronization process. The synchronization trigger controller generates multiple synchronization pulse signals to drive the industrial camera and LiDAR for synchronous sampling. In a preferred embodiment, the synchronization trigger controller uses a microcontroller (such as an STM32) to generate multiple synchronization pulse signals to trigger sampling by the industrial camera and LiDAR, and provides precise timestamp information to the host processing unit via serial port, Ethernet, shared memory, or hardware timestamp synchronization mechanism. The host processing unit establishes a precise correspondence between image frames and point cloud frames based on these timestamps, thereby establishing a unified time reference. For data with slight time offsets, time compensation, interpolation resampling, or alignment correction can be used to achieve time synchronization with second-level or higher precision.

[0036] Step S3: Joint calibration and spatial alignment.

[0037] Figure 3This diagram illustrates the joint calibration and spatial alignment process for an industrial camera and LiDAR. To eliminate spatial inconsistencies among multiple sensors, joint calibration is performed on the industrial camera and LiDAR to obtain cross-sensor spatial transformation parameters. The specific process includes: First, using the fast joint calibration method FAST-Calib, feature points are selected in the image manually or semi-automatically, and corresponding 3D points are extracted from the point cloud, establishing a correspondence between image pixels and point cloud 3D points. Based on this correspondence, initial values ​​of the extrinsic parameter matrix between the camera and LiDAR are obtained by solving the camera projection model. After obtaining the initial extrinsic parameter values, the LiDAR point cloud is projected onto the image plane, and the image grayscale information at the corresponding pixel positions and the laser reflection intensity information of the point cloud points are extracted. A cross-modal similarity error function is constructed, and the extrinsic parameter matrix is ​​iteratively updated using a nonlinear least squares optimization method to minimize the difference between image grayscale and point cloud intensity, thereby obtaining high-precision extrinsic parameter calibration results. During system operation, industrial camera image frames and lidar point cloud frames are paired according to timestamps, and the spatial transformation parameters are used to align the image space and point cloud space. For example, the point cloud is projected onto the image plane, or the image features are back-projected onto three-dimensional space to obtain multimodal data in a unified coordinate system.

[0038] Joint calibration is performed on the industrial camera and LiDAR to obtain joint calibration parameters, namely cross-sensor spatial transformation parameters, and the alignment of the image space and point cloud space is completed based on the spatial transformation parameters. The joint calibration includes: establishing the correspondence between image pixels and 3D points in the point cloud; solving the initial value of the extrinsic parameter matrix based on the correspondence; projecting the point cloud data collected by the LiDAR onto the image plane; extracting the image grayscale information of the corresponding pixel region in the image plane; simultaneously extracting the laser reflection intensity information of the corresponding point cloud points; optimizing the initial value of the extrinsic parameter matrix by minimizing the cross-modal similarity error between the image grayscale information and the point cloud laser reflection intensity information; constructing a similarity error function; and thus obtaining the extrinsic parameter calibration result. The similarity error function is as follows:

[0039] ;

[0040] in The grayscale value of the image pixels. The laser reflection intensity corresponds to the point cloud point. Let the external parameters to be optimized be the rotation matrix and the translation vector. This is the intensity normalization coefficient;

[0041] Specifically, the initial values ​​of the extrinsic parameters are obtained by utilizing the 3D-2D correspondence, and more accurate camera-LiDAR extrinsic parameters are obtained through cross-modal similarity optimization.

[0042] Step S4: Multimodal data preprocessing.

[0043] The synchronized and aligned image data and point cloud data are preprocessed separately to generate standardized multimodal input data.

[0044] Image data preprocessing includes scaling the image while maintaining its aspect ratio, padding the scaled image's boundaries, generating an effective region mask, and normalizing the image pixel values. The effective region mask is used to suppress interference from the padded regions during subsequent recognition calculations.

[0045] Point cloud data preprocessing includes at least filtering the point cloud to remove outliers, performing coordinate transformation on the point cloud (e.g., converting to the camera coordinate system), and projecting the point cloud onto the image space based on joint calibration parameters. The point cloud filtering includes at least one of statistical outlier filtering, radius outlier filtering, and voxel downsampling. Statistical outlier filtering removes outliers based on the distance distribution between a point and its neighbors; radius outlier filtering removes isolated points based on the number of neighbors within a preset radius; and voxel downsampling reduces point cloud density and improves subsequent processing efficiency. Based on the spatial alignment results, a joint data representation, such as a depth map or dense point cloud map corresponding to image pixels, can be further generated.

[0046] Step S5: Multimodal feature extraction and fusion.

[0047] like Figure 4 The diagram shows the multimodal data preprocessing, feature extraction and fusion process. The preprocessed data is sent to the image branch and the point cloud branch for feature extraction.

[0048] The image branch includes an image feature extraction network, which employs at least one of convolutional neural networks (such as ResNet, DenseNet), lightweight convolutional networks (such as MobileNet), or Transformer networks (such as ViT) to encode features of the image and generate a high-dimensional image feature map.

[0049] The point cloud branch includes a point cloud geometric feature extraction network, which uses at least one of the following: PointNet, PointNet++, or SparseConvNet, to extract the geometric structural features of the point cloud.

[0050] Subsequently, at least one of the following fusion methods—data-level fusion, feature-level fusion, or decision-level fusion—is used to fuse the extracted image features and point cloud geometric features.

[0051] Data-level fusion scheme: When using data-level fusion, spatial alignment of the image branch and point cloud branch is required to establish a spatial correspondence between image pixels and corresponding 3D points, thereby generating a multimodal data representation containing color and depth information. Specifically, by projecting the point cloud onto the image plane, each image pixel simultaneously possesses corresponding depth information, forming fused data containing color and depth information, i.e., a color point cloud. In data-level fusion, image prediction heads or point cloud prediction heads can be used for target recognition and spatial information estimation. When using an image prediction head, the 2D position or feature line position of the target in the image is first predicted through an image detection network, and then the spatial parameters of the target are obtained by combining the spatially aligned point cloud depth information. When using a point cloud prediction head, the fused color point cloud can be input into a 3D detection network, and the center position, axial direction, attitude angle, and end face boundary information of the pipe joint can be directly obtained through the prediction head.

[0052] Decision-level fusion scheme: When using decision-level fusion, the image branch and the point cloud branch each have independent recognition heads. The image branch outputs the target feature line position based on the image and the confidence level of the recognition head for that position simultaneously through the image recognition head; the higher the confidence level, the closer it is to the true position. The point cloud branch outputs the target feature line based on the point cloud and the confidence level of the recognition head for that position through the point cloud recognition head. The final decision-level fusion includes weighting the feature line position based on the confidence levels of the prediction results of the image recognition result and the point cloud recognition result, or using geometric consistency for rule-constrained fusion. Specifically, the target feature line position results output by the image branch are compared with the target feature line position results output by the point cloud branch and projected onto a unified coordinate system. The consistency comparison includes a judgment based on at least one geometric consistency index, such as feature line center position deviation and orientation angle deviation. A threshold can be set according to the allowable error limit. When the geometric consistency index meets the preset threshold condition, the target feature line position is weighted and fused based on the confidence scores output by the image branch and the point cloud branch. When the geometric consistency index does not meet the preset threshold condition, the recognition result with higher confidence is selected as the final result, or the candidate results that do not meet the consistency condition are filtered out.

[0053] Feature-level fusion scheme: When feature-level fusion is used, after the initial feature extraction is completed in the image branch and the point cloud branch, the feature maps output by the two branches are fused in the intermediate feature layer. Specifically, the image branch outputs image feature maps. Point cloud branch outputs point cloud feature map Both can be adjusted to have the same spatial resolution through interpolation or convolution transformation. This achieves spatial alignment.

[0054] In one implementation, the image feature map and the point cloud feature map can be concatenated along the channel dimension to form a fused feature map:

[0055] ;

[0056] in This represents a feature concatenation operation along the channel dimension. The concatenated fused feature map contains both image texture information and point cloud geometric information, which can be further input into subsequent network layers for feature encoding and target prediction.

[0057] In another implementation, fusion can be performed using a weighted summation method, i.e.:

[0058] ;

[0059] in and The weighting coefficients can be fixed parameters or automatically generated by the network through learning, and are used to represent the importance ratio of image features and point cloud features in different scenarios.

[0060] In a further implementation, an attention mechanism can be introduced to achieve adaptive fusion. For example, an attention module can calculate attention weights for image features and point cloud features respectively, and dynamically weight the two branches of features according to the weights, thereby enhancing the feature response related to the target and suppressing background noise. The attention mechanism can be channel attention, spatial attention, or a combination thereof.

[0061] In another implementation, a cross-modal interaction structure can be used to achieve deep fusion, such as a Transformer-based cross-attention mechanism. Specifically, image features are used as query vectors and point cloud features as key-value pairs, or image features are used as key-value pairs and point cloud features as query vectors. Cross-attention calculation is used to achieve interaction between the two modalities, thereby obtaining a feature map containing multimodal semantic information. The fused feature map... This will serve as the input for subsequent prediction heads, thereby providing a more stable and accurate feature base for the identification and localization of targets in the tubular column.

[0062] This invention employs a differentiable spatial-to-numerical transformation method, such as DSNT or soft-argmax, in the prediction stage after feature-level fusion to convert dense probability distributions into target coordinates. Specifically, in the network structure of this invention, the fused feature map is processed by the prediction head to generate a heatmap or probability distribution map representing the probability distribution of the target's spatial location. Each pixel in this probability map has a value representing the probability of the target appearing at the corresponding spatial location. Subsequently, this invention uses a differentiable spatial-to-numerical transformation method to estimate the coordinates of the probability distribution map. The specific process involves first normalizing the probability distribution map to form a probability distribution; then performing a weighted summation operation based on each spatial location coordinate and its corresponding probability value to obtain the continuous spatial coordinate estimation result of the target. Through this process, the dense probability map output by the network can be converted into two-dimensional coordinates, three-dimensional coordinates, or keypoint coordinates of the target. DSNT establishes a corresponding spatial coordinate grid from the probability distribution map output by the prediction head and performs a weighted summation operation on the probability distribution to obtain the continuous coordinate estimation result of the target position. The soft-argmax method, on the other hand, normalizes the probability distribution map and performs a weighted summation of the corresponding coordinates based on the probability values ​​of each spatial location to obtain the continuous coordinates of the target. This achieves the conversion from spatial probability distribution to continuous coordinates, avoiding the discretization error caused by the traditional argmax operation and improving the accuracy and stability of target localization.

[0063] Step S6: Target identification and localization of the tubing.

[0064] Based on the fusion result obtained in step S5, the joint area of ​​the tubing is identified and located. Specifically, this includes: firstly, processing the fused feature map using a prediction head to generate a spatial location probability distribution or candidate region for the joint area, and then determining the joint parameters of the target in a unified spatial coordinate system using one of the following methods: direct regression, a combination of detection and regression, or a method based on the probability distribution mapping in step S5. The joint parameters are at least one of the following: spatial location of the joint, axial direction, joint feature line, end face boundary position, and dimensional parameters. Based on the joint parameters, the relative positional deviation, axial deviation, angular deviation, or spacing deviation between the tubing to be connected is determined.

[0065] When it is necessary to output the location of boundary feature lines related to the joint target (such as the lower boundary of the joint), the image branch can be specifically used for this task. The specific implementation is as follows: The image branch adopts an encoder-decoder structure (such as U-Net) and outputs a two-dimensional response heatmap. In the decoding stage, a directional spatial enhancement module is introduced to enhance the target's features. This module allows the model to focus more on lines while effectively suppressing other noise. Specifically, the directional spatial enhancement module is set at the intermediate feature map output end of the decoding stage to perform directional enhancement processing on the intermediate feature map. Let the intermediate feature map output by the decoding stage be... Where C represents the number of channels, H represents the feature map height, and W represents the feature map width. When the preset direction is the width direction, pooling is first performed along the width direction on the intermediate feature map F, resulting in a size of... The directional statistical features; correspondingly, when the preset direction is the height direction, the intermediate feature map F is first pooled along the height direction to obtain a size of The directional statistical features are then used. Pooling can be implemented using average pooling. Subsequently, the directional statistical features are input into a weight generation unit, which consists of convolutional layers or fully connected layers. This weight generation unit maps the directional statistical features and outputs a directional weight map corresponding to the spatial dimension of the intermediate feature map. Generally, the weight generation unit includes one or more linear mapping layers and non-linear activation layers to convert the statistically derived directional features into spatial weights with values ​​within a preset range. After obtaining the directional weight map, it is broadcast-expanded along the unpooled spatial dimension to make it correspond to the original intermediate feature map. They have the same spatial dimensions; then the expanded orientation weight map is multiplied element-wise with the original intermediate feature map to obtain the enhanced feature map. :

[0066] ;

[0067] in, Represents the directional weighting graph. This represents element-wise multiplication. Through this element-wise reweighting process, the response in the direction of the target boundary feature line can be enhanced, while background noise unrelated to that direction can be suppressed. Enhanced feature map. The input continues to subsequent convolutional layers or prediction heads to generate the final 2D response heatmap. After obtaining the heatmap, directional aggregation (e.g., any or a combination of logarithmic summation exponential aggregation, summation aggregation, weighted summation aggregation, or maximum aggregation along the width or height directions), Softmax probability normalization with mask constraints, and coordinate expectation calculation are performed to obtain sub-pixel precision continuous coordinate estimation results. After scale mapping, these coordinates are used to obtain the alignment of the header boundary line position with the original image.

[0068] Step S7: Sending and linking results.

[0069] The identification and positioning results of the target tubing are sent to the iron drill terminal. The identification and positioning results are calculated based on the center position, axial direction, attitude angle, and end face boundary information of the two tubing joints to determine the radial deviation, axial spacing, and tilt angle deviation between the joints to be connected, as well as the height of the joints from the ground. This information is used to assist the iron drill operator in intelligently tightening and loosening the joints.

[0070] Based on the identification and positioning results, the iron driller performs intelligent joint tightening and loosening linkage control of the joint, including: guiding the upper and lower tubing strings to center and guide them closer according to the spatial position and axial direction of the joint; determining the initial engagement position according to the joint end face boundary and joint feature line, and controlling the screwing mechanism to perform low-speed screwing; during the screwing process, dynamically correcting the movement trajectory of the actuator according to the real-time changes in the joint posture and position; when abnormal changes in joint axis deviation or boundary misalignment are detected, determining that there is a risk of oblique or incorrect screwing, and controlling the actuator to perform deceleration, stop, or retraction operations; during the unscrewing process, determining the clamping position and rotation direction according to the joint position and posture information, controlling the actuator to rotate in the opposite direction to achieve joint separation.

[0071] Training phase: Before training, the recognition model in the above method needs to construct a multimodal training dataset and be trained using optimization strategies.

[0072] Training dataset construction: First, 2D annotations (such as bounding boxes or keylines) are performed on the pipe targets in the image; then, the point cloud is projected onto the image space according to the calibration parameters, and the point cloud belonging to the target is selected using the 2D annotation regions; finally, cylindrical fitting is performed on the target point cloud, or corresponding geometric fitting (such as conical fitting, composite fitting) is performed according to the shape characteristics of the pipe joint, automatically generating 3D labels for supervised training, including at least one of position, pose, and size parameters. This method can effectively reduce the cost of purely manual 3D annotation.

[0073] Model Training: During the training phase, at least one of the following loss methods is jointly optimized: coordinate regression loss, distribution fitting loss, classification loss, bounding box regression loss, pose loss, and regularization term. For key boundary feature line localization scenarios, a one-dimensional Gaussian distribution can be generated as a soft label based on the labeled coordinates, and an entropy regularization term can be added to adjust the distribution sharpness. To improve the model's generalization ability, at least one data augmentation strategy can be used, including brightness enhancement, contrast enhancement, blur perturbation, noise injection, and occlusion simulation. In addition, parameter update strategies such as exponential moving average can be used to improve model stability.

[0074] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art should understand that any equivalent substitutions, modifications, or variations made to the synchronization method, calibration algorithm, fusion strategy, network architecture, probabilistic regression form, post-processing strategy, and deployment method without departing from the spirit and substance of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for intelligent identification and positioning of multimodal tubing strings for intelligent joint unscrewing, characterized in that, Includes the following steps: S1. Construct a pipe column recognition system, configuring an industrial camera, LiDAR, synchronous trigger controller, and host processing unit; the pipe column comprises several sections of pipe connected by joints, the joints being the connection points between adjacent pipe sections for mechanical connection and torque transmission; the industrial camera is used to acquire image data of the drilling operation area, and the LiDAR is used to acquire 3D point cloud data of the operation area; the synchronous trigger controller is used for unified triggering and timing control of the industrial camera and LiDAR; the host processing unit is used for data reception, spatiotemporal alignment, feature fusion, and recognition and positioning calculations; the host processing unit is configured with an image branch and a point cloud branch; the image branch is used for feature extraction, coordinate transformation, and joint boundary feature line positioning processing of the image data; the point cloud branch is used for geometric feature extraction and spatial parameter estimation of the point cloud data. S2. The synchronous trigger controller generates multiple synchronous pulse signals to drive the industrial camera and lidar to acquire data synchronously, and establishes a unified time reference to establish the time correspondence between image frames and point cloud frames. S3. Perform joint calibration on the industrial camera and lidar to obtain joint calibration parameters, namely cross-sensor spatial transformation parameters, and complete the alignment of the image space and point cloud space based on the spatial transformation parameters. The joint calibration includes: establishing a correspondence between image pixels and 3D points in the point cloud; solving for the initial value of the extrinsic parameter matrix based on the correspondence; projecting the point cloud data acquired by the lidar onto the image plane; extracting the image grayscale information of the corresponding pixel region in the image plane; simultaneously extracting the laser reflection intensity information of the corresponding point cloud points; optimizing the initial value of the extrinsic parameter matrix by minimizing the cross-modal similarity error between the image grayscale information and the point cloud laser reflection intensity information; constructing a similarity error function; and thus obtaining the extrinsic parameter calibration result. The similarity error function is as follows: ; in The grayscale value of the image pixels. The laser reflection intensity corresponds to the point cloud point. Let the external parameters to be optimized be the rotation matrix and the translation vector. This is the intensity normalization coefficient; S4. Preprocess the synchronized image data and point cloud data to obtain multimodal input data for recognition; S5. Extract image features and point cloud geometric features through the image branch and point cloud branch respectively, and fuse the extracted image features and point cloud geometric features using at least one of the following fusion methods: data-level fusion, feature-level fusion, or decision-level fusion; wherein, the image branch includes an image feature extraction network, which uses at least one of the following methods to encode the image features and generate an image feature map; the point cloud branch includes a point cloud geometric feature extraction network, which uses at least one of the following methods to extract point cloud geometric features: PointNet, PointNet++, or sparse convolutional network. S6. Based on the fusion results, identify and locate the joint area of ​​the tubing, and output at least one of the following: spatial position of the joint, axial direction, joint feature line, end face boundary position, and size parameters. Determine the relative position deviation, axial deviation, angular deviation, or spacing deviation between the tubing to be connected based on the joint parameters. S7. The joint identification and positioning results are sent to the iron drill terminal to control the iron drill clamping mechanism, screwing mechanism or moving mechanism to perform joint centering, guiding and approaching, thread engagement start position adjustment, screwing and tightening or reverse unscrewing actions, and to correct the execution trajectory in real time according to the joint position change during the operation.

2. The method according to claim 1, characterized in that, The synchronization trigger controller in step S2 is configured to use a microcontroller to achieve time synchronization. The microcontroller is configured to generate multiple synchronization pulse signals to trigger sampling by industrial cameras and lidar, and to provide timestamp information to the host processing unit through serial port, Ethernet, shared memory or hardware timestamp timing mechanism.

3. The method according to claim 1, characterized in that, Step S4 involves preprocessing the synchronized image data and point cloud data, including: scaling the image while maintaining its aspect ratio, filling the scaled image with boundaries, and labeling the real regions of the filled image as effective region masks, where the effective region masks are used to suppress interference from the filled regions on subsequent recognition calculations. Point cloud data preprocessing includes at least: filtering the point cloud to remove outliers, performing coordinate transformation on the point cloud, and projecting the point cloud onto the image space based on joint calibration parameters, so that each image pixel simultaneously has corresponding depth information, forming fused data containing color and depth information to obtain a color point cloud.

4. The method according to claim 1, characterized in that, In step S5, when decision-level fusion is used, the image branch and the point cloud branch each have independent recognition heads. The image branch recognition head outputs the image recognition result, and the point cloud branch recognition head outputs the point cloud recognition result. The decision-level fusion includes confidence-weighted fusion of the image recognition result and the point cloud recognition result.

5. The method according to claim 1, characterized in that, In step S5, when feature-level fusion is used, the image branch and / or point cloud branch contains a prediction head. The prediction head is used to generate a heat map or probability distribution map of the target location, and converts the heat map or probability distribution map into two-dimensional coordinates, three-dimensional coordinates, or key point coordinates of the target through a differentiable space-to-numerical transformation module. The differentiable space-to-numerical transformation module includes the differentiable space-to-numerical transformation DSNT method or the soft-argmax method.

6. The method according to claim 1, characterized in that, The identification and location of the target column in step S6 includes: generating target location information or target area information based on fused features, and determining the three-dimensional parameters of the target in a unified spatial coordinate system using one of the following methods: direct regression, detection and regression combination, or probability distribution mapping; wherein the three-dimensional parameters include at least one of position parameters, attitude parameters, and size parameters.

7. The method according to claim 1, characterized in that, The image branch is also used for locating the top edge, boundary line, or key structural line of the pipe joint. By outputting a two-dimensional response heatmap, and performing directional aggregation, mask constraint probability normalization, and coordinate expectation calculation on the heatmap, continuous coordinate estimation results are obtained. The image branch also includes a directional spatial enhancement module in the decoding stage. The directional spatial enhancement module generates spatial weights by statistical pooling the intermediate feature map output in the decoding stage in a preset direction, and reweights the intermediate feature map to enhance the response in the direction of the target before generating the two-dimensional response heatmap. The identification and positioning results in step S6 are used for intelligent joint tightening and loosening control, specifically including: calculating the radial deviation, axial spacing, and tilt angle deviation between the joints to be connected based on the center position, axial direction, attitude angle, and end face boundary information of the two pipe joints, and sending the deviations to the iron drill terminal to control the clamping mechanism and the screwing mechanism to perform joint centering, guiding approach, and thread engagement; during tightening or loosening, the execution trajectory is corrected based on the real-time updated joint position and attitude information.

8. The method according to claim 7, characterized in that, The directional aggregation is any one or a combination of logarithmic summation exponential aggregation, summation aggregation, weighted summation aggregation, or maximum value aggregation along the width or height direction; The probability normalization employs Softmax normalization with mask constraints.

9. The method according to claim 1, characterized in that, It also includes the training dataset construction steps: two-dimensional annotation of the joint area of ​​the pipe column in the image; projection of the point cloud onto the image space according to the calibration parameters and selection of target point cloud; performing cylindrical fitting on the target point cloud, or performing corresponding geometric fitting according to the shape characteristics of the pipe column joint, to generate three-dimensional labels for supervised training.

10. The method according to claim 9, characterized in that, The model training phase is configured to perform joint optimization using at least one of the following: coordinate regression loss, distribution fitting loss, classification loss, bounding box regression loss, pose loss, and regularization term. It may also employ at least one of the following data augmentation strategies: brightness enhancement, contrast enhancement, blur perturbation, noise injection, and occlusion simulation.