A positioning method and system for the drill pipe of a trenchless drilling rig based on image segmentation
By constructing and training the drill pipe panoramic image segmentation network model, the accuracy of drill pipe positioning in complex backgrounds is solved, and efficient and reliable drill pipe identification and positioning is achieved.
Patent Information
- Application Number
- CN202510206761.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The prior art is difficult to accurately detect the position and attitude of the drill pipe in complex contexts, and lacks the accuracy of ensuring coordinate system calibration.
By constructing a drill pipe panoramic image segmentation network model, using the image data set for training, generating segmentation result images, extracting edge features, calculating the furthest point pair and the minimum area rectangle, and converting the coordinate system to obtain the position information of the drill pipe.
Accurate detection of drill pipe position and attitude in complex backgrounds is achieved, ensuring the accuracy of coordinate system calibration, and improving the efficiency and reliability of drill pipe identification.
Smart Images

Figure CN119693461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and particularly to a positioning method and system for the drill pipe of a trenchless drilling rig based on image segmentation. Background Art
[0002] The drill pipe is a key component in trenchless drilling rig engineering construction. Its precise positioning and rapid grasping are crucial for the efficiency and safety of automated operations. In traditional operations, the assembly of drill pipes usually relies on manual operations, which leads to various problems. First, manual operations are time-consuming, especially when multiple drill pipes need to be operated or in frequent operations, significantly reducing the overall work efficiency. Second, since drill pipes are usually large in volume and heavy in weight, safety accidents are likely to occur during manual handling, especially in harsh construction environments. Finally, it is difficult for manual operations to ensure the high-precision positioning of drill pipes during assembly or placement, which may lead to the accumulation of errors in subsequent construction.
[0003] With the rapid development of industrial automation, the automated positioning and installation of drill pipes have gradually become a research hotspot. However, existing positioning and grasping technologies usually rely on pre-calibrated drill pipe positions and lack the ability of real-time detection and positioning, making it difficult to adapt to the complex and changeable on-site environment of trenchless drilling rigs. In recent years, object detection and semantic segmentation technologies based on computer vision have been widely used. Especially with the support of deep learning, high-precision recognition and positioning of target objects in complex scenes can be achieved. Nevertheless, in combining computer vision technology with mechanical control to achieve the automatic positioning of drill pipes of trenchless drilling rigs, existing technologies still face some challenges. These challenges include accurately detecting the position and posture of drill pipes in complex backgrounds, as well as adapting to different lighting conditions, occlusion situations, and environmental changes. In addition, existing technologies also have deficiencies in accurately converting the pixel coordinates of drill pipes in images into three-dimensional world coordinates and ensuring the accuracy of coordinate system calibration. Summary of the Invention
[0004] Therefore, the technical problem to be solved by the present invention is to overcome the inability to accurately detect the position and posture of drill pipes in complex backgrounds and ensure the accuracy of coordinate system calibration in the prior art.
[0005] In a first aspect, to solve the above technical problem, the present invention provides a positioning method for the drill pipe of a trenchless drilling rig based on image segmentation, including:
[0006] Obtaining an image data set of the drill pipe;
[0007] Constructing a panoramic image segmentation network model for the drill pipe, and training the panoramic image segmentation network model for the drill pipe with the image data set to obtain an optimized image segmentation network model;
[0008] Obtain the panoramic image of the drill pipe to be located, input the panoramic image of the drill pipe into the optimized image segmentation network model, and generate a segmentation result image;
[0009] Extract the edge features of the segmentation result image to obtain the edge coordinate point set of the drill pipe to be located;
[0010] Calculate the edge coordinate point set to obtain the farthest point pair; calculate the distance between the farthest point pair to obtain the minimum area rectangle;
[0011] Obtain the center coordinates of the minimum area rectangle, convert the center coordinates into camera coordinates, and then convert the camera coordinates into world coordinates to obtain the position information of the drill pipe to be located.
[0012] In an embodiment of the present invention, the drill pipe panoramic image segmentation network model includes an encoder, a multi-scale feature reconstruction module, and a decoder. The encoder performs downsampling to obtain a drill pipe feature map; the multi-scale feature reconstruction module performs multi-scale reconstruction on the drill pipe feature maps of different scales to obtain a reconstructed feature map; the decoder fuses the reconstructed feature map.
[0013] In an embodiment of the present invention, the method for the multi-scale feature reconstruction module to perform multi-scale reconstruction on the drill pipe feature maps of different scales is as follows:
[0014] ; ;
[0015] Among them, represents the multi-scale reconstruction feature, represents the fusion feature, represents upsampling, represents convolution, represents the dilated convolution with a dilation rate of a and a convolution kernel size of b, F 1 、 F 2 、 F 3 and F 4 respectively come from the features of the 1st to 4th convolutional layers in the encoder.
[0016] In an embodiment of the present invention, the encoder includes a depthwise separable convolution module, a convolution module, and a hybrid multi-scale attention module; the hybrid multi-scale attention module includes a channel attention module and a spatial attention module; the channel attention module is connected to the spatial attention module.
[0017] In an embodiment of the present invention, the method for training the panoramic image segmentation network model of the drill pipe using the image dataset is as follows: The image dataset is divided into a training set according to a preset ratio, and the panoramic image segmentation network model of the drill pipe is iteratively optimized based on the training set using the Adam backpropagation algorithm until the total loss function converges.
[0018] In an embodiment of the present invention, the calculation formula of the total loss function is:
[0019] ;
[0020] Wherein, represents the total loss, X represents the true value, Y represents the predicted value.
[0021] In an embodiment of the present invention, obtaining the image dataset of the drill pipe includes preprocessing each image in the image dataset, and the preprocessing method is:
[0022] Grayscale the image, perform data augmentation, unify the size, and enhance the contrast to obtain a first image;
[0023] Use the labelme script to perform semantic segmentation annotation on the first image.
[0024] In an embodiment of the present invention, calculating the edge coordinate point set to obtain the farthest point pair; calculating the distance between the farthest point pair to obtain the minimum area rectangle is as follows:
[0025] Identify the minimum x-coordinate point and the maximum x-coordinate point in the edge coordinate point set, and connect these two points to form a horizontal line;
[0026] Calculate the vertical distance from each point in the edge coordinate point set to the horizontal line, and mark the farthest point pair; rotate the horizontal line counterclockwise by an angle so that the line segment between the farthest point pair coincides; repeat the rotation process until the horizontal line is rotated one week, record the distance between the farthest point pairs during each rotation process, and calculate the corresponding rectangle area;
[0027] Compare the rectangle areas obtained at all rotation angles to obtain the minimum area rectangle.
[0028] In an embodiment of the present invention, the method for obtaining the center coordinates of the minimum area rectangle, converting the center coordinates to camera coordinates, and then converting the camera coordinates to world coordinates is as follows:
[0029] Use the internal parameter matrix to convert the center coordinates to a point in the normalized camera coordinate system, and the calculation formula is:
[0030] ;
[0031] Among them, represents a point in the normalized camera coordinate system, represents the abscissa of the center point, represents the ordinate of the center point, represents the internal parameter matrix;
[0032] Convert the point in the normalized camera coordinate system to a point in the camera coordinate by using the external parameter matrix of the camera;
[0033] According to the conversion of the camera coordinate to the world coordinate, the calculation formula of the world coordinate is:
[0034] ;
[0035] Among them, represents the rotation matrix, represents the translation vector; , and represent the coordinates of the point in the camera coordinate; , and represent the coordinates of the point in the world coordinate.
[0036] In an embodiment of the present invention, the method for extracting the edge features of the segmented result image is the Canny edge detection algorithm.
[0037] Second, to solve the above technical problems, the present invention provides a positioning system for a trenchless drilling rig drill pipe based on image segmentation, including:
[0038] A dataset acquisition module for acquiring an image dataset of the drill pipe;
[0039] A model construction module for constructing a drill pipe panoramic image segmentation network model, training the drill pipe panoramic image segmentation network model by using the image dataset, and obtaining an optimized image segmentation network model;
[0040] A segmented image output module for acquiring a drill pipe panoramic image of the drill pipe to be positioned, inputting the drill pipe panoramic image into the optimized image segmentation network model, and generating a segmented result image;
[0041] A feature extraction module for extracting the edge features of the segmented result image and obtaining an edge coordinate point set of the drill pipe to be positioned;
[0042] A rectangle acquisition module, configured to calculate the edge coordinate point set to obtain the pair of farthest points; calculate the distance between the pair of farthest points to obtain a rectangle with the minimum area.
[0043] A positioning module, configured to obtain the central coordinates of the rectangle with the minimum area, convert the central coordinates into camera coordinates, and then convert the camera coordinates into world coordinates to obtain the position information of the drill pipe to be positioned.
[0044] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0045] (1) For a positioning method and system of a trenchless drilling rig drill pipe based on image segmentation according to the present invention, by constructing an image segmentation network model, the drill pipe can be automatically and accurately segmented from a panoramic image, significantly reducing the need for manual operation and ensuring the consistency and accuracy of the segmentation result. The application of this method greatly improves the efficiency and reliability of drill pipe recognition. Based on the segmentation result, by calculating the edge coordinate point set of the image, the pair of farthest points and the rectangle with the minimum area of the drill pipe can be determined. In addition, the present invention provides a reliable basis for the precise positioning of the drill pipe through an accurate mapping process from image coordinates to camera coordinates and then from camera coordinates to world coordinates, ensuring an accurate conversion from the image space to the actual physical space. The present invention is not only applicable to a simple background, but also can accurately detect the position and attitude of the drill pipe in a complex background, ensuring the accuracy of coordinate system calibration.
[0046] (2) The segmentation network of the present invention is based on the UNet architecture and incorporates a multi-scale feature reconstruction module, enabling the drill pipe panoramic image segmentation network model to extract multi-scale receptive field information from each layer of the encoder, thereby significantly enhancing the feature expression ability of the network model. In the encoder, a depthwise separable convolution module is used to replace part of the traditional convolution operations, significantly reducing the computational amount of the model. In addition, the present invention also introduces a hybrid multi-scale attention module to process the features of each layer of the encoder and transfer the processed features to the corresponding layer of the decoder. This design enables the network model to generate a more accurate drill pipe segmentation image, thereby improving the accuracy and efficiency of drill pipe segmentation in a changing environment and background. After segmentation, the present invention uses the Canny edge detection algorithm to extract the edge coordinate point set of the drill pipe, further enhancing the ability to accurately identify and position the drill pipe. Description of the Drawings
[0047] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in combination with the drawings, where:
[0048] Figure 1Flowchart of a positioning method for a trenchless drilling rig drill pipe based on image segmentation in a preferred embodiment of the present invention;
[0049] Figure 2 Structural diagram of a drill pipe panoramic image segmentation network model in a preferred embodiment of the present invention;
[0050] Figure 3 Structural diagram of a depthwise separable convolution module in a preferred embodiment of the present invention;
[0051] Figure 4 Structural diagram of a hybrid multi-scale attention module in a preferred embodiment of the present invention;
[0052] Figure 5 Flowchart of multi-scale feature reconstruction in a preferred embodiment of the present invention;
[0053] Figure 6 Flowchart of the complete visual processing from the original image to the final positioning in a preferred embodiment of the present invention. Detailed implementation manners
[0054] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited are not intended to limit the present invention. Embodiment 1
[0055] Refer to Figure 1 As shown, an embodiment of the present invention provides a positioning method for a trenchless drilling rig drill pipe based on image segmentation, including but not limited to the following steps:
[0056] S1. Obtain an image dataset of the drill pipe;
[0057] S2. Construct a drill pipe panoramic image segmentation network model, and use the image dataset to train the drill pipe panoramic image segmentation network model to obtain an optimized image segmentation network model;
[0058] S3. Obtain a drill pipe panoramic image of the drill pipe to be positioned, and input the drill pipe panoramic image into the optimized image segmentation network model to generate a segmentation result image;
[0059] S4. Extract the edge features of the segmentation result image to obtain an edge coordinate point set of the drill pipe to be positioned;
[0060] S5. Calculate the edge coordinate point set to obtain the farthest point pair; calculate the distance between the farthest point pair to obtain the minimum area rectangle;
[0061] S6. Obtain the center coordinates of the minimum area rectangle, convert the center coordinates to camera coordinates, and then convert the camera coordinates to world coordinates to obtain the position information of the drill pipe to be positioned.
[0062] An embodiment of the present invention provides a positioning method for a non-excavation drill rod based on image segmentation. Through the constructed image segmentation network model, the drill rod can be automatically and accurately segmented from the panoramic image, reducing manual operations and ensuring the consistency and accuracy of the segmentation results. By calculating the edge coordinate point set of the segmented result image, the farthest point pair and the minimum area rectangle can be determined, which is crucial for accurately calculating the position and attitude of the drill rod. In addition, the accurate mapping from image coordinates to camera coordinates and then to world coordinates provides a reliable basis for the precise positioning of the drill rod. The embodiment of the present invention realizes high-precision image recognition through deep learning, and combines coordinate transformation algorithms to obtain the spatial position information of the drill rod, realizing the accurate detection of the position and attitude of the drill rod under complex backgrounds and ensuring the accuracy of coordinate system calibration. In addition, the obtained spatial position information can be used for the subsequent automatic positioning and assembly of the drill rod, thereby greatly improving the operation efficiency and safety.
[0063] Specifically, for step S1, an image dataset is constructed and image preprocessing is performed. The specific steps for preprocessing the dataset are as follows:
[0064] S101. Image grayscale conversion and data augmentation. All images in the dataset are converted into grayscale images to meet the single-channel input requirement of the network model. Then, data augmentation strategies are implemented, including operations such as image rotation, translation, cropping, and scaling, to increase the diversity of the dataset and improve the generalization ability of the model.
[0065] S102. Image size unification and contrast enhancement. All images in the dataset are adjusted to a unified size to ensure the consistency of the input data, thereby simplifying the model training process. At the same time, the image brightness and contrast are enhanced to improve the visibility of the drill rod features in the image and further improve the effect of feature extraction.
[0066] S103. Semantic segmentation annotation. Semantic segmentation is a technique for image segmentation at the pixel level. It classifies each pixel in the image separately to achieve precise distinction between the drill rod and the background. This detailed segmentation method greatly improves the accuracy of the target detection task and the robustness of the system, making the detection under complex backgrounds more reliable. In this embodiment, the labelme script is used to perform detailed semantic segmentation annotation on the enhanced panoramic image of the drill rod, that is, the first image obtained from steps S101 to S102.
[0067] S104. Dataset division. The annotated panoramic images of drill pipes and their corresponding label data are divided into a training set and a test set according to a preset ratio, such as a 9:1 ratio. This division ratio helps to effectively evaluate the generalization performance of the model while ensuring that the model can fully learn. In this step, the preset ratio (i.e., the division ratio) of the training set to the test set can be flexibly adjusted according to specific application requirements and the characteristics of the dataset.
[0068] Specifically, for step S2, the drill pipe panoramic image segmentation network model (abbreviated as MDH-UNet segmentation model) includes an encoder, a multi-scale feature reconstruction module (Multi-Scale Feature Reconstruction Module, abbreviated as MFRM), and a decoder. Refer to Figure 2 , the construction method and running steps of the drill pipe panoramic image segmentation network model are as follows:
[0069] S201. The encoder includes 4 layers of depthwise separable convolution modules (Depthwise Separable Convolution, abbreviated as DwConv), 3×3 convolution modules, and hybrid multi-scale attention modules (Hybrid Multi-Scale AttentionModule, abbreviated as HMAM). These modules perform downsampling in sequence to obtain the drill pipe feature map.
[0070] Specifically, refer to Figure 3 , the depthwise separable convolution module includes 3×3 DwConv, a pointwise convolution 1×1Conv, two batch normalization BNs, and two activation functions ReLU. Refer to Figure 4 , the hybrid multi-scale attention module includes a channel attention module and a spatial attention module connected in series.
[0071] Furthermore, refer to Figure 4 , the working steps of the channel attention module are as follows: The feature image F 1 First, the input feature map is processed by convolutional layers of 3×3 and 5×5 respectively to capture receptive fields of different scales. Next, these feature maps of different scales are fused and further refined through an average pooling layer (Average Pool) and a fully connected layer (Fully Connectedlayer, abbreviated as FC). Subsequently, the refined feature map is remapped back to the original 3×3 and 5×5 receptive field feature maps. Finally, through an element-wise summation operation, the mapped feature layers are fused to output an enhanced channel feature map.
[0072] Furthermore, refer to Figure 4 , the working steps of the spatial attention module are as follows: The feature image F 2First, the channel dimension of the feature map is reduced by the Max Pool and Average Pool operations to extract more representative features. Then, the feature map is convolved using 3×3, 5×5, and 7×7 convolution kernels to capture spatial features of different scales. Subsequently, these multi-scale features are fused through element-level summation to form a comprehensive multi-scale feature representation. Then, the multi-dimensional feature map is converted into a one-dimensional feature vector using the Flatten operation, and the feature vector is normalized using the SoftMax function to highlight important spatial locations. The normalized feature vector is fused element-by-element with the original feature map to further enhance the spatial information of the feature map. Finally, the fused feature map is integrated through the element-level summation operation to output the final spatial feature map.
[0073] S202, MFRM is responsible for upsampling the drill rod feature maps from different layers of the encoder to the same size. This step is to ensure that feature maps of different scales can be compared and fused at the same resolution. The upsampled feature map is then reconstructed at multiple scales in MFRM, which helps to integrate feature information of different scales and enhance the expressiveness of features. The reconstructed feature map (i.e., the reconstructed feature map) is finally input to the corresponding layer of the decoder and fused with the feature map in the decoder to achieve more accurate image segmentation.
[0074] Further, refer to Figure 5 , the multi-scale reconstruction method is:
[0075] ; ;
[0076] in, represents the multi-scale reconstruction feature, represents the fused features of convolutional layers 1 to 4 in the encoder, represents upsampling, represents a convolution with a kernel size of 1, represents a dilated convolution with a dilation rate of a and a kernel size of b, and F 1 , F 2 , F 3 and F 4 They come from the features of convolutional layers 1 to 4 in the encoder respectively. For example, the value of a can be 8, 4, 2, 1; the value of b can be 3. Figure 5 middle, Represents the final multi-scale reconstruction features.
[0077] S203. The decoder includes four convolutional layers and performs upsampling in sequence. At the same time, the output of each layer in the encoder is directly passed to the corresponding layer in the decoder through skip connections. This design allows the decoder to utilize high-resolution feature information from the encoder when reconstructing the image, thereby retaining more detail and edge information in the final segmentation result.
[0078] S204. The output of the last upsampling is fed into the first convolutional module, which includes a BatchNorm regularization layer, a ReLU activation function, and a 3×3 convolutional layer. The feature map output by the first convolutional module then passes through a Sigmoid activation function to obtain the final segmentation result.
[0079] Through the design of steps S201 to S204, the network model realizes an automated feature extraction, fusion, and segmentation process, reducing manual intervention and improving the automation and intelligence level of detection.
[0080] Furthermore, the above-constructed drill pipe panoramic image segmentation network model is trained using the image dataset preprocessed in step S1 to obtain an optimized image segmentation network model. The specific training method is as follows: In the model training stage, the MDH-UNet segmentation model is iteratively trained using the Adam optimization algorithm and backpropagation technology. Using the training set data, continue until the total loss function reaches a convergence state, thereby obtaining a well-trained drill pipe panoramic image segmentation model. Given that the Dice loss function can effectively measure the similarity between the prediction result and the true label, and then guide model training to improve the segmentation accuracy. Therefore, the Dice loss function is preferably used in the training process. The Dice loss function has the following calculation formula:
[0081] ;
[0082] where X represents the true value, Y represents the predicted value.
[0083] Specifically, considering that the Canny edge detection algorithm adopts a multi-level detection mechanism, including steps such as Gaussian filtering, gradient calculation, non-maximum suppression, and double-threshold detection, these steps work together to ensure the accuracy of edge detection. The Gaussian filtering step effectively reduces the noise in the image and avoids the influence of noise on the edge detection results. Secondly, the algorithm can calculate the intensity of the edges, which helps to distinguish edges of different importance and thus more accurately identify the key edges in the image. Therefore, in step S4, the Canny edge detection algorithm is preferably used to extract the edge feature of the segmentation result image output by the optimized image segmentation network model in step S3. Among them, obtaining the edge coordinate point set of the drill pipe is expressed as P={(u 1 ,v 1 ),(u 2 ,v 2 ),..,(u n ,v n )}.
[0084] Specifically, for step S5, the rotating calipers method is used to locate the minimum area rectangle according to the edge coordinate point set. The specific steps are as follows:
[0085] S501. Identify the minimum x-coordinate point (i.e., the leftmost point) and the maximum x-coordinate point (i.e., the rightmost point) in the edge coordinate point set, and connect these two points to form a horizontal line.
[0086] S502. For each point in the point set, calculate its perpendicular distance to this line, and mark the point with the farthest distance, and then obtain the two points with the farthest perpendicular distances on both sides of this line, that is, the pair of points with the farthest distance.
[0087] S503. Rotate this horizontal line counterclockwise by an angle so that the line segment formed by the previously marked pair of farthest points coincides with the new line. Repeat this rotation process until the line rotates one week. During each rotation process, record the distance between the pair of farthest points, and calculate the corresponding rectangle area accordingly. By comparing the rectangle areas obtained at all rotation angles, select the rectangle with the smallest area, and its corresponding rotation angle is the optimal angle sought. The minimum area rectangle obtained in this way has the minimum area A = w×h, where w and h are the width and height of the rectangle respectively. This step ensures that the minimum area rectangle that best fits the edge of the drill pipe can be found, providing accurate geometric information for subsequent analysis and processing.
[0088] S504. Obtain the center point coordinates of the minimum area rectangle frame and the relative rotation angle , and the geometric center coordinates of the minimum rectangle frame are expressed as:
[0089] ;
[0090] Among them, represents the abscissa of the center point, represents the ordinate of the center point, and respectively represent the minimum and maximum x - coordinate values of the rectangle in the horizontal direction. and respectively represent the minimum and maximum y - coordinate values of the rectangle in the vertical direction.
[0091] The rotation angle of the minimum rectangle , that is, the angle of counter - clockwise rotation of the longer side relative to the horizontal line, and the mathematical expression is:
[0092] ;
[0093] Among them, represents the change in the y - coordinate, represents the change in the x - coordinate. represents the arctangent function.
[0094] Through steps S501 to S504, the optimized detection and positioning of the drill pipe edge are achieved. This series of steps uses mathematical calculations and optimization strategies, significantly enhancing the efficiency and accuracy of detection and positioning, and laying a solid foundation for automated processing and analysis. Among them, the mathematical calculations in step S504 ensure the correct alignment and positioning of the drill pipe in the image, improving the accuracy and reliability of the system.
[0095] In this embodiment, the complete visual processing flow from the original image to the final positioning can be referred to Figure 6 as shown. From Figure 6 it can be clearly seen the visualization results of each step: (a) presents the original drill pipe image, (b) shows the result after semantic segmentation processing, where different regions are clearly distinguished and marked; (c) shows the image edges obtained by the edge detection algorithm, and these edges outline the important features in the image; finally, (d) depicts the minimum rectangle obtained based on the previous steps, and this rectangle precisely encloses the drill pipe area, facilitating subsequent analysis and processing.
[0096] Specifically, for step S6, the center point coordinates of the obtained minimum - area rectangle are converted into a point in the camera coordinate system. Through this conversion, the accurate position information of the object in the three - dimensional space can be obtained from the two - dimensional image space, providing the necessary geometric data for subsequent analysis and processing. The specific steps of the conversion are:
[0097] S601. Use the internal parameter matrix of the camera for conversion. The internal parameter matrix is a key conversion tool that contains the intrinsic parameters of the camera, such as focal length, principal point coordinates, and pixel size. These parameters are crucial for accurately mapping image coordinates to the camera coordinate system. In this embodiment, the internal parameter matrix has the following expression:
[0098] ;
[0099] where and represent the number of pixels in the x-axis and y-axis directions per millimeter respectively; and represent the positions of the principal point of the image (usually the center of the image) in the pixel coordinate system respectively.
[0100] S602. Use the internal parameter matrix to convert the point in the image coordinate system to the point in the normalized camera coordinate system. The calculation formula is:
[0101] .
[0102] S603. Convert the point in the normalized camera coordinate system to the actual camera coordinate system. In this process, use the external parameter matrix of the camera for conversion. This matrix includes the rotation matrix and the translation vector . The mathematical expression is: M = R | t . The mathematical expressions of the rotation matrix and the translation vector are:
[0103] ;
[0104] where with different subscripts represents different components of the matrix; with different subscripts represents different components of the vector.
[0105] S604. Further convert the point in the camera coordinate system to the point in the world coordinate system. This process involves the conversion from camera coordinates to world coordinates, which is crucial for achieving accurate spatial positioning and 3D reconstruction. In this embodiment, assume that the depth is known. Then the mathematical expression of the point in the camera coordinate system is:
[0106] .
[0107] Therefore, the point in the world coordinate system is obtained The mathematical expression is:
[0108] .
[0109] Through the above steps S601 to S604, the center point of the drill pipe identified in the image coordinate system is accurately transformed into the world coordinate system, so as to obtain the specific position coordinates of the drill pipe in the world coordinate system . This process ensures the accuracy of drill pipe positioning and provides a solid foundation for subsequent automated operations and analyses.
[0110] The segmentation network based on the UNet architecture is used in the embodiment of the present invention, and a multi-scale feature reconstruction module is incorporated. This module can extract multi-scale receptive field information from each layer of the encoder and enhance the feature expression ability. To improve the computational efficiency of the network, a depthwise separable convolution module is used in the encoder to replace some traditional convolution operations. This replacement significantly reduces the computational amount while extracting spatial features and channel features. In addition, a hybrid multi-scale attention module is introduced into the model to process the features of each layer of the encoder, and then the processed features are passed to the corresponding layer of the decoder. This process can finally generate an accurate drill pipe segmentation image, thereby improving the accuracy and efficiency of drill pipe segmentation in different scenarios and backgrounds. After segmentation, the Canny edge detection algorithm is further used to extract the edge coordinate point set of the drill pipe. Through these point sets, the minimum circumscribed rectangle of the drill pipe can be determined, and the center point coordinates, width and height dimensions, and relative rotation angle of the rectangle can be obtained. Finally, based on the internal parameters of the camera, this information is converted into the coordinates of the drill pipe in the real world to achieve accurate recognition and positioning of the drill pipe. This series of processes not only improves the accuracy of segmentation but also provides a solid foundation for subsequent recognition and positioning. Embodiment 2
[0111] Based on the same inventive concept, this embodiment provides a positioning system for a trenchless drilling rig drill pipe based on image segmentation. The principle of solving the problem is similar to that of a positioning method for a trenchless drilling rig drill pipe provided in Embodiment 1, and the repeated parts will not be elaborated.
[0112] This embodiment provides a positioning system for a trenchless drilling rig drill pipe based on image segmentation, including:
[0113] A data set acquisition module, configured to acquire an image data set of the drill pipe;
[0114] A model construction module, configured to construct a drill pipe panoramic image segmentation network model, and train the drill pipe panoramic image segmentation network model by using the image data set to obtain an optimized image segmentation network model;
[0115] The segmented image output module is used to obtain the panoramic image of the drill pipe to be located, input the panoramic image of the drill pipe into the optimized image segmentation network model, and generate a segmented result image;
[0116] The feature extraction module is used to extract the edge features of the segmented result image and obtain the edge coordinate point set of the drill pipe to be located;
[0117] The rectangle acquisition module is used to calculate the edge coordinate point set to obtain the farthest point pair; calculate the distance between the farthest point pairs to obtain the minimum area rectangle;
[0118] The positioning module is used to obtain the center coordinates of the minimum area rectangle, convert the center coordinates into camera coordinates, and then convert the camera coordinates into world coordinates to obtain the position information of the drill pipe to be located.
[0119] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0120] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0121] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks. Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for realizing the functions specified in one block or a plurality of blocks.
[0123] Obviously, the above-described embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to exhaustively list all the implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. A method for positioning a drill rod of a trenchless drilling rig based on image segmentation, characterized in that: include: Acquire an image dataset of the drill pipe; A drill rod panoramic image segmentation network model is constructed, and the drill rod panoramic image segmentation network model is trained using the image data set to obtain an optimized image segmentation network model; wherein the drill rod panoramic image segmentation network model includes an encoder and a multi-scale feature reconstruction module; the encoder includes a depthwise separable convolution module, and the depthwise separable convolution module includes depthwise convolution and pointwise convolution, which are used to obtain a drill rod feature map; the multi-scale feature reconstruction module is used to perform multi-scale reconstruction on the drill rod feature maps of different scales; the multi-scale reconstruction method of the multi-scale feature reconstruction module is: ; ; in, represents the multi-scale reconstruction feature, represents the fusion feature, represents upsampling, represents a dilated convolution with a dilation rate of a and a convolution kernel size of b. F 1 , F 2 , F 3 and F 4 They come from the features of convolutional layers 1 to 4 in the encoder respectively; Acquire a drill rod panoramic image of the drill rod to be located, input the drill rod panoramic image into the optimized image segmentation network model, and generate a segmentation result image; Extracting edge features of the segmentation result image to obtain an edge coordinate point set of the drill rod to be positioned; Calculating the edge coordinate point set to obtain the farthest point pair; calculating the distance between the farthest point pairs to obtain a minimum area rectangle; The center coordinates of the minimum area rectangle are obtained, the center coordinates are converted into camera coordinates, and then the camera coordinates are converted into world coordinates to obtain the position information of the drill rod to be positioned.
2. The method for positioning a drill rod of a trenchless drilling rig based on image segmentation according to claim 1, characterized in that: The drill rod panoramic image segmentation network model also includes a decoder; the relationship between the decoder, the encoder and the multi-scale feature reconstruction module is: the encoder performs downsampling to obtain a drill rod feature map; the multi-scale feature reconstruction module performs multi-scale reconstruction on the drill rod feature maps of different scales to obtain a reconstructed feature map; the decoder fuses the reconstructed feature map.
3. The method for positioning a drill rod of a trenchless drilling rig based on image segmentation according to claim 1, characterized in that: The encoder also includes a convolution module and a hybrid multi-scale attention module; the hybrid multi-scale attention module includes a channel attention module and a spatial attention module; the channel attention module is connected to the spatial attention module.
4. The method for positioning a drill rod of a trenchless drilling rig based on image segmentation according to claim 1, characterized in that: The method for training the drill rod panoramic image segmentation network model using the image data set is: dividing the image data set into a training set according to a preset ratio, and iteratively optimizing the drill rod panoramic image segmentation network model based on the training set using the Adam back-propagation algorithm until the total loss function converges.
5. The method for positioning a drill rod of a trenchless drilling rig based on image segmentation according to claim 4 is characterized in that: The calculation formula of the total loss function is: ; in, represents the total loss, X represents the true value, Y Represents the predicted value.
6. The method for positioning a drill rod of a trenchless drilling rig based on image segmentation according to claim 1, characterized in that: The obtaining of the drill pipe image data set includes preprocessing each image in the image data set, and the preprocessing method is: Gray-scaling, data enhancement, size unification and contrast enhancement are performed on the image to obtain a first image; Use the labelme script to perform semantic segmentation annotation on the first image.
7. The method for positioning a drill rod of a trenchless drilling rig based on image segmentation according to claim 1, characterized in that: The edge coordinate point set is calculated to obtain the farthest point pair; the distance between the farthest point pairs is calculated to obtain the minimum area rectangle as follows: Identify the minimum x-coordinate point and the maximum x-coordinate point in the edge coordinate point set, and connect the two points to obtain a horizontal straight line; Calculate the vertical distance from each point in the edge coordinate point set to the horizontal line, mark the point pair with the farthest distance, where the point pair with the farthest distance is the two points on both sides of the horizontal line with the farthest vertical distance; rotate the horizontal line counterclockwise by an angle so that the rotated horizontal line coincides with the line segment between the farthest point pair; repeat the rotation process until the horizontal line is rotated once, record the distance between the farthest point pairs during each rotation process, and calculate the corresponding rectangular area; Compare the areas of rectangles obtained at all rotation angles and get the rectangle with the smallest area.
8. The method for positioning a drill rod of a trenchless drilling rig based on image segmentation according to claim 1, characterized in that: The method of obtaining the center coordinates of the minimum area rectangle, converting the center coordinates into camera coordinates, and then converting the camera coordinates into world coordinates is: The center coordinates are converted into points in the normalized camera coordinate system using the intrinsic parameter matrix. The calculation formula is: ; in, represents a point in the normalized camera coordinate system, represents the horizontal coordinate of the center point, represents the vertical coordinate of the center point, represents the internal parameter matrix; Using the camera's extrinsic matrix, the point in the normalized camera coordinate system is converted into a point in the camera coordinate system; According to the conversion of the camera coordinates into the world coordinates, the calculation formula of the world coordinates is: ; in, represents the rotation matrix, represents the translation vector; , and represents the coordinates of the midpoint of the camera coordinates; , and Represents the coordinates of the midpoint in the world coordinate system.
9. The method for positioning a drill rod of a trenchless drilling rig based on image segmentation according to claim 1, characterized in that: The method for extracting edge features of the segmentation result image is the Canny edge detection algorithm.
10. A positioning system for a trenchless drilling rig drill rod based on image segmentation, characterized in that: include: A data set acquisition module, used to acquire an image data set of the drill pipe; A model building module is used to build a drill rod panoramic image segmentation network model, and the drill rod panoramic image segmentation network model is trained using the image data set to obtain an optimized image segmentation network model; wherein the drill rod panoramic image segmentation network model includes an encoder and a multi-scale feature reconstruction module; the encoder includes a depthwise separable convolution module, and the depthwise separable convolution module includes depthwise convolution and pointwise convolution, which are used to obtain a drill rod feature map; the multi-scale feature reconstruction module is used to perform multi-scale reconstruction on the drill rod feature maps of different scales; the multi-scale reconstruction method of the multi-scale feature reconstruction module is: ; ; in, represents the multi-scale reconstruction feature, represents the fusion feature, represents upsampling, represents a dilated convolution with a dilation rate of a and a convolution kernel size of b. F 1 , F 2 , F 3 and F 4 They come from the features of convolutional layers 1 to 4 in the encoder respectively; A segmented image output module is used to obtain a drill rod panoramic image of the drill rod to be located, input the drill rod panoramic image into the optimized image segmentation network model, and generate a segmentation result image; A feature extraction module, used to extract edge features of the segmentation result image and obtain an edge coordinate point set of the drill rod to be positioned; A rectangle acquisition module is used to calculate the edge coordinate point set to obtain the farthest point pair; calculate the distance between the farthest point pairs to obtain the minimum area rectangle; The positioning module is used to obtain the center coordinates of the minimum area rectangle, convert the center coordinates into camera coordinates, and then convert the camera coordinates into world coordinates to obtain the position information of the drill rod to be positioned.
Citation Information
Patent Citations
Remote sensing image road segmentation method based on contextual information and multi-scale feature fusion
CN113850825A
Manipulator grabbing method based on deep learning target detection and image segmentation
CN115816460A
Crop leaf disease image semantic segmentation method
CN119169289A