Underground pipeline abnormal area positioning method based on stereoscopic visual perception

Through a method based on stereoscopic visual perception, using binocular cameras and IGEV networks for image processing and stereo matching, the problems of insufficient ambient light interference, stability and flexibility in the prior art are solved, and more efficient and accurate positioning of abnormal areas of underground pipelines is achieved.

CN120198503AInactive Publication Date: 2025-06-24YANCHENG POWER SUPPLY CO STATE GRID JIANGSU ELECTRIC POWER CO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510264120.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing underground pipeline abnormal area positioning technology cannot eliminate ambient light interference, system stability and flexibility, resulting in low detection efficiency and safety hazards.

Method used

Using a method based on stereoscopic visual perception, image clarity correction, multi-level feature-guided abnormal area recognition and binocular stereo matching are performed through binocular cameras and IGEV networks, and a full resolution parallax map is output to locate foreign objects.

Benefits of technology

It achieves faster and accurate abnormal area positioning of underground pipelines, improves the degree of automation and environmental adaptability, and reduces the safety hazards and inefficiency of manual inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198503A_ABST
    Figure CN120198503A_ABST
Patent Text Reader

Abstract

The invention discloses an underground pipeline abnormal area positioning method based on stereoscopic visual perception. Firstly, a binocular camera is used for collecting an image and carrying out sharpness correction on the image; performing foreign matter identification on the corrected image by using a multi-level feature guided underground pipeline abnormal area identification method; positioning a foreign matter according to a foreign matter recognition result, and performing binocular stereo matching by using an IGEV network to output a full-resolution disparity map C; and finding out a foreign matter contour in the full-resolution disparity map C, positioning coordinates of four points at the outermost edge in the contour, obtaining three-dimensional coordinates based on a coordinate conversion principle, and completing foreign matter measurement. According to the invention, foreign matters in the cable duct can be timely and accurately detected and positioned, key data information is provided for subsequent operation and maintenance of the underground cable duct, safe operation of the underground cable duct is better guaranteed, and the method has very important significance and value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method for locating abnormal areas of underground pipelines based on stereo vision perception. Background Art

[0002] With the rapid development and construction of cities, the number of underground cable pipelines has been increasing year by year. They undertake the key tasks of transmitting electricity and information, which are directly related to people's electricity use safety and the stable operation of the power grid. However, during the laying process of underground cables, foreign objects such as urban garbage and small stones are likely to enter the pipeline interior, resulting in pipeline blockage. In severe cases, it may even cause cable damage, leading to serious consequences such as power outages and short circuits, severely restricting the development pace of cities. Therefore, it is of great significance and value to detect and remove foreign objects inside the cable pipelines in a timely and accurate manner.

[0003] Currently, the detection and location of foreign objects are mainly manual methods, which require a large amount of manpower and material resources, and there are problems such as low efficiency and certain safety hazards. With the development of intelligent technologies, pipeline abnormal area detection schemes based on stereo vision systems have gradually emerged. By fusing cameras and sensors, accurate capture of abnormal positions is achieved, significantly improving the pipeline detection efficiency. However, existing detection technologies still have problems such as being unable to exclude the interference of ambient light, and the system stability and flexibility are limited. Therefore, how to improve the automation level and environmental adaptability while ensuring the detection accuracy of foreign objects, making the detection more efficient and convenient, is the key to solving the application problem of underground pipeline abnormal area location technology in actual scenarios. Summary of the Invention

[0004] In view of this, the technical problem actually to be solved by the present invention is: how to develop a method for locating abnormal areas of underground pipelines based on stereo vision perception in a faster and more accurate manner to better ensure the safe operation of underground cable pipelines.

[0005] The above object is achieved through the following technical solutions:

[0006] A method for locating abnormal areas of underground pipelines based on stereo vision perception, comprising the following steps:

[0007] Step (1) Perform image sharpening and calibration;

[0008] Step (2) Use a method for identifying abnormal areas of underground pipelines guided by multi-level features to identify foreign objects in the image calibrated in step (1);

[0009] Step (3) Locate foreign objects according to the foreign object identification result in step (2), and use the IGEV network for binocular stereo matching to output a full-resolution disparity map C;

[0010] Step (4) finds the foreign object contour in the full-resolution disparity map C output in step (3) and locates the coordinates of the four outermost points in the contour, and obtains the three-dimensional coordinates based on the coordinate transformation principle to complete the foreign object measurement.

[0011] Further, the image sharpening and correction in step (1) specifically includes the following steps:

[0012] 11) Use a binocular camera to capture checkerboard images at different positions, obtain 20 left-eye and right-eye images each, and based on the Zhang Zhengyou calibration method, obtain the internal parameter matrices M1, M2 and rotation matrices R1, R2 of the left and right cameras;

[0013] 12) Use a binocular camera to collect foreign object images inside the pipeline, where the left-eye image is C left , and the right-eye image is C right . Then use an epipolar correction method based on the singular value decomposition of the camera translation matrix for epipolar correction to obtain the corrected images C left ′, C right ′. The specific steps of the epipolar correction method based on the singular value decomposition of the camera translation matrix are as follows:

[0014] a) Calculate a new rotation matrix R n , and its formula is as follows. The corrected rotation matrices of the left and right cameras are R n , the image poles are at infinity and the epipolar lines are parallel;

[0015]

[0016] where T is the relative translation vector between the binocular cameras, and R x is the rotation matrix rotating by w angle around the X axis, and its size is U1, U2 are the abscissas of the actual principal points of the left and right eyes, and Q is the correction matrix, and its size is where

[0017] b) Determine the new internal parameter matrix M n of the corrected binocular camera;

[0018] First, determine the tilt factor S. Assume that the origin in C left is O(0,0,1), the row corner point A(M,0,1), the column corner point B(0,N,1), and the center point After these 4 points are transformed by H1, they are O1, A1, B1, G1 respectively. A1O1 = (x1, y1, 0), B1O1 = (x2, y2, 0), then

[0019] Then adjust the focal length, move the corrected image to the center of the field of view, and let The G1 coordinates are (p, q, l). L, H, X, and Y are intermediate parameters without actual physical meaning. Then, the new internal parameter matrix M n is calculated as follows:

[0020]

[0021] c) For each image point P1, P2 in the original pixel coordinate system, the image points m1 = M1 -1 P1, m2 = M2 -1 P2 in the form of normalized coordinates are obtained. Then, through the transformation matrices R L , R R the image points m1, m2 in the form of normalized coordinates are transformed into the corrected camera coordinate system, and multiplied by the reciprocals α1, α2 of the Z vectors of the object points in the corrected camera coordinate system, obtaining the normalized coordinates m1' = α1R L m1, m2' = α2R R m2, where the transformation matrices R L , R R are the results obtained by the left and right eye cameras rotating by angles around the Z axis, then rotating by θ angles around the Y axis, and finally rotating by w angles around the X axis.;

[0022] d) Through the new internal parameter matrix M n the points m1', m2' are transformed onto the new pixel coordinate system, obtaining the corrected points P1' = M n m1', P2' = M n m2', where the expressions of the transformation matrices H1, H2 for the left and right image corrections are respectively:

[0023]

[0024] Finally, the corrected images C left ', C right ' are obtained.

[0025] Furthermore, step (2) specifically includes the following steps:

[0026] 21) In the search stage, first, feature extraction is performed, that is, based on the ResNet50 neural network, multi-level features f left are extracted from the corrected left-eye image C m , m ∈ {1, 2, 3, 4, 5}, and the resolution of each feature is respectively m ∈ {1, 2, 3, 4, 5}, where H1 and W respectively represent the sizes of three scales of height and width; then, the last three layers of features f3, f4, and f5 are respectively input into the texture enhancement module, and through a group of dilated convolutions with different dilation rates, the corresponding features are extracted by imitating the grouped receptive fields of different sizes in the human visual system, and f3′, f4′, and f5′ are output;

[0027] 22) Input f3′, f4′, and f5′ into the adjacent connection decoding module to obtain the rough map C6 of the target. The specific steps are as follows:

[0028] a) The three layers of features f3′, f4′, and f5′ are input into the adjacent connection decoding module for the following processing to obtain

[0029]

[0030] where are the weights of each layer, u ∈ {1, 2, 3}; represents a 3×3 convolution and normalization operation, and Upsample 2 represents two upsamplings, represents pixel-wise multiplication;

[0031] b) Combine to obtain the rough map C6 of the target;

[0032] 23) In the confirmation stage, first obtain the reverse attention guidance from the rough map C6 through the Sigmoid activation function and the reverse operation The formula is as follows;

[0033]

[0034] where Upsample 2 represents two upsamplings, Downsample 4 represents four downsamplings, Reverse represents the inversion operation, E is the identity matrix with all element values being 1, and σ is the Sigmoid activation function.

[0035] 24) Take and f5′ as inputs, and successively pass through three grouped - reverse attention modules, combine the reverse attention guidance and the grouped guidance with the residual learning process, combine multiple grouped - reverse attention modules, gradually refine the rough prediction through feature layers of different depths, and then combine with the rough map C6 of the target to output the result C5;

[0036] 25) Repeat this step twice based on the output of step 24), and sequentially output results C4 and C3. Finally, normalize C3 using the Sigmoid function to output the recognition result GT of the multi-level feature-guided underground pipeline anomaly area recognition method. left ; right-eye image C right ′ Similarly, operate steps 21)-24) to output GT right ;

[0037] 26) Input GT left and GT right into the classification module, including a fully connected layer and a Softmax activation function layer, to obtain the foreign object classification result. The specific steps are as follows:

[0038] a) Flatten GT left and GT right into one-dimensional data respectively and input them into 3 fully connected layers, gradually adjusting the output size to be equal to the number of categories;

[0039] b) Input into the Softmax layer for classification to obtain the predicted category of each pixel, and output the final classification result map GT left ′ and GT ri ht ′:

[0040]

[0041] where z i is the input of the Softmax layer, Q is the number of label types, y i is the probability that the predicted pixel belongs to the Qth class. Finally, select the class with the highest probability as the final prediction result;

[0042] In the multi-level feature-guided underground pipeline anomaly area recognition method, the cross-entropy loss function is selected as the loss function. The cross-entropy function CE(h, s) and the loss function loss calculation formulas are as follows:

[0043]

[0044] where Q is the number of label types, a is the number of input pixels, h is the true label distribution of the pixel, s is the predicted label result of the pixel, represents the j′th category of the i′th sample, represents the probability that the predicted i′th sample belongs to the j′th class.

[0045] Further, step (3) specifically includes the following steps:

[0046] 31) First, locate the estimated area of the foreign object according to the recognition result of the multi-level feature-guided underground pipeline anomaly area recognition method, that is, GT left″ = C left ′ · GT left ′ and GT right ″ = C right ′ · GT right ′;

[0047] 32) Input GT left ″ and GT right ″ into the feature extractor of the IGEV network, which includes a feature extraction network and a context network. The specific steps are as follows:

[0048] a) In the feature extraction network, downsample the input graphs GT left ″ and GT right ″ to 1 / 32, and then obtain multi-scale features through upsampling Select f l,4 and f r,4 to construct the cost volume, C e 、H1、W represent the sizes of the three scales of channels, height, and width respectively;

[0049] b) In the context network, through a series of residual blocks and downsampling layers, generate multi-scale context features at 1 / 4, 1 / 8, and 1 / 16 of the image resolution with 128 channels as input, denoted as c k 、c r 、c h ;

[0050] 33) Construct the grouped correlation cost volume, and the steps are as follows:

[0051] a) Divide f l,4 and f r,4 into N g = 8 groups according to the channel dimension, and calculate the correlation maps of each group to form a 4D correlation cost volume C corr , and its formula is as follows:

[0052]

[0053] where <·,·> is the vector inner product, d is the disparity index, N c is the number of channels, N g is the number of groups, and x, y represent the horizontal and vertical coordinates of the pixels;

[0054] b) Use a lightweight 3D regularization network to perform cost aggregation on C corr and output the geometric encoding volume C G ; At the same time, insert the guiding cost volume excitation operation to update C e; The 3D regularization network contains 3 downsampling modules and 3 upsampling modules. The downsampling module contains two 3D convolutions of 3×3×3, with the number of channels being 16, 32, and 48 respectively. The upsampling module contains a transposed convolution of 4×4×4 and two 3D convolutions of 3×3×3. The formula for cost aggregation is as follows:

[0055] C G = R(C corr )

[0056] where C G represents the geometric encoding volume obtained after cost aggregation, and R(·) represents the 3D regularization network;

[0057] During the cost aggregation process, for a cost volume C (e = 4, 8, 16, 32), the guiding cost volume excitation formula is as follows: e (e = 4, 8, 16, 32), the guiding cost volume excitation formula is as follows:

[0058]

[0059] where D represents the total number of disparity points, C e ' represents the result obtained after the guiding cost volume excitation of the cost volume C e , σ represents the Sigmoid function, is the Hadamard product;

[0060] c) Calculate the corresponding left and right image feature differences to obtain the local feature correlation cost volume C A , and then expand the receptive field, that is, use 1D average pooling with both the size and stride of 2 to obtain two levels of C G pyramid and C A pyramid;

[0061] 34) Use a 3-level ConvGRU-based Update Operator (ConvGRU) for iterative update to output the final disparity d k+1 , and the specific steps are as follows:

[0062] 35) Convolve the hidden state h k to generate features, then upsample them to 1 / 2 resolution, and cascade them with f i,2 from the left image as the weight μ ∈ R H×W×9 , and finally output the full-resolution disparity map C through the weighted combination of the predicted disparity d k at 1 / 4 resolution.

[0063] Furthermore, the specific steps in step (4) include the following steps:

[0064] 41) Find the foreign object contour in the disparity map C and locate the coordinates of the four outermost points on the contour, denoted as L(x A , y A ), R(x B , y B ), U(x C , y C ), D(x D , y D ) in sequence;

[0065] 42) Based on the coordinate transformation principle, transform the four points L(x A , y A ), R(x B , y B ), U(x C , y C ), D(x D , y D ) from the camera coordinate system to the world coordinate system to obtain their true coordinates;

[0066] 43) Use the Euclidean distance calculation formula to calculate the distances between the four coordinates respectively, obtain the three-dimensional information of the foreign object, and complete the measurement of the foreign object.

[0067] Further, in the step 24), it specifically includes the following steps:

[0068] a) In a grouped - reverse attention module, first separate the guiding prior feature and the candidate feature through grouped guidance, that is, split the candidate feature along the channel dimension into m i groups, where g i represents the size of each group after splitting;

[0069] b) Interpolate the reverse attention guidance periodically between the grouped features where i ∈ {1, 2, 3}, j ∈ {1, …, m }, k ∈ {3, 4, 5}; i}

[0070] c) Use the residual stage to generate the refined feature The formula is as follows:

[0071]

[0072] where represents pixel - level addition, Conv 3×3 represents a 3×3 convolution operation; ReLU(*) represents the ReLU activation function;

[0073] d) In Perform convolution on the basis of to obtain a single-channel residual guidance Its formula is as follows:

[0074]

[0075] e) Repeat steps a)-d) three times, that is, pass through the three-group - reverse attention module in sequence, and then combine the output of the third time and the target rough map C6, and then go through an upsampling to restore the size, and output the result C5. This process can be expressed by the general formula as:

[0076]

[0077] Furthermore, in the said step 34), it specifically includes the following steps:

[0078] a) Use the Softmax activation function to regress the initial disparity d0 from C G , and its formula is as follows, where D represents the total number of disparity points:

[0079]

[0080] b) Start iteration. At each iteration, use the current disparity d k to index from C G and the local feature correlation cost volume C A to generate a set of geometric features G f , and its formula is as follows:

[0081]

[0082] where d k represents the current disparity, r is the index radius, τ represents the pooling operation, and Concat represents concatenation.

[0083] c) Update the hidden state h k-1 , G f , c k , c r , ch and the current disparity d k by cascading through two encoder layers with d k to form x k , and then use ConvGRU to update the hidden state h k-1 , and its formula is as follows:

[0084] x k = [Encoder g (G f ), Encoder d ((d) k ), d k ,

[0085] z k = σ(Conv([h k-1 , x k , W z ) + c k ),

[0086] r k = σ(Conv[Conv([h k-1 , x k , W r )] + c r ),

[0087]

[0088] where c k , c r , c h are context features, W z , W r , W h are corresponding weights, σ represents the Sigmoid function, is the Hadamard product, tanh is the tangent function, Conv 3×3 represents the convolution operation;

[0089] d) Based on the final hidden state h k after iteration, the disparity residual Δd k is decoded through two convolutional layers, and then the current disparity is updated to obtain the final disparity d k+1 , and the formula is as follows:

[0090] d k+1 = d k + Δd k .

[0091] Compared with the prior art, the beneficial effects of the present invention are:

[0092] (1) The epipolar rectification method based on the singular value decomposition of the camera translation matrix adopted has certain advantages in rectification accuracy, image deformation, and running speed compared with the traditional method, meets the requirements of epipolar rectification, optimizes the stereo matching algorithm, solves the errors and cumbersome calculation processes caused by the mechanical deviation of the camera during the stereo matching process of binocular cameras, and has certain application value.

[0093] (2) A method for identifying abnormal areas of underground pipelines guided by multi-level features is proposed. By identifying foreign objects in the corrected image, it can initially frame the image area, reducing the computational complexity of subsequent stereo matching. At the same time, this method can automatically learn and extract important features. Compared with traditional methods, it has higher accuracy, enhanced feature extraction ability, can effectively handle various interference factors that may occur in practical applications, and has good application value.

[0094] (3) The IGEV network is used for binocular stereo matching, which combines existing cost filtering-based methods and iterative optimization-based methods. It can improve the recognition ability of pathological areas (such as occlusion, repetitive texture, low texture, high reflection, etc.) and the problem of excessive smoothing at boundaries and fine details. Brief Description of the Drawings

[0095] Figure 1 This is the overall flowchart of a method for locating abnormal areas of underground pipelines based on stereo vision perception according to the present invention.

[0096] Figure 2 This is the target recognition flowchart of a method for locating abnormal areas of underground pipelines based on stereo vision perception according to the present invention.

[0097] Figure 3 This is the group - reverse attention module diagram of a method for locating abnormal areas of underground pipelines based on stereo vision perception according to the present invention. Detailed Embodiments

[0098] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention with reference to the accompanying drawings.

[0099] Please refer to Figure 1 , the present invention provides a method for locating abnormal areas of underground pipelines based on stereo vision perception. The specific process is as follows:

[0100] 1) The image is processed for clarity, including binocular camera calibration and epipolar rectification. The specific steps are as follows:

[0101] 11) The checkerboard images at different positions are captured by the binocular camera, obtaining 20 left-eye images and 20 right-eye images respectively. Based on Zhang's calibration method, the internal parameter matrices M1, M2 and rotation matrices R1, R2 of the left and right cameras are obtained;

[0102] 12) The images of foreign objects in the pipeline are collected by the binocular camera. The left-eye image is C left , and the right-eye image is C right . Then, an epipolar rectification method based on the singular value decomposition of the camera translation matrix is used for epipolar rectification to obtain the rectified images C left ′, C right′, the specific steps of the epipolar rectification method based on the singular value decomposition of the camera translation matrix are as follows:

[0103] a) Calculate a new rotation matrix R n , and its formula is as follows. The rotation matrices of the corrected left and right cameras are R n , the image poles are at infinity and the epipolar lines are parallel;

[0104]

[0105] where T is the relative translation vector between the binocular cameras, and R x is the rotation matrix rotating by w angle around the X axis, and its size is U1 and U2 are the abscissas of the actual principal points of the left and right eyes, and Q is the correction matrix, and its size is where

[0106] b) Determine the new internal parameter matrix M of the corrected binocular cameras n ;

[0107] First, determine the tilt factor s. Assume that the origin in C left is O(0,0,1), the row corner point A(M,0,1), the column corner point B(0,N,1), and the center point After the transformation by H1, these 4 points become O1, A1, B1, and G1 respectively. A1O1 = (x1,y1,0), B1O1 = (x2,y2,0), then

[0108] Then adjust the focal length, move the corrected image to the center of the field of view, and let The coordinates of G1 are (p,q,l), L, H, X, and Y are intermediate parameters and have no actual physical meaning. Then the formula for the new internal parameter matrix M n is as follows:

[0109]

[0110] c) For each image point P1, P2 in the original pixel coordinate system, obtain the image points m1 = M1 -1 P1, m2 = M2 -1 P2 in the normalized coordinate representation through the internal parameter matrices M1 and M2, and then through the transformation matrices R L , R R Convert the image points m1, m2 in the normalized coordinate representation to the corrected camera coordinate system, multiply by the reciprocals α1, α2 of the Z vectors of the object points in the corrected camera coordinate system, and obtain the normalized coordinates m1' = α1R L m1, m2' = α2RR m2, where the transformation matrix R L and R R are the results obtained by the left and right eye cameras rotating around the Z-axis by angle, then rotating around the Y-axis by θ angle, and finally rotating around the X-axis by ω angle respectively;

[0111] d) Through the new intrinsic matrix M n the points m1′ and m2′ are transformed onto a new pixel coordinate system, obtaining the corrected points P1′ = M n m1′, P2′ = M n m2′, where the expressions of the transformation matrices H1 and H2 for the left and right image corrections are respectively:

[0112]

[0113] Finally, the corrected images C left ′ and C right ′ are obtained.

[0114] 2) Please refer to Figure 2 and send the corrected left and right eye images C left ′ and C right ′ into the multi-level feature-guided underground pipeline anomaly area recognition method to identify foreign objects in the corrected images, respectively obtaining the classification result ground truth maps GT left ″ and GT right ″. This method includes a search stage and a confirmation stage. The specific steps are as follows:

[0115] 21) In the search stage, first, feature extraction is performed, that is, based on the ResNet50 neural network, multi-level features f left are extracted from the corrected left eye image C m , m ∈ {1, 2, 3, 4, 5}, and the resolution of each feature is respectively m ∈ {1, 2, 3, 4, 5}, where H1 and W represent the sizes of three scales of height and width respectively; then the last three layers of features f3, f4, and f5 are respectively input into the texture enhancement module, and through a set of dilated convolutions with different dilation rates, the corresponding features are extracted by mimicking the different-sized grouped receptive fields in the human visual system, and f3′, f4′, and f5′ are output;

[0116] 22) Input f3′, f4′, and f5′ into the adjacent connection decoding module to obtain a rough map C6 of the target. The specific steps are as follows:

[0117] a) The following processing is performed on the three layers of features f3′, f44′, and f55′ input into the adjacent connection decoding module to obtain

[0118]

[0119] wherein is the weight of each layer, and \(u\in\{1,2,3\}\); represents a \(3\times3\) convolution and normalization operation, and Upsample 2 represents two upsamplings, represents pixel-level multiplication;

[0120] b) Combine to obtain a rough graph \(C6\) of the target;

[0121] 23) In the confirmation stage, first obtain the reverse attention guidance from the rough graph \(C6\) through the Sigmoid activation function and reverse operation The formula is as follows;

[0122]

[0123] wherein Upsample 2 represents two upsamplings, Downsample 4 represents four downsamplings, Reverse represents the negation operation, \(E\) is the identity matrix with all element values being 1, and \(\sigma\) is the Sigmoid activation function.

[0124] 24) Take and \(f5'\) as inputs, and successively pass through three grouped - reverse attention modules. Combine the reverse attention guidance and grouped guidance with the residual learning process. Please refer to Figure 3 , combine multiple grouped - reverse attention modules, gradually refine the rough prediction through feature layers of different depths, and then combine with the rough graph \(C6\) of the target to output the result \(C5\);

[0125] a) In a grouped - reverse attention module, first separate the guiding prior feature and candidate feature through the grouped guidance, that is, split the candidate feature along the channel dimension into \(m\) i groups, where \(g\) i represents the size of each group after splitting;

[0126] b) Interpolate the reverse attention guidance periodically among the grouped features where \(i\in\{1,2,3\}\), \(j\in\{1,\cdots,m i \}\), \(k\in\{3,4,5\}\);

[0127] c) Use the residual stage to generate refined features The formula is as follows:

[0128]

[0129] Among them denoted as pixel-level addition, Conv 3×3 represents a 3×3 convolution operation; ReLU(*) represents the ReLU activation function;

[0130] d) On the basis of perform convolution to obtain a single-channel residual guidance The formula is as follows:

[0131]

[0132] e) Repeat steps a)-d) three times, that is, pass through the grouping-inverse attention module three times in sequence, and then combine the output of the third time and the target rough map C6, and then go through an upsampling to restore the size, and output the result C5. This process can be expressed by the general formula as:

[0133]

[0134] 25) Repeat this step twice on the basis of the output of step 24), and output the results C4 and C3 in sequence. Finally, use the Sigmoid function to normalize C3, and output the recognition result GT of the multi-level feature-guided underground pipeline anomaly area recognition method left ; The right-eye image C right ′ performs the same operations as steps 21)-24), and outputs GT right ;

[0135] 26) Input GT left and GT right into the classification module, including a fully connected layer and a Softmax activation function layer, to obtain the foreign object classification result. The specific steps are as follows:

[0136] a) Flatten GT left and GT right into one-dimensional data respectively and input them into 3 fully connected layers, gradually adjusting the output size to be equal to the number of categories;

[0137] b) Input into the Softmax layer for classification to obtain the predicted category of each pixel, and output the final classification result maps GT left ′ and GT right ′:

[0138]

[0139] Among them, z i is the input of the Softmax layer, Q is the number of label types, and y i is the probability that the predicted pixel belongs to the Qth category. Finally, select the category with the highest probability as the final prediction result;

[0140] In the method for identifying abnormal areas of underground pipelines guided by multi-level features, the cross-entropy loss function is selected as the loss function. The calculation formulas for the cross-entropy function CE(h, s) and the loss function loss are as follows:

[0141]

[0142] where Q is the number of label categories, a is the number of input pixels, h is the true label distribution of the pixel, and s is the predicted label result of the pixel. represents the j'-th category of the i'-th sample, represents the probability that the predicted i'-th sample belongs to the j'-th category.

[0143] 3) Use the IGEV (Iterative Geometry Encoding Volume, IGEV) network to perform binocular stereo matching on the recognition and classification results GT left ′ and GT right ′, and output the full-resolution disparity map C, as Figure 3 . The specific steps are as follows:

[0144] 31) First, locate the estimated area of the foreign object according to the recognition result of the method for identifying abnormal areas of underground pipelines guided by multi-level features, that is, GT left ″ = C left ′ · Gt left ′ and GT right ″ = C right ′ · GT right ′;

[0145] 32) Input GT left ″ and GT right ″ into the feature extractor of the IGEV network, which includes a feature extraction network and a context network. The specific steps are as follows:

[0146] a) In the feature extraction network, downsample the input images GT left ″ and GT right ″ to 1 / 32, and then obtain multi-scale features through upsampling Select f l,4 and f r,4 to construct the cost volume, C r , H1, and W represent the sizes of the three scales of channels, height, and width respectively;

[0147] b) In the context network, through a series of residual blocks and downsampling layers, generate multi-scale context features at 1 / 4, 1 / 8, and 1 / 16 of the input image resolution of 128 channels, denoted as c k 、cr , c h ;

[0148] 33) Construct the grouping-related cost volume, and the steps are as follows:

[0149] a) Divide f l,4 and f r,4 into N g = 8 groups according to the channel dimension, and calculate the correlation mapping of each group to form a 4D correlation cost volume C corr , and its formula is as follows:

[0150]

[0151] where <·,·> is the vector inner product, d is the disparity index, N c is the number of channels, N g is the number of groups, and x, y represent the horizontal and vertical coordinates of the pixel;

[0152] b) Use a lightweight 3D regularization network to perform cost aggregation on C corr and output the geometric encoding volume C G ; at the same time, insert a guided cost volume excitation operation to update C e during each operation; the 3D regularization network includes 3 downsampling modules and 3 upsampling modules. The downsampling module includes two 3×3×3 3D convolutions with the number of channels being 16, 32, and 48 respectively. The upsampling module includes a 4×4×4 transposed convolution and two 3×3×3 3D convolutions. The formula for cost aggregation is as follows:

[0153] C G = R(C corr )

[0154] where C G represents the geometric encoding volume obtained after cost aggregation, and R(·) represents the 3D regularization network;

[0155] During the cost aggregation process, for a cost volume C e (e = 4, 8, 16, 32), the formula for the guided cost volume excitation is as follows:

[0156]

[0157] where D represents the total number of disparity points, C e ' represents the result obtained after the guided cost volume excitation of the cost volume C e , σ represents the Sigmoid function, is the Hadamard product;

[0158] c) Calculate the corresponding left and right image feature differences to obtain the local feature correlation cost volume CA , and then expand the receptive field, that is, use 1D average pooling with a size and stride of 2 to obtain two levels of C G pyramid and C A pyramid;

[0159] 34) Use a 3-level ConvGRU-based Update Operator (ConvGRU) for iterative update and output the final disparity d k+1 , and the specific steps are as follows:

[0160] a) Use the Softmax activation function to regress the initial disparity d0 from C G , and its formula is as follows, where D represents the total number of disparity points:

[0161]

[0162] b) Start the iteration. At each iteration, use the current disparity d k to index from C G and the local feature correlation cost volume C A to generate a set of geometric features G f , and its formula is as follows:

[0163]

[0164] where d k represents the current disparity, r is the indexing radius, τ represents the pooling operation, and Concat represents concatenation.

[0165] c) Update the hidden state h k-1 , G f , c k , c r , c h and the current disparity d k are cascaded with d k to form x k , and then use ConvGRU to update the hidden state h k-1 , and its formula is as follows:

[0166] x k = [Encoder g (G f ), Encoder d ((d) k ), d k ,

[0167] z k = σ(Conv([h k-1 , x k , Wz ) + c k ),

[0168] r k = σ(Conv[Conv([h k-1 , x k , W r )] + c r ),

[0169]

[0170] where c k , c r , c h are context features, W z , W r , W h are corresponding weights, σ represents the Sigmoid function, is the Hadamard product, tanh is the tangent function, Conv 3×3 represents the convolution operation;

[0171] d) Based on the final hidden state h k after iteration, the disparity residual Δd k is decoded through two convolutional layers, and then the current disparity is updated to obtain the final disparity d k+1 , and the formula is as follows:

[0172] d k+1 = d k + Δd k

[0173] 35) Convolve the final hidden state h k to generate features, then upsample them to 1 / 2 resolution, and concatenate them with f l,2 from the left image as the weight μ ∈ R H×W×9 , and finally output the full-resolution disparity map C through the weighted combination of the predicted disparity d k at 1 / 4 resolution.

[0174] 4) Obtain the three-dimensional coordinates based on the coordinate transformation principle to complete the foreign object ranging. The specific steps are as follows:

[0175] 41) Find the foreign object contour in the disparity map C and locate the coordinates of the four outermost points on the contour, denoted as L(x A , y A ), R(x B , y B ), U(x C , y C ), D(x D , y D );

[0176] 42) Based on the coordinate transformation principle, transform the four points L(x A , y A ), R(x B , y B ), U(x C , y C ), and D(x D , y D ) from the camera coordinate system to the world coordinate system to obtain their true coordinates;

[0177] 43) Use the Euclidean distance calculation formula to calculate the distances between the four coordinates respectively to obtain the three-dimensional information of the foreign object and complete the measurement of the foreign object.

[0178] As described above, it is only the specific implementation manner in the present invention, and the scope protected by the present invention is not limited thereto. Anyone familiar with the technology is within the technical scope disclosed by the present invention.

Claims

1. A method for locating abnormal areas of underground pipelines based on stereoscopic visual perception, characterized in that: The following steps are involved: Step (1) performing a sharpening correction on the image; Step (2) using a multi-level feature-guided underground pipeline abnormal area recognition method to identify foreign objects in the image corrected in step (1); Step (3) locates the foreign object according to the foreign object recognition result of step (2), and uses the IGEV network to perform binocular stereo matching to output a full-resolution disparity map C; Step (4) finds the foreign body contour in the full-resolution disparity map C output in step (3) and locates the coordinates of the four most edge points in the contour, obtains the three-dimensional coordinates based on the coordinate transformation principle, and completes the foreign body measurement.

2. The method for locating abnormal areas of underground pipelines based on stereoscopic visual perception according to claim 1 is characterized in that: The step (1) of performing image sharpening correction specifically comprises the following steps: 11) Use a binocular camera to shoot chessboard images at different positions, obtain 20 left and right eye images, and based on the Zhang Zhengyou calibration method, obtain the intrinsic parameter matrices M1, M2 and rotation matrices R1, R2 of the left and right cameras; 12) Use binocular cameras to collect images of foreign objects in the pipeline, where the left image is C left , the right image is C right , and then use an epipolar correction method based on the singular value decomposition of the camera translation matrix to perform epipolar correction to obtain the corrected image C left ′、C right ', the epipolar correction method based on singular value decomposition of camera translation matrix has the following specific steps: a) Calculate a new rotation matrix R n , the formula is as follows, the corrected left and right camera rotation matrix is ​​M n , the image poles are at infinity and the epipolar lines are parallel; Where T is the relative translation vector between the binocular cameras, R x is the rotation matrix of angle w around the X axis, and its size is U1 and U2 are the horizontal coordinates of the actual principal points of the left and right eyes, and Q is the correction matrix, whose size is in b) Determine the new internal parameter matrix M of the corrected binocular camera n ; First determine the tilt factor S, assuming C left The origin is O(0,0,1), the row corner is A(M,0,1), the column corner is B(0,N,1), and the center is After H1 transformation, these four points are O1, A1, B1, G1, A1O1=(x1, y1, 0), B1O1=(x2, y2, 0), then Then adjust the focal length and move the corrected image to the center of the field of view. The coordinates of G1 are (p,q,l), L, H, X, and Y are intermediate parameters and have no actual physical meaning. The new internal parameter matrix M n The calculation formula is as follows: c) For each pixel P1, P2 in the original pixel coordinate system, the normalized coordinate representation of the image point m1=M1 is obtained through the internal parameter matrices M1, M2 -1 P1, m2 = M2 -1 P2, and then through the transformation matrix R L , R R The image points m1 and m2 represented by the normalized coordinates are converted to the calibrated camera coordinate system, and multiplied by the inverse of the Z vector α1 and α2 of the object point in the calibrated camera coordinate system to obtain the normalized coordinates m1′=α1R in the calibrated camera coordinate system. L m1、m2′=α2R R m2, where the transformation matrix R L , R R Rotate the left and right cameras around the Z axis respectively Angle, then rotate around the Y axis by angle θ, and finally rotate around the X axis by angle w. ; d) Through the new internal parameter matrix M n Convert points m1′ and m2′ to the new pixel coordinate system to obtain the correction point P1′=M n m1′、P2′=M n m2′, where the transformation matrices H1 and H2 for left and right image correction are expressed as: Finally, the corrected image C is obtained left ′、C right ′.

3. The method for locating abnormal areas of underground pipelines based on stereoscopic visual perception according to claim 1 is characterized in that: Step (2) specifically includes the following steps: 21) In the search phase, feature extraction is first performed, that is, based on the ResNet50 neural network, from the corrected left image C left ′ to extract multi-level features f m , m∈{1,2,3,4,5}, the resolution of each feature is m∈{1,2,3,4,5}, H1 and W represent the height and width of the three scales respectively; then the last three layers of features f3, f4, and f5 are input into the texture enhancement module respectively, and a set of dilated convolutions with different dilation rates are used to imitate the different sizes of grouped receptive fields in the human visual system to extract the corresponding features, and output f3′, f4′, and f5′; 22) Input f3′, f4′, and f5′ into the neighbor connection decoding module to obtain the target rough graph C6. The specific steps are as follows: a) The three-layer feature input of f3′, f4′, and f5′ is processed as follows in the neighbor connection decoding module to obtain in is the weight of each layer, u∈{1,2,3}; Represents a 3×3 convolution and normalization operation, Upsample 2 represents two upsamplings, represents pixel-level multiplication; b) Combination Get a rough image C6 of the target; 23) In the confirmation stage, the reverse attention guidance is first obtained from the rough image C6 through the Sigmoid activation function and the reverse operation The formula is as follows; Upsample 2 Indicates two upsampling, Downsample 4 Represents four times of downsampling, Reverse represents the inversion operation, E is the unit matrix whose element values ​​are all 1, and σ is the Sigmoid activation function. 24) and f5′ as input, and pass through three group-reverse attention modules in sequence, combining reverse attention guidance and group guidance with the residual learning process, combining multiple group-reverse attention modules, gradually refining the rough prediction through feature layers of different depths, and then combining the rough map C6 of the target to output the result C5; 25) Repeat this step twice based on the output of step 24), output the results C4 and C3 in turn, and finally normalize C3 using the Sigmoid function to output the recognition result GT of the underground pipeline abnormal area recognition method guided by multi-level features left ; Right eye image C right 'Same as steps 21)-24), output GT right ; 26) GT left and GT right Input the classification module, including the fully connected layer and the Softmax activation function layer, to obtain the foreign body classification results. The specific steps are as follows: a) GT left and GT right Flatten them into one-dimensional data and input them into three fully connected layers, gradually adjusting the output size to be equal to the number of categories; b) Input the Softmax layer for classification, obtain the predicted category of each pixel, and output the final classification result map GT left ′ and GT right ′: Among them, z i is the input of the Softmax layer, Q is the number of label types, y i It is the probability of predicting that the pixel belongs to the Qth category, and finally the category with the highest probability is selected as the final prediction result; The loss function in the underground pipeline abnormal area recognition method guided by multi-level features selects the cross entropy loss function. The cross entropy function CE(h,s) and the loss function loss calculation formula are as follows: Among them, Q is the number of label types, a is the number of input pixels, h is the true label distribution of the pixel, and s is the predicted label result of the pixel. represents the j′th category of the i′th sample, It represents the predicted probability that the i′th sample belongs to the j′th category.

4. The method for locating abnormal areas of underground pipelines based on stereoscopic visual perception according to claim 1, characterized in that: Step (3) specifically includes the following steps: 31) First, the estimated area of ​​the foreign body is located according to the recognition results of the underground pipeline abnormal area recognition method guided by multi-level features, that is, GT left =C left 'GT left ′ and GT right =C right 'GT right ′; 32) GT left ″ and GT right ″Input the feature extractor of the IGEV network, which contains the feature extraction network and the context network. The specific steps are as follows: a) In the feature extraction network, the input graph GT left ″ and GT right ″Downsample to 1 / 32, then upsample to get multi-scale features Select f l,4 and f r,4 Construct the cost body, C e , H1, and W represent the sizes of the three scales of channel, height, and width respectively; b) In the context network, through a series of residual blocks and downsampling layers, multi-scale context features are generated at 1 / 4, 1 / 8 and 1 / 16 of the input 128-channel image resolution, denoted by c k 、c r 、c h ; 33) Construct the group-related cost body, the steps are as follows: a) f l,4 and f r,4 Divide into N according to the channel dimension g = 8 groups, and calculate the correlation mapping of each group to form a 4-dimensional correlation cost body C corr , the formula is as follows: where <·,·> is the vector inner product, d is the disparity index, and N c is the number of channels, N g is the grouping number, x and y represent the horizontal and vertical coordinates of the pixel; b) Use a lightweight 3D regularization network for C corr Perform cost aggregation and output the geometric encoding body C G ; At the same time, insert the guided cost body incentive operation and update C at each operation e ; The 3D regularization network contains 3 downsampling modules and 3 upsampling modules. The downsampling module contains two 3×3×3 3D convolutions with 16, 32, and 48 channels respectively. The upsampling module contains a 4×4×4 transposed convolution and two 3×3×3 3D convolutions. The formula for cost aggregation is as follows: C F =R(C corr ) Among them, C G represents the geometric encoding body obtained after cost aggregation, and R(·) represents the 3D regularized network; In the cost aggregation process, for a The cost body C e (e=4,8,16,32), the guiding cost body incentive formula is as follows: Where D represents the total number of disparity points, C e ′ represents the cost body C e The result obtained after guiding the cost body excitation, σ represents the Sigmoid function, and is the Hadamard product; c) Calculate the corresponding left and right image feature difference to obtain the local feature association cost volume C A , and then expand the receptive field, that is, use 1D average pooling with a size and stride of 2 to obtain two levels of C G Pyramid and C A pyramid; 34) Use the 3-level ConvGRU iterative updater (ConvGRU-based Update Operator, ConvGRU) to iteratively update and output the final disparity d k+1 , the specific steps are as follows: 35) For the hidden state h k Convolution is performed to generate features, which are then upsampled to 1 / 2 resolution and combined with f from the left image. l,2 Cascade as weight μ∈R H×W×9 , and finally the predicted disparity d at 1 / 4 resolution k The weighted combination of outputs the full-resolution disparity map C.

5. The method for locating abnormal areas of underground pipelines based on stereoscopic visual perception according to claim 1, characterized in that: The step (4) specifically comprises the following steps: 41) Find the contour of the foreign object in the disparity map C and locate the coordinates of the four points on the edge of the contour, which are recorded as L(x A ,y A )、R(x B ,y B )、U(x C ,y C )、D(x D ,y D ); 42) Based on the principle of coordinate transformation, L(x A ,y A )、R(x B ,y B )、U(x C ,y C )、D(x D ,y D ) The four points are transformed from the camera coordinate system to the world coordinate system to obtain their real coordinates; 43) Use the Euclidean distance calculation formula to calculate the distances between the four coordinates respectively, obtain the three-dimensional information of the foreign body, and complete the foreign body measurement.

6. The method for locating abnormal areas of underground pipelines based on stereoscopic visual perception according to claim 3 is characterized in that: In the step 24), the following steps are specifically included: a) In a group-reverse attention module, the guided prior features and candidate features are first separated by group guidance. Split into m along the channel dimension i Group, where g i Represents the size of each group after splitting; b) In grouping features Periodically interpolate reverse attention guidance where i∈{1,2,3}, j∈{1,…,m i }, k∈{3,4,5}; c) Use the residual stage to generate refined features The formula is as follows: in Represented as pixel-level addition, Conv 3×3 Represents a 3×3 convolution operation; ReLU(*) represents the ReLU activation function; d) Convolution is performed on the basis of single channel residual guidance The formula is as follows: e) Repeat steps a)-d) three times, i.e., pass through the grouping-reverse attention module three times in sequence, and then combine the output of the third time And the target rough image C6, after an upsampling to restore the size, the output result C5, the process can be expressed as:

7. The method for locating abnormal areas of underground pipelines based on stereoscopic visual perception according to claim 4 is characterized in that: In the step 34), the following steps are specifically included: a) Use Softmax activation function from C G The initial disparity d0 is regressed in the following formula, where D represents the total number of disparity points: b) Start iteration, and at each iteration, use the current disparity d k By linear interpolation from C G Associated cost volume C with local features A Indexing is performed to generate a set of geometric features G f , the formula is as follows: where d k represents the current disparity, r is the index radius, τ represents the pooling operation, and Concat represents the connection. c) Update the hidden state h k-1 , G f 、c k 、c r 、c h and the current disparity d k Through two encoder layers with d k Cascade formation x k , and then use ConvGRU to update the hidden state h k-1 , the formula is as follows: x k =[Encoder g (G f ),Encoder d ((d) k ),d k ], z k =σ(Conv([h k-1 ,x k ],W z )+c k ), r k =σ(Conv[Conv([h k-1 ,x k ],W r )]+c r ), where c k 、c r 、c h is the context feature, W z , W r , W h is the corresponding weight, σ represents the Sigmoid function, is the Hadamard product, tanh is the tangent function, Conv 3×3 Represents the convolution operation; d) Based on the final hidden state h after iteration k , decoded through two convolutional layers to obtain the disparity residual Δd k , then update the current disparity to get the final disparity d k+1 , the formula is as follows: d k+1 =d k +Δd k 。

Citation Information

Patent Citations

  • Pipeline defect detecting, positioning and ranging system based on binocular stereoscopic vision

    CN115272271A