A pipeline defect detection, positioning and ranging system based on binocular stereo vision
Through binocular stereoscopic vision and deep learning technology, three-dimensional reconstruction and object detection of pipeline defects are achieved, and the problem of low efficiency and accuracy of pipeline defect detection in traditional methods is solved, and high-precision and real-time pipeline defect detection and positioning is achieved.
Patent Information
- Application Number
- CN202210947571.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-08-09
AI Technical Summary
Traditional manual pipeline inspection is low efficiency and accuracy, and existing robots rely on monocular cameras to accurately detect, measure and locate pipeline defects.
The pipeline defect detection and positioning distance measurement system based on binocular stereo vision is adopted, and the images are captured by binocular cameras are used for three-dimensional reconstruction, combined with deep learning technology to achieve target detection, and the three-dimensional coordinates of pipeline defects are obtained using stereo matching and depth calculation modules.
It realizes fully automatic and contactless measurement, with accurate positioning and high real-time performance, improving the efficiency and accuracy of defect identification.
Smart Images

Figure CN115272271B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of object detection and multi - camera vision positioning in computer vision, and particularly relates to a pipeline defect detection, positioning and ranging system based on binocular stereo vision. Background Art
[0002] Pipeline transportation is an essential part of the urbanization process. To ensure the safe and smooth operation of pipeline transportation, it is necessary to regularly detect pipeline defects. However, traditional manual pipeline inspection mainly relies on human vision. In the case of long - term work, the efficiency and accuracy of defect recognition will decline, and there are also potential safety hazards. Most of the existing pipeline detection robots only rely on a single - camera to collect images, and cannot accurately detect, measure and locate the defects existing in the pipeline.
[0003] With the continuous improvement of computer image - processing capabilities, computer vision has been widely applied, and the combination of computer vision and robot systems has become an important means to improve the intelligence of robots. In current actual industrial applications, two - dimensional vision technology is usually combined with robots. However, two - dimensional images can hardly obtain the depth information of objects and it is difficult to obtain the three - dimensional information of targets. Therefore, it is necessary to reconstruct the three - dimensional information of the target from two - dimensional images in order to more comprehensively and realistically reflect objective objects. Binocular stereo vision is an important branch in the field of computer vision. It captures a scene simultaneously using two cameras separated by a short distance, simulating the human eye to obtain images of the same scene from two different viewing angles, and restores the three - dimensional information of each point in the original image by analyzing the disparity between the two images. Due to the advantages of high efficiency, high precision, non - contact measurement, etc., binocular vision can be widely applied to target recognition and positioning.
[0004] The object - detection technology based on deep learning is an important branch in the field of computer vision. Compared with traditional object - detection methods, it not only improves the recognition speed but also the recognition accuracy. Deep learning conducts self - training and learning through a convolutional neural network, can automate the feature - extraction process, complete the construction and training of the deep - learning network, and obtain a weight file, thus realizing the detection and recognition of targets, with the advantages of stability, high precision, high efficiency, automatic detection, etc.
[0005] Aiming at the problems faced by traditional pipeline defect - detection methods, how to accurately detect, measure and locate the defects existing in the pipeline using a binocular vision system is an issue that needs to be solved currently. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a pipeline defect detection, positioning and ranging system based on binocular stereo vision. Through binocular stereo vision technology, three-dimensional reconstruction of the images captured by the binocular cameras and positioning and ranging of the detection targets are realized. Combining deep learning technology to achieve target detection of pipeline defects, making the system have the characteristics of full automation, non-contact measurement, accurate positioning and high real-time performance, and improving the efficiency and accuracy of defect recognition.
[0007] The present invention provides the following technical solutions:
[0008] A pipeline defect detection, positioning and ranging system based on binocular stereo vision, comprising a binocular camera image capture module, a camera calibration module, a stereo rectification module, a deep learning target detection module, and a stereo matching and depth calculation module;
[0009] The binocular camera image capture module is used to capture and collect the image data of the left and right cameras, realize the real-time capture of pipeline images, and at the same time collect pipeline defect image data, providing a data source for the pipeline defect data set;
[0010] The camera calibration module is used to correct the images captured by the binocular cameras to obtain images with relatively small distortion. The input is the image coordinates and world coordinates of the known feature points of the calibration board, establish a geometric model of camera imaging, determine the relationship between the camera image pixel coordinates and the three-dimensional coordinates of the scene points, and the output is the internal and external parameters and distortion coefficients of the camera;
[0011] The stereo rectification module adopts a vision system with intersecting optical axes. By decomposing the rotation matrix and row-aligned rectification rotation matrix, the relative positions of the two cameras are changed so that the corresponding points in the two images are on the same horizontal epipolar line. Based on the original data of the cameras, mathematical and physical methods are used to obtain the corrected camera parameters, change the two-dimensional search to one-dimensional search, reduce the matching search space, improve the search rate of stereo matching, and finally complete image rectification through a rectification mapping lookup table, crop and save the images;
[0012] The deep learning target detection module, through the pipeline defect data collected by the binocular camera image capture module, builds its own pipeline defect data set, and trains a deep learning network based on the target detection algorithm. The network infers and predicts the images. The input is the RGB three-channel image captured by the left camera after stereo rectification, and the output is the defect semantic label class, the center coordinates (x, y) of the recognition box, and the width and height (w, h) recognized in the left camera image. The semantic label, the center coordinates of the recognition box, and the width and height data are used as object recognition information;
[0013] The stereo matching and depth calculation module takes the left and right camera images of the binocular camera after stereo rectification and the pipeline defect detection information obtained by deep learning object recognition as inputs. After being processed by the stereo matching algorithm, it obtains the disparity map of the left camera of the binocular camera after stereo rectification, and then converts it into a depth map through depth calculation. Finally, combined with the pipeline defect detection information, it outputs the spatial three-dimensional coordinates of the recognized object in the left camera image of the binocular camera after stereo rectification mapped to the actual three-dimensional space.
[0014] Preferably, the binocular camera image capture module captures and collects the image data of the left and right cameras; the camera calibration module establishes a camera imaging geometric model and corrects lens distortion to obtain the internal and external parameters and distortion coefficients of the camera; the stereo rectification module realizes the coplanar row alignment of the left and right images, making the left and right image planes parallel to the baseline, and the corresponding points in the left and right images on the same horizontal epipolar line; the deep learning object detection module trains a deep learning network based on the object detection algorithm to realize the detection and recognition of pipeline defects and obtain object recognition information; the stereo matching and depth calculation module obtains the spatial three-dimensional coordinates of pipeline defects to realize the precise positioning and ranging of pipeline defects.
[0015] Preferably, between the binocular camera image capture module, the camera calibration module, the stereo rectification module, the deep learning object detection module, and the stereo matching and depth calculation module, the following steps are used to realize the object detection and positioning and ranging of pipeline defects:
[0016] Step 1: Calibrate the binocular camera, establish a geometric model of camera imaging, determine the mutual relationship between the three-dimensional geometric position of a certain point on the surface of a spatial object and its corresponding point in the image, and solve the internal parameters, external parameters, and distortion coefficients of the left and right cameras of the binocular camera from the image coordinates and world coordinates of the known feature points of the calibration board. The internal parameters, external parameters, and distortion coefficients of the left and right cameras of the binocular camera are used as camera calibration parameters. By adjusting the internal parameters, external parameters, and distortion coefficients, the images captured by the binocular camera can be corrected to obtain images with relatively small distortion.
[0017] Preferably, the specific implementation steps of Step 1 are as follows:
[0018] Step 1.1: Make a checkerboard calibration board composed of alternating black and white squares, and use the binocular camera to take multiple shots of the checkerboard calibration board at multiple positions, angles, and postures, so that the single-plane checkerboard is clearly imaged in both the left and right cameras. For each calibration picture, extract its corner information to obtain the image coordinates of all interior corner points on the calibration image and the three-dimensional spatial coordinates of all interior corner points on the calibration board image.
[0019] Step 1.2: Establish a geometric model of camera imaging, determine the mutual relationship between the three-dimensional geometric position of a certain point on the surface of a spatial object and its corresponding point in the image. These geometric model parameters are the camera calibration parameters, including internal and external parameters and distortion coefficients;
[0020] The external parameter matrix W reflects the transformation between the camera coordinate system and the world coordinate system. Among them, R is the rotation matrix of the right camera of the binocular camera relative to the left camera, t is the translation vector of the right camera of the binocular camera relative to the left camera, and the internal parameter matrix M reflects the transformation between the pixel coordinate system and the camera coordinate system. Among them, f is the lens focal length, (u0, v0) is the coordinate of the origin of the image coordinate system in the pixel coordinate system, and d x 、f y are the sizes of each pixel point in the x-axis and y-axis directions of the image coordinate system:
[0021]
[0022]
[0023] Step 1.3: Take the image coordinates of all inner corner points on the calibration image obtained in Step 1.1 and the three-dimensional spatial coordinates of all inner corner points on the calibration plate image as inputs. According to the geometric model of camera imaging, solve through experiments and calculations and output the internal parameters, external parameters, and distortion coefficients of the left and right cameras of the binocular camera;
[0024] Step 1.4: Take the internal and external parameters of the binocular camera calibration in Step 1.3 as known constants, and use the coordinate information obtained through Step 1.1 and the coordinate relationship before and after correction to solve the five distortion parameters k1, k2, k3, p1, and p2 for distortion correction:
[0025]
[0026] where (x p , y p ) is the original coordinate of the image, and (x tcorr , y tcorr ) is the corrected coordinate of the image, which is approximately described by the Taylor series expansion at r = 0.
[0027] Step Two: Perform stereo calibration through epipolar constraint to make the corresponding points in the two images on the same horizontal epipolar line, obtain the calibrated camera parameters, call OpenCV to obtain the parameters of the calibrated left and right cameras in real time to complete the calibration, and finally obtain the calibrated image through the calibration mapping;
[0028] Preferably, the specific implementation steps of Step Two are as follows:
[0029] Step 2.1: Divide the binocular camera rotation matrix R into two parts, the combined rotation matrices r1 and r2 of the left and right cameras. Each of the left and right cameras rotates by half to make the optical axes of the left and right cameras parallel, so as to make the imaging planes of the left and right cameras coplanar.
[0030] Step 2.2: Input the combined rotation matrices r1 and r2 of the left and right cameras, the original internal parameter matrices of the left and right cameras, the translation vector t, and the size of the checkerboard image, and call the cvStereoRectify function in OpenCV to output the row-aligned rectification rotation matrices R1 and R2 of the left and right cameras, the rectified internal parameter matrices M l and M r of the rectified left and right cameras, the projection matrices P l and P r of the rectified left and right cameras, and the reprojection matrix Q;
[0031] Step 2.3: Take the output matrices in Step 2.2 as known constants, find the mapping tables through the rectification of the left and right views, use inverse mapping to find the floating-point positions on the source image corresponding to each integer pixel position on the target image, and interpolate each integer value of the surrounding source pixels. After the rectified images are all assigned values, crop the images and save the rectification results.
[0032] Step Three: Pre-collect a large number of pipeline defect images through the binocular camera, build a self-built pipeline defect dataset, screen and enhance the dataset to optimize the dataset, perform image annotation on the pipeline defect images based on the self-built dataset. After the annotation is completed, start training the deep convolutional neural network model. After the training is completed, infer the obtained weights, analyze the detection results, and obtain a weight file available for pipeline defect detection. During actual use, the image capture module of the binocular camera can capture images in real time, and perform pipeline defect target detection based on the weight file to obtain pipeline defect detection information;
[0033] Preferably, the specific implementation steps of Step Three are as follows:
[0034] Step 3.1: Data collection, pre-collect a large number of images through the binocular camera, take images containing pipeline defects, and build a self-built pipeline defect dataset;
[0035] Step 3.2: Data screening and enhancement, preliminarily screen the pipeline defect dataset, eliminate invalid data, and perform data enhancement operations on the preliminarily screened images to optimize the pipeline defect dataset for improving the training effect;
[0036] Step 3.3: Image annotation, annotate possible targets, generate annotation files, organize the training directory, and construct the training set, validation set, and test set required for training;
[0037] Step 3.4: Image training. Based on the pre-trained model, perform model iteration, detect the training results through Jupter, and adjust the parameters to prevent overfitting;
[0038] Step 3.5: Image inference. Infer the weights obtained after image training. The actual captured images should be used to analyze the detection results. If the detection effect is good, the weight file obtained from training can be used for object detection;
[0039] Step 3.6: Object detection. Through the images captured in real time by the binocular camera, a convolutional neural network that aggregates and forms image features at different image fine-grained levels, mixes and combines the image features, and passes the image features to the prediction layer to predict the image features, generate bounding boxes and predict categories, and finally obtain the pipeline defect detection information.
[0040] Preferably, in the neural network of the object detection algorithm, it specifically includes:
[0041] I. Data augmentation. After shrinking the images, randomly paste them onto the COCO 2017 dataset to increase the number of the dataset. At the same time, further generalize the augmented dataset by means of random scaling, random cropping, and random permutation;
[0042] II. Focus interlaced sampling and splicing. Uniformly scale the images to the size of (3, 640, 640) as the input, copy four copies, and cut these four pictures into four slices of (3, 320, 320) through slicing operations. Next, use Concat to connect these four slices from the depth, and the output is (12, 20, 320). Then, through a convolutional layer with 32 convolutional kernels, generate an output of (32, 320, 320). Finally, input the result into the next convolutional layer through batch_borm and leaky_relu;
[0043] III. Backbone. A convolutional neural network that aggregates and forms image features at different image fine-grained levels. Among them, Bottlenneck is a classic residual structure, first a 1×1 convolutional layer, then a 3×3 convolutional layer, and finally add the result to the initial input through the residual structure. CSP divides the original input into two branches, and performs convolutional operations on each branch to halve the number of channels. Branch 1 performs Bottlenneck×N operations, and then Concat branch 1 and branch 2, so that the input and output of BottlenneckCSP are of the same size, aiming to enable the model to learn more features;
[0044] 4. Neck: A network layer that mixes and combines image features and passes them to the prediction layer. The most important one is the SPP structure. The input of SPP is 512×20×20, and after a 1×1 convolution layer, the output is 256×20×20, and then it is sampled by three parallel MaxPools, and the result is added to its initial features to output 1024×20×20. Finally, a 512 convolution kernel is used to restore it to 512×20×20;
[0045] 5. Head: Predict image features, generate bounding boxes and predict categories.
[0046] Step 4: The left and right camera images of the binocular camera after stereo correction and the pipeline defect detection information obtained through deep learning object recognition are transmitted as input to the stereo matching and depth calculation module. The center coordinates of the identification box in the pipeline defect detection information are used as reference. After the stereo matching and depth calculation module processes the image, the output is obtained, and the spatial three-dimensional coordinates of the identified pipeline defect in the left camera image of the binocular camera after stereo correction are mapped to the actual three-dimensional space.
[0047] Preferably, the specific implementation steps of step 4 are as follows:
[0048] Step 4.1: Take the left and right camera images of the stereo camera after stereo correction as input, and calculate the matching cost within the preset parallax range. The purpose of the matching cost calculation is to measure the correlation between the pixels to be matched and the candidate pixels. Whether the two pixels are homonymous points or not, the matching cost can be calculated by the matching cost function. The smaller the cost, the greater the correlation and the greater the probability of being homonymous points. The BT algorithm is used as the calculation method for the matching cost. The calculation formula is:
[0049]
[0050] The formula can be used to calculate the matching cost of the left and right camera images of the stereo camera after stereo correction within the preset disparity range, and obtain the matching cost of each pixel point of the original image within the preset disparity range;
[0051] Step 4.2: Take the matching cost of each pixel calculated within the preset disparity range as input, perform cost aggregation, and adopt the idea of global stereo matching algorithm, that is, global energy optimization strategy. In simple terms, it is to find the optimal disparity of each pixel so that the global energy function of the whole image is minimized. The definition of global energy function is as follows:
[0052] E(d)=E data (d)+E smooth (d);
[0053] Adopt the method of path cost aggregation, that is, aggregate the matching costs at all disparities of a pixel along all one-dimensional paths around the pixel to obtain the path cost value under the path, and then add up all the path cost values to obtain the aggregated matching cost value of the pixel. The path cost calculation method for pixel p along a certain path r is as follows:
[0054]
[0055] Among them, p represents the pixel, r represents the path, d represents the disparity, p-r represents the pixels within the path of pixel p, L represents the aggregation cost value of a certain path, Lr(p-r,d) represents the cost value when the disparity of the previous pixel within the path is d, Lr(p-r,d-1) represents the cost value when the disparity of the previous pixel within the path is d-1, Lr(p-r,d+1) represents the cost value when the disparity of the previous pixel within the path is d+1, and min i L r (p-r, i) represents the minimum value of all cost values of the previous pixel within the path;
[0056] The first term is the matching cost C, which belongs to the data item;
[0057] The second term is the smooth term. The value accumulated to the path cost takes the minimum value among the three cases of no penalty, P1 penalty, and P2 penalty; P1 is to adapt to inclined or curved surfaces, and P2 is to preserve discontinuities. P2 is often dynamically adjusted according to the gray level difference of adjacent pixels, as shown in the following formula:
[0058]
[0059] P2′ is the initial value of P2, which is generally set to a number much larger than P1, I bq and and I bq represent the gray level values of pixels p and q respectively;
[0060] The third term is to ensure that the new path cost value Lr does not exceed a certain numerical upper limit,
[0061] The total path cost value S can be calculated by the following formula:
[0062]
[0063] Through the above calculations, the cost aggregation of multiple paths can be realized, and the multi-path cost aggregation values of each pixel within the preset disparity range can be obtained;
[0064] Step 4.3: Using the multi-path cost aggregation values of each pixel within the preset disparity range as input, perform disparity calculation. Disparity calculation determines the optimal disparity value for each pixel through the cost matrix S after cost aggregation. The Winner-take-all algorithm is adopted, that is, among the cost values for all disparities of a certain pixel, select the disparity corresponding to the minimum cost value as the optimal disparity. Finally, obtain the disparities of each pixel after cost aggregation;
[0065] Step 4.4: Using the disparities of each pixel after cost aggregation as input, perform disparity optimization. The purpose of disparity optimization is to further optimize the disparity map obtained in the previous step and improve the quality of the disparity map, including eliminating incorrect matches, improving disparity accuracy, and suppressing noise;
[0066] To eliminate incorrect matches, the left-right consistency check method is used. It is based on the uniqueness constraint of disparity, that is, each pixel has at most one correct disparity. The specific steps are to swap the positions of the left and right images, that is, the left image becomes the right image, and the right image becomes the left image, and then perform stereo matching again to obtain another disparity map. Since each value in the disparity map reflects the corresponding relationship between two pixels, according to the uniqueness constraint of disparity, through the disparity map of the left image, find the corresponding pixel and its corresponding disparity value in the right image for each pixel. If the absolute value of the difference between these two disparity values is less than 1, it satisfies the uniqueness constraint and is retained, otherwise it does not satisfy the uniqueness constraint and is eliminated. At the same time, the method of connected component detection is used to eliminate isolated outliers, remove small clusters in the disparity map caused by incorrect matches, and filter small isolated speckles. The formula for consistency check is as follows:
[0067]
[0068] To improve disparity accuracy, sub-pixel optimization technology is adopted. The method of quadratic curve interpolation is used to obtain sub-pixel accuracy. Perform quadratic curve fitting on the cost value of the optimal disparity and the cost values of the two adjacent disparities. The disparity value corresponding to the extreme point of the curve is the new sub-pixel disparity value; To suppress noise, median filtering is used to make the disparity result smoother, eliminate noise in the disparity map to a certain extent, and at the same time play a role in disparity filling. Through the above steps, disparity optimization can be achieved, and finally obtain the left camera disparity map of the stereo-calibrated binocular camera;
[0069] Step 4.5: Using the left camera disparity map of the stereo-calibrated binocular camera as input, depth calculation can be performed. The pixel depth calculation formula is as follows:
[0070]
[0071] where f is the focal length, b is the baseline length, d is the disparity, c xr and c xlis the column coordinates of the principal points of the two cameras. After depth calculation, a depth map of the left camera image of the binocular camera after stereo rectification can be obtained. Combining the pipeline defect detection information obtained by deep learning object detection, the spatial three-dimensional coordinates of the identified pipeline defects in the left camera image of the binocular camera after stereo rectification mapped to the actual three-dimensional space can be finally obtained.
[0072] Compared with the prior art, the present invention has the following beneficial effects:
[0073] (1) The pipeline defect detection, positioning and ranging system based on binocular stereo vision of the present invention realizes the three-dimensional reconstruction of the images captured by the binocular camera and the positioning and ranging of the detection target through binocular stereo vision technology, and combines deep learning technology to realize the object detection of pipeline defects, with the characteristics of full automation, non-contact measurement, accurate positioning and high real-time performance.
[0074] (2) The pipeline defect detection, positioning and ranging system based on binocular stereo vision of the present invention can realize non-contact measurement of pipeline defects, and has the characteristics of wide monitoring range, good real-time performance, high accuracy and accurate positioning.
[0075] (3) The pipeline defect detection, positioning and ranging system based on binocular stereo vision of the present invention trains a deep learning network based on object detection algorithms through a deep learning object detection module to realize the detection and recognition of pipeline defects and obtain object recognition information.
[0076] (4) The pipeline defect detection, positioning and ranging system based on binocular stereo vision of the present invention obtains the spatial three-dimensional coordinates of pipeline defects through a stereo matching and depth calculation module, and realizes accurate positioning and ranging of pipeline defects. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0078] Figure 1 is the overall framework diagram of the present invention.
[0079] Figure 2 is the working flow chart of the present invention.
[0080] Figure 3 is the working flow chart of the deep learning object detection module of the present invention.
[0081] Figure 4It is the flowchart of the stereo matching and depth calculation module of the present invention. Specific Embodiments
[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0083] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0084] Example 1:
[0085] As Figure 1 shown, a pipeline defect detection, positioning, and ranging system based on binocular stereo vision includes a binocular camera image capture module, a camera calibration module, a stereo rectification module, a deep learning object detection module, and a stereo matching and depth calculation module;
[0086] The binocular camera image capture module is used to capture and collect the image data of the left and right cameras, realize the real-time capture of the pipeline image, and also collect the pipeline defect image data to provide a data source for the pipeline defect data set;
[0087] The camera calibration module is used to correct the images captured by the binocular camera to obtain images with relatively small distortion. The input is the image coordinates and world coordinates of the known feature points of the calibration board, establish the geometric model of camera imaging, determine the relationship between the camera image pixel coordinates and the three-dimensional coordinates of the scene points, and the output is the internal and external parameters and distortion coefficients of the camera;
[0088] The stereo rectification module adopts a vision system with intersecting optical axes. By decomposing the rotation matrix and row-aligned rectification rotation matrix, the relative positions of the two cameras are changed so that the corresponding points in the two images are on the same horizontal epipolar line. Based on the original data of the cameras, mathematical and physical methods are used to obtain the corrected camera parameters, change the two-dimensional search to one-dimensional search, reduce the matching search space, and improve the search rate of stereo matching. Finally, the image rectification is completed through the rectification mapping lookup table, and the image is cropped and saved;
[0089] The deep learning object detection module builds its own pipeline defect dataset through the pipeline defect data collected by the binocular camera image capture module, and trains a deep learning network based on the object detection algorithm. This network infers and predicts the image. The input is the RGB three-channel image captured by the left camera after stereo rectification, and the output is the defect semantic label class, the center coordinates (x, y) of the recognition box, and the width and height (w, h) identified in the left camera image. The semantic label, the center coordinates of the recognition box, and the width and height data are used as object recognition information;
[0090] The stereo matching and depth calculation module takes the left and right camera images of the binocular camera after stereo rectification and the pipeline defect detection information obtained through deep learning object recognition as inputs. After being processed by the stereo matching algorithm, it obtains the disparity map of the left camera of the binocular camera after stereo rectification, and then converts it into a depth map through depth calculation. Finally, combined with the pipeline defect detection information, the output is the spatial three-dimensional coordinates of the recognized object in the left camera image of the binocular camera after stereo rectification mapped to the actual three-dimensional space.
[0091] Through the binocular camera image capture module, the image data of the left and right cameras are captured and collected; through the camera calibration module, a camera imaging geometric model is established and lens distortion is corrected to obtain the internal and external parameters and distortion coefficients of the camera; through the stereo rectification module, coplanar row alignment of the left and right images is achieved, so that the left and right image planes are parallel to the baseline, and the corresponding points in the left and right images are on the same horizontal epipolar line; through the deep learning object detection module, a deep learning network based on the object detection algorithm is trained to realize the detection and recognition of pipeline defects and obtain object recognition information; through the stereo matching and depth calculation module, the spatial three-dimensional coordinates of pipeline defects are obtained to realize the precise positioning and ranging of pipeline defects.
[0092] Combined Figure 2 As shown in
[0093] Step 1: Calibrate the binocular camera, establish the geometric model of camera imaging, determine the mutual relationship between the three-dimensional geometric position of a certain point on the surface of the spatial object and its corresponding point in the image, and solve the internal parameters, external parameters and distortion coefficients of the left and right cameras of the binocular camera from the image coordinates and world coordinates of the known feature points of the calibration board. The internal parameters, external parameters and distortion coefficients of the left and right cameras of the binocular camera are used as camera calibration parameters. By adjusting the internal parameters, external parameters and distortion coefficients, the images captured by the binocular camera can be corrected to obtain images with relatively small distortion;
[0094] The specific implementation steps of Step 1 are as follows:
[0095] Step 1.1: Make a checkerboard calibration board composed of alternating black and white squares. Use a binocular camera to take multiple shots of the checkerboard calibration board at multiple positions, angles, and poses, so that the single-plane checkerboard is clearly imaged in both the left and right cameras. For each calibration image, extract its corner information to obtain the image coordinates of all interior corner points on the calibration image, as well as the three-dimensional spatial coordinates of all interior corner points on the calibration board image;
[0096] Step 1.2: Establish a geometric model of camera imaging to determine the mutual relationship between the three-dimensional geometric position of a certain point on the surface of a spatial object and its corresponding point in the image. These geometric model parameters are the camera calibration parameters, including internal and external parameters and distortion coefficients;
[0097] The external parameter matrix W reflects the conversion between the camera coordinate system and the world coordinate system. Among them, R is the rotation matrix of the right camera of the binocular camera relative to the left camera, and t is the translation vector of the right camera of the binocular camera relative to the left camera. The internal parameter matrix M reflects the conversion between the pixel coordinate system and the camera coordinate system. Among them, f is the lens focal length, (u0, v0) is the coordinate of the origin of the image coordinate system in the pixel coordinate system, and d x 、d y are the sizes of each pixel point in the x-axis and y-axis directions of the image coordinate system:
[0098]
[0099]
[0100] Step 1.3: Take the image coordinates of all interior corner points on the calibration image obtained in Step 1.1 and the three-dimensional spatial coordinates of all interior corner points on the calibration board image as inputs. According to the geometric model of camera imaging, solve and output the internal parameters, external parameters, and distortion coefficients of the left and right cameras of the binocular camera through experiments and calculations;
[0101] Step 1.4: Take the calibrated internal and external parameters of the binocular camera in Step 1.3 as known constants, and use the coordinate information obtained in Step 1.1 and the coordinate relationship before and after correction to solve the five distortion parameters k1, k2, k3, p1, and p2 for distortion correction:
[0102]
[0103] where (x p , y p ) is the original coordinate of the image, and (x tcorr , y tcorr ) is the corrected coordinate of the image, which is approximately described by the Taylor series expansion at r = 0.
[0104] Step 2: Perform stereo rectification through epipolar constraint to make corresponding points in the two images lie on the same horizontal epipolar line, obtain the rectified camera parameters, call OpenCV to obtain the parameters of the rectified left and right cameras in real time, complete the rectification, and finally obtain the rectified images through rectification mapping;
[0105] The specific implementation steps of Step 2 are as follows:
[0106] Step 2.1: Divide the binocular camera rotation matrix R into two parts, the combined rotation matrices r1 and r2 of the left and right cameras. Each of the left and right cameras rotates by half to make the optical axes of the left and right cameras parallel, and achieve coplanarity of the imaging planes of the left and right cameras;
[0107] Step 2.2: Input the combined rotation matrices r1 and r2 of the left and right cameras, the original internal parameter matrices of the left and right cameras, the translation vector t, and the size of the checkerboard image, call the cvStereoRectify function in OpenCV, and output the row-aligned rectification rotation matrices R1 and R2 of the left and right cameras, the rectified internal parameter matrices M l and M r 、the rectified projection matrices P l and P r as well as the reprojection matrix Q;
[0108] Step 2.3: Take the output matrices in Step 2.2 as known constants, through the rectification lookup mapping tables of the left and right views, use inverse mapping to find the floating-point positions on the source images corresponding to each integer pixel position on the target image, and interpolate each integer value of the surrounding source pixels. After all the rectified images are assigned values, crop the images and save the rectification results.
[0109] Step 3: Pre-collect a large number of pipeline defect images through the binocular camera, build a self-built pipeline defect dataset, and screen and enhance the dataset to optimize the dataset. Based on the self-built dataset, perform image annotation on the pipeline defect images. After the annotation is completed, start training the deep convolutional neural network model. After the training is completed, perform inference on the obtained weights, analyze the detection results, and obtain a weight file available for pipeline defect detection. During actual use, the binocular camera image capture module can be used to capture images in real time, and pipeline defect target detection can be performed based on the weight file to obtain pipeline defect detection information;
[0110] Combined Figure 3 As shown, the specific implementation steps of Step 3 are as follows:
[0111] Step 3.1: Data collection, pre-collect a large number of images through the binocular camera, take images containing pipeline defects, and build a self-built pipeline defect dataset;
[0112] Step 3.2: Data Screening and Enhancement. Initially screen the pipeline defect dataset, remove invalid data, and perform data enhancement operations on the images after the initial screening to optimize the pipeline defect dataset for improving the training effect;
[0113] Step 3.3: Image Annotation. Label possible targets, generate annotation files, organize the training directory, and construct the training set, validation set, and test set required for training;
[0114] Step 3.4: Image Training. Perform model iteration based on the pre-trained model, detect the training results through Jupter, and adjust the parameters to prevent overfitting;
[0115] Step 3.5: Image Inference. Infer the weights obtained after image training. The actual captured images should be used to analyze the detection results. If the detection effect is good, the weight file obtained from training can be used for object detection;
[0116] Step 3.6: Object Detection. Through the images captured in real time by the binocular camera, a convolutional neural network that aggregates and forms image features at different image fine-grained levels, mixes and combines image features, and passes the image features to the prediction layer to predict the image features, generate bounding boxes and predict categories, and finally obtain the pipeline defect detection information.
[0117] In the neural network of the YOLOv5 object detection algorithm disclosed in the present invention, it specifically includes:
[0118] I. Data Enhancement. Shrink the image and randomly paste it onto the COCO 2017 dataset to increase the number of the dataset. At the same time, further generalize the enhanced dataset by means of random scaling, random cropping, and random permutation;
[0119] II. Focus Interleaved Sampling and Stitching. Uniformly scale the image to a size of (3, 640, 640) as the input, copy it four times, and cut these four pictures into four slices of (3, 320, 320) through slicing operations. Next, use Concat to connect these four slices from the depth, and the output is (12, 20, 320). Then, through a convolutional layer with 32 convolutional kernels, an output of (32, 320, 320) is generated. Finally, the result is input into the next convolutional layer through batch_borm and leaky_relu;
[0120] 3. Backbone, a convolutional neural network that aggregates and forms image features at different image granularities. Bottlenneck is a classic residual structure, which starts with a 1×1 convolution layer, then a 3×3 convolution layer, and finally adds the residual structure to the initial input. CSP divides the original input into two branches, and performs convolution operations on each branch to halve the number of channels. Branch 1 performs a Bottlenneck×N operation, and then concats branch 1 and branch 2, so that the input and output of BottlenneckCSP are the same size, in order to allow the model to learn more features.
[0121] 4. Neck: A network layer that mixes and combines image features and passes them to the prediction layer. The most important one is the SPP structure. The input of SPP is 512×20×20, and after a 1×1 convolution layer, the output is 256×20×20, and then it is sampled by three parallel MaxPools, and the result is added to its initial features to output 1024×20×20. Finally, a 512 convolution kernel is used to restore it to 512×20×20;
[0122] 5. Head: Predict image features, generate bounding boxes and predict categories.
[0123] For YOLOv5, the Backbone, Neck and Head of various weight files are the same. The only difference is the depth and width settings of the model. You only need to modify these two parameters to adjust the network structure of the model. The parameters of yolov5l are the default parameters.
[0124] Step 4: The left and right camera images of the binocular camera after stereo correction and the pipeline defect detection information obtained through deep learning object recognition are transmitted as input to the stereo matching and depth calculation module. The center coordinates of the identification box in the pipeline defect detection information are used as reference. After the stereo matching and depth calculation module processes the image, the output is obtained, and the spatial three-dimensional coordinates of the identified pipeline defect in the left camera image of the binocular camera after stereo correction are mapped to the actual three-dimensional space.
[0125] Combination Figure 4 As shown, the specific implementation steps of step 4 are as follows:
[0126] Step 4.1: Take the left and right camera images of the stereo camera after stereo correction as input, and calculate the matching cost within the preset parallax range. The purpose of the matching cost calculation is to measure the correlation between the pixels to be matched and the candidate pixels. Whether the two pixels are homonymous points or not, the matching cost can be calculated by the matching cost function. The smaller the cost, the greater the correlation and the greater the probability of being homonymous points. The BT algorithm is used as the calculation method for the matching cost. The calculation formula is:
[0127]
[0128] The matching cost calculation can be realized for the left and right camera images of the binocular camera after stereo rectification within a preset disparity range through a formula, and the matching cost of each pixel point in the original image within the preset disparity range can be obtained;
[0129] Step 4.2: Using the matching cost of each pixel point calculated within the preset disparity range as the input, perform cost aggregation. Adopt the idea of the global stereo matching algorithm, that is, the global energy optimization strategy. Simply put, it is to find the optimal disparity of each pixel to minimize the global energy function of the entire image. The definition of the global energy function is as follows:
[0130] E(d) = E data (d) + E smooth (d);
[0131] Adopt the method of path cost aggregation, that is, perform one-dimensional aggregation of the matching costs under all disparities of the pixel on all paths around the pixel to obtain the path cost value under the path, and then add up all the path cost values to obtain the aggregated matching cost value of the pixel. The calculation method of the path cost of pixel p along a certain path r is as follows:
[0132]
[0133] Among them, p represents the pixel, r represents the path, d represents the disparity, p - r represents the pixel within the path of pixel p, L represents the aggregated cost value of a certain path, Lr(p - r, d) represents the cost value when the disparity of the previous pixel within the path is d, Lr(p - r, d - 1) represents the cost value when the disparity of the previous pixel within the path is d - 1, Lr(p - r, d + 1) represents the cost value when the disparity of the previous pixel within the path is d + 1, min i L r (p - r, i) represents the minimum value of all cost values of the previous pixel within the path;
[0134] The first term is the matching cost C, which belongs to the data term; the second term is the smooth term, and the value accumulated on the path cost takes the minimum value among the three cases of no penalty, P1 penalty, and P2 penalty; P1 is to adapt to inclined or curved surfaces, and P2 is to preserve discontinuities. P2 is often dynamically adjusted according to the gray difference of adjacent pixels, as shown in the following formula:
[0135]
[0136] P2′ is the initial value of P2, generally set to a number much larger than P1, I bp and and Ibq represent the grayscale values of pixels p and q respectively; the third term is to ensure that the new path cost value Lr does not exceed a certain numerical upper limit.
[0137] The total path cost value S can be calculated by the following formula:
[0138]
[0139] Through the calculations of the above formulas, cost aggregation of multiple paths can be achieved, and the multi-path cost aggregation values of each pixel within the preset disparity range can be obtained.
[0140] Step 4.3: Using the multi-path cost aggregation values of each pixel within the preset disparity range as input, perform disparity calculation. Disparity calculation is to determine the optimal disparity value of each pixel through the cost matrix S after cost aggregation. The Winner-take-all algorithm is adopted, that is, among the cost values under all disparities of a certain pixel, select the disparity corresponding to the minimum cost value as the optimal disparity, and finally obtain the disparities of each pixel after cost aggregation.
[0141] Step 4.4: Using the disparities of each pixel after cost aggregation as input, perform disparity optimization. The purpose of disparity optimization is to further optimize the disparity map obtained in the previous step and improve the quality of the disparity map, including eliminating incorrect matches, improving disparity accuracy, and suppressing noise.
[0142] To eliminate incorrect matches, the left-right consistency check method is adopted. It is based on the uniqueness constraint of disparity, that is, each pixel has at most one correct disparity. The specific steps are to swap the positions of the left and right images, that is, the left image becomes the right image, and the right image becomes the left image, and then perform stereo matching again to obtain another disparity map. Since each value in the disparity map reflects the corresponding relationship between two pixels, according to the uniqueness constraint of disparity, through the disparity map of the left image, find the corresponding pixel and its corresponding disparity value in the right image for each pixel. If the absolute value of the difference between these two disparity values is less than 1, it satisfies the uniqueness constraint and is retained; otherwise, it does not satisfy the uniqueness constraint and is eliminated. At the same time, the method of connected component detection is used to eliminate isolated outliers, remove small clusters in the disparity map caused by incorrect matches, and filter small isolated speckles. The formula for consistency check is as follows:
[0143]
[0144] To improve the parallax accuracy, a sub-pixel optimization technique is adopted. The sub-pixel accuracy is obtained by the method of quadratic curve interpolation. The cost values of the optimal parallax and the cost values of the two adjacent parallaxes are fitted with a quadratic curve. The parallax value corresponding to the extreme point of the curve is the new sub-pixel parallax value. To suppress noise, median filtering is used to make the parallax result smoother, eliminate the noise in the parallax map to a certain extent, and play a role in parallax filling at the same time. Through the above steps, parallax optimization can be achieved, and finally the left camera parallax map of the stereo-calibrated binocular camera is obtained.
[0145] Step 4.5: Taking the left camera parallax map of the stereo-calibrated binocular camera as the input, depth calculation can be carried out. The pixel depth calculation formula is as follows:
[0146]
[0147] where f is the focal length, b is the baseline length, d is the parallax, c xr and c xl are the column coordinates of the principal points of the two cameras. After depth calculation, the depth map of the left camera image of the stereo-calibrated binocular camera can be obtained. Combining the pipeline defect detection information obtained by deep learning object detection, the spatial three-dimensional coordinates of the identified pipeline defects mapped to the actual three-dimensional space in the left camera image of the stereo-calibrated binocular camera can be finally obtained.
[0148] Through the above steps, the spatial three-dimensional coordinates of the pipeline defects are obtained, that is, the object detection, positioning and ranging of the pipeline defects are realized. Through binocular stereo vision technology, the three-dimensional reconstruction of the images captured by the binocular camera and the positioning and ranging of the detection target are realized. Combining deep learning technology, the object detection of the pipeline defects is realized, which has the characteristics of full automation, non-contact measurement, accurate positioning and high real-time performance.
[0149] The device obtained by the above technical solution is a pipeline defect detection, positioning and ranging system based on binocular stereo vision. By setting up a binocular camera image capture module, a camera calibration module, a stereo calibration module, a deep learning object detection module, and a stereo matching and depth calculation module, through the working process of the present invention: Step 1, Step 2, Step 3, Step 4, the object detection, positioning and ranging of the pipeline defects can be realized.
[0150] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications; any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A pipeline defect detection, location and ranging system based on binocular stereo vision, characterized in that, It includes a binocular camera image capture module, a camera calibration module, a stereo rectification module, a deep learning object detection module, and a stereo matching and depth calculation module; Between the binocular camera image capture module, the camera calibration module, the stereo rectification module, the deep learning object detection module, and the stereo matching and depth calculation module, the following steps are used to achieve the object detection and positioning ranging of pipeline defects: Step 1: Calibrate the binocular camera, establish a geometric model of camera imaging, determine the mutual relationship between the three-dimensional geometric position of a certain point on the surface of a spatial object and its corresponding point in the image, and solve the internal parameters, external parameters, and distortion coefficients of the left and right cameras of the binocular camera from the image coordinates and world coordinates of the known feature points of the calibration board. The internal parameters, external parameters, and distortion coefficients of the left and right cameras of the binocular camera are used as camera calibration parameters. By adjusting the internal parameters, external parameters, and distortion coefficients, the images captured by the binocular camera can be corrected to obtain images with relatively small distortion. Step 2: Perform stereo rectification through epipolar constraint to make the corresponding points in the two images on the same horizontal epipolar line, obtain the corrected camera parameters, call OpenCV to obtain the parameters of the corrected left and right cameras in real time, complete the correction, and finally obtain the corrected images through the correction mapping. Step 3: Use the binocular camera to pre-collect a large number of pipeline defect images, build a self-built pipeline defect dataset, and screen and enhance the dataset to optimize the dataset. Based on the self-built dataset, perform image annotation on the pipeline defect images. After the annotation is completed, start training the deep convolutional neural network model. After the training is completed, infer the obtained weights, analyze the detection results, and obtain a weight file available for pipeline defect detection. In actual use, the binocular camera image capture module can be used to capture images in real time, and pipeline defect object detection can be performed based on the weight file to obtain pipeline defect detection information. Step 4: Transmit the left and right camera images of the binocular camera after stereo rectification and the pipeline defect detection information obtained through deep learning object recognition to the stereo matching and depth calculation module. Taking the center coordinates of the recognition box in the pipeline defect detection information as the reference quantity, after the stereo matching and depth calculation module processes the image, the spatial three-dimensional coordinates of the recognized pipeline defect mapped to the actual three-dimensional space in the left camera image of the binocular camera after stereo rectification are output. The specific implementation steps of Step 2 are as follows: Step 2.1: Divide the rotation matrix R of the binocular camera into two parts, the combined rotation matrices r1 and r2 of the left and right cameras. Each of the left and right cameras rotates by half to make the optical axes of the left and right cameras parallel, so as to make the imaging planes of the left and right cameras coplanar. Step 2.2: Input the composite rotation matrices r1 and r2 of the left and right cameras, the original intrinsic matrices of the left and right cameras, the translation vector t, and the size of the checkerboard image, and call the cvStereoRectify function in OpenCV to output the row-aligned rectification rotation matrices R1 and R2 of the left and right cameras, the rectified intrinsic matrices M l and M r of the rectified left and right cameras, the projection matrices P l and P r of the rectified left and right cameras, as well as the reprojection matrix Q; Step 2.3: Take the output matrix in Step 2.2 as a known constant, use the correction lookup table of the left and right views, adopt inverse mapping, find the floating-point position on the source image corresponding to each integer pixel position on the target image, and interpolate each integer value of the surrounding source pixels. After the corrected images are all assigned values, crop the images and save the correction results.
2. The pipeline defect detection, location and ranging system based on binocular stereo vision according to claim 1, characterized in that, The specific implementation steps of Step 1 are as follows: Step 1.1: Make a chessboard calibration plate composed of black and white squares, and use a binocular camera to take multiple shots of the chessboard calibration plate at multiple positions, angles, and postures, so that the single-plane chessboard is clearly imaged in both the left and right cameras. For each calibration image, extract its corner point information to obtain the image coordinates of all inner corner points on the calibration image, as well as the spatial three-dimensional coordinates of all inner corner points on the calibration plate image; Step 1.2: Establish a geometric model of camera imaging and determine the relationship between the 3D geometric position of a point on the surface of a spatial object and its corresponding point in the image. These geometric model parameters are the camera calibration parameters, including intrinsic and extrinsic parameters and distortion coefficients. The external parameter matrix W reflects the transformation between the camera coordinate system and the world coordinate system. Among them, R is the rotation matrix of the right camera of the binocular camera relative to the left camera, and t is the translation vector of the right camera of the binocular camera relative to the left camera. The internal parameter matrix M reflects the transformation between the pixel coordinate system and the camera coordinate system. Among them, f is the lens focal length, and (u0, v0) is the coordinate of the origin of the image coordinate system in the pixel coordinate system, d x , d y are the sizes of each pixel point in the x-axis and y-axis directions of the image coordinate system: Step 1.3: Take the image coordinates of all the inner corner points on the calibration image obtained in step 1.1 and the spatial three-dimensional coordinates of all the inner corner points on the calibration plate image as input, and solve and output the intrinsic parameters, extrinsic parameters, and distortion coefficients of the left and right cameras of the binocular camera through experiments and calculations based on the geometric model of camera imaging; Step 1.4: Take the internal and external parameters of the binocular camera calibration in step 1.3 as known constants, and use the coordinate information obtained in step 1.1, and the coordinate relationship before and after correction to solve the five distortion parameters k1, k2, k3, p1, and p2 for distortion correction: where (x p , y p ) are the original coordinates of the image, and (x tcorr , y tcorr ) are the coordinates of the image after correction, which are approximately described by the Taylor series expansion at r = 0.
3. The pipeline defect detection, location and ranging system based on binocular stereo vision according to claim 1, characterized in that, The specific implementation steps of step three are as follows: Step 3.1: Data collection: collect a large number of images in advance through a binocular camera, take images containing pipeline defects, and build a pipeline defect dataset; Step 3.2: Data screening and enhancement: Preliminary screening of the pipeline defect dataset, elimination of invalid data, and data enhancement operations on the images after preliminary screening to optimize the pipeline defect dataset in order to improve the training effect; Step 3.3: Annotate the image, annotate possible objects, generate annotation files, organize the training directory, and build the training set, validation set, and test set required for training; Step 3.4: Image training: perform model iteration based on the pre-trained model, detect the training results through Jupter, and adjust the parameters to prevent overfitting; Step 3.5: Image inference, infer the weights obtained after image training. The actual captured images should be used to analyze the detection results. If the detection effect is good, the trained weight file can be used for target detection. Step 3.6: Target detection, through the images captured by the binocular camera in real time, aggregate and form a convolutional neural network of image features at different image granularities, mix and combine image features, and pass the image features to the prediction layer to predict the image features, generate bounding boxes and predict categories, and finally obtain pipeline defect detection information.
4. The pipeline defect detection, location and ranging system based on binocular stereo vision according to claim 1, characterized in that, Step 4: The specific implementation steps are as follows: Step 4.1: Take the stereo-corrected left and right camera images of the binocular camera as input, and calculate the matching cost within the preset parallax range. The purpose of the matching cost calculation is to measure the correlation between the pixels to be matched and the candidate pixels. Whether the two pixels are homonymous points or not, the matching cost can be calculated by the matching cost function. The smaller the cost, the greater the correlation and the greater the probability of being homonymous points. The BT algorithm is used as the calculation method of the matching cost. Implement the calculation of the matching cost for the left and right camera images of the binocular camera after stereo rectification within a preset disparity range, and obtain the matching cost of each pixel point in the original image within the preset disparity range; Step 4.2: Use the matching cost of each pixel point calculated within the preset disparity range as the input for cost aggregation. Adopt the idea of the global stereo matching algorithm, that is, the global energy optimization strategy. Simply put, it is to find the optimal disparity for each pixel to minimize the global energy function of the entire image. The definition of the global energy function is as follows: E(d) = E data (d) + E smooth (d); Adopt the method of path cost aggregation, that is, perform one-dimensional aggregation of the matching costs under all disparities of the pixel along all paths around the pixel to obtain the path cost value under the path, and then add up all the path cost values to obtain the aggregated matching cost value of the pixel. The calculation method of the path cost of pixel p along a certain path r is as follows: Wherein, p represents a pixel, r represents a path, d represents a disparity, p-r represents the pixels within the path of pixel p, L represents an aggregation cost value of a certain path, Lr(p-r, d) represents the cost value when the disparity of the previous pixel within the path is d, Lr(p-r, d-1) represents the cost value when the disparity of the previous pixel within the path is d-1, Lr(p-r, d+1) represents the cost value when the disparity of the previous pixel within the path is d+1, and min i L r (p-r, i) represents the minimum value of all cost values of the previous pixel within the path; The first term is the matching cost C, which belongs to the data term; The second term is the smooth term. The value accumulated to the path cost takes the minimum value among the three cases of no penalty, P1 penalty, and P2 penalty; P1 is to adapt to inclined or curved surfaces, and P2 is to preserve discontinuities. P2 is dynamically adjusted according to the gray difference of adjacent pixels, as shown in the following formula: P2' is the initial value of P2, set to a number much larger than P1, I bp and and l bq represent the grayscale values of pixels p and q respectively; The third term is to ensure that the new path cost value Lr does not exceed a certain numerical upper limit, The total path cost value S can be calculated by the following formula: Through the above calculations, the cost aggregation of multiple paths can be realized, and the multi-path cost aggregation value of each pixel within the preset disparity range can be obtained; Step 4.3: Use the multi-path cost aggregation value of each pixel within the preset disparity range as the input for disparity calculation. Disparity calculation is to determine the optimal disparity value of each pixel through the cost matrix S after cost aggregation. Adopt the Winner-take-all algorithm, that is, among the cost values under all disparities of a certain pixel, select the disparity corresponding to the minimum cost value as the optimal disparity, and finally obtain the disparity of each pixel after cost aggregation; Step 4.4: Use the disparity of each pixel after cost aggregation as the input for disparity optimization. The purpose of disparity optimization is to further optimize the disparity map obtained in the previous step and improve the quality of the disparity map, including eliminating incorrect matches, improving disparity accuracy, and suppressing noise; The method of eliminating incorrect matches adopts the left-right consistency check method, which is based on the uniqueness constraint of disparity, that is, each pixel has at most one correct disparity. The specific steps are to swap the positions of the left and right images, that is, the left image becomes the right image and the right image becomes the left image, and then perform stereo matching again to obtain another disparity map. Because each value in the disparity map reflects the corresponding relationship between two pixels, according to the uniqueness constraint of disparity, through the disparity map of the left image, find the corresponding pixel and its corresponding disparity value of each pixel in the right image. If the absolute value of the difference between these two disparity values is less than 1, it satisfies the uniqueness constraint and is retained, otherwise it does not satisfy the uniqueness constraint and is eliminated. At the same time, the method of connected component detection is used to eliminate isolated outliers and remove small clusters in the disparity map caused by incorrect matches and filter small isolated speckles; To improve the parallax accuracy, a sub-pixel optimization technique is adopted. The sub-pixel accuracy is obtained by using the quadratic curve interpolation method. The cost values of the optimal parallax and the cost values of the two adjacent parallaxes are fitted with a quadratic curve. The parallax value corresponding to the extreme point of the curve is the new sub-pixel parallax value. To suppress noise, median filtering is used to make the parallax result smoother, eliminate the noise in the parallax map to a certain extent, and at the same time play a role in parallax filling. Through the above steps, parallax optimization can be achieved, and finally the parallax map of the left camera of the binocular camera after stereo correction can be obtained; Step 4.5: Using the parallax map of the left camera of the binocular camera after stereo correction as the input, depth calculation can be performed. The pixel depth calculation formula is as follows: where f is the focal length, b is the baseline length, d is the parallax, and c xr and c xl are the column coordinates of the principal points of the two cameras. After depth calculation, the depth map of the left camera image of the binocular camera after stereo rectification can be obtained. Combining the pipeline defect detection information obtained through deep learning object detection, the spatial three-dimensional coordinates of the identified pipeline defects in the left camera image of the binocular camera after stereo rectification mapped to the actual three-dimensional space can be finally obtained.
Citation Information
Patent Citations
Method and device for detecting internal defects of pipeline through stereoscopic vision imaging
CN113487490A
Pipeline defect identification and positioning method based on target detection and binocular vision
CN114067197A
Cited By
PCB defect detection method and device based on binocular stereo vision and improved YOLO network
CN120912553A