Pipeline fault intelligent detection method based on multi-model fusion
Through multi-model fusion technology and combined with YOLOv5 model, intelligent detection of pipeline faults is solved, and the shortcomings of the existing technology in subtle defect identification and complex environments are achieved, and efficient and accurate pipeline fault detection is achieved.
Patent Information
- Application Number
- CN202510185785.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
Existing pipeline fault detection technology has limitations in identifying subtle defects and defect information extraction in complex environments, and it is difficult to meet the real-time detection needs.
Using intelligent pipeline fault detection method based on multi-model fusion, we use multiple types of training data sets and expert detection models, and combined with YOLOv5 object detection model, we realize multi-model collaborative prediction of internal images of the pipeline, reducing the error rate and enhancing the stability and reliability of detection.
It effectively avoids the preference limitations of a single model for specific defects, achieves comprehensive coverage and accuracy of detection, significantly reduces the error prediction rate, and enhances the robustness and stability of the detection system.
Smart Images

Figure CN120125527A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing and intelligent detection, and specifically to a pipeline fault intelligent detection method based on multi-model fusion, which is suitable for the identification and diagnosis of pipeline faults. Background Art
[0002] With the accelerated pace of urbanization and the continuous expansion of impermeable pavement coverage, coupled with the increasing frequency of heavy rainfall events caused by climate change, the drainage pipe system frequently encounters hidden dangers such as blockage, deformation, and leakage under high-intensity operation. These hidden dangers not only weaken the drainage efficiency, but also disrupt the normal operation of the underground pipe network, thereby inducing secondary disasters such as waterlogging and road collapse, bringing a heavy burden to the social economy. Therefore, exploring efficient and accurate drainage pipe detection technology has become a key issue that needs to be urgently addressed in current urban construction.
[0003] However, with the continuous expansion of urban pipeline networks, a comprehensive survey of pipeline health status faces the dual challenges of time and human resources, and its urgency cannot be ignored. In recent years, the emergence of advanced technologies such as periscope detection, sonar detection, and CCTV pipeline endoscopes has improved detection efficiency and operational safety to a certain extent, and can intuitively display the internal conditions of pipelines. However, such technologies still have limitations, such as the difficulty in identifying subtle defects and insufficient defect information extraction in complex pipeline environments, and urgently need technical iteration and optimization.
[0004] Vision-based pipeline endoscopy detection technology, as the fastest-growing pipeline detection method in the past three decades, uses robots equipped with high-definition cameras to capture images inside pipelines and transmit them to the ground in real time for defect analysis. It has been widely used in Europe, the United States, Japan and other places, and has gradually penetrated into the Chinese market. It has been recognized as an effective detection method by the "Technical Regulations for the Inspection and Evaluation of Urban Drainage Pipelines (CJJ-181-2012)". Compared with sonar and pipeline periscopes, this technology, with the combination of high-definition cameras and robots, has demonstrated higher detection accuracy and a wider range of operations, indicating that it will become the mainstream technology in the field of pipeline detection in the future.
[0005] Looking back at past research, in 2002, Fieguth et al. proposed morphological segmentation and fuzzy neural network recognition methods in their literature, but they were limited by the single segmentation type and the long time consumption, making it difficult to meet the needs of real-time detection. In 2008, Ming-DerYang's team used wavelet transform and machine learning technology to perform defect segmentation on pipeline images with a single background, but the performance on complex backgrounds and low-quality images was not satisfactory. In 2012, Sun Wenya et al. realized the segmentation of pipeline crack images based on BP neural network, but the effect was greatly reduced when faced with defects with small grayscale differences. Since then, Alam et al. (2014), Liu Zhe et al. (in the same year), and Huang Yulong's team (2018) have respectively proposed methods based on the Sobel operator, spatial and frequency domain denoising combined with the Canny operator, image denoising and dynamic threshold processing. Although each has its own advantages, they still face difficulties such as high image quality requirements, weak generalization ability, and insufficient support for multiple defect detection. They are easily affected by distortion, noise and changes in lighting conditions, which limits their application scope in actual drainage pipe inspection.
[0006] With the rapid development of convolutional neural network (CNN) technology, CNN-based drainage pipe defect detection algorithms have made significant breakthroughs. In 2018, Kumar et al. used a deep CNN system to train a large number of pipe images and achieved high accuracy and recall rates. In the same year, Cheng et al., based on the optimization of the Faster R-CNN algorithm, proved that increasing the data set and the number of convolutional layers can effectively improve detection performance. Since then, Wang Qing et al. (2019), Zhou Qianqian et al. (2021), and Wang Dacheng et al. (2022) have further promoted the development of this field by introducing K-means clustering optimization, expanding the types of defects to be detected, improving detection efficiency, and realizing automatic report generation. Li Wei et al. combined the Mask R-CNN technology to achieve accurate identification and rating of various defects such as cracks, corrosion, and obstacles.
[0007] Although deep learning has shown great potential in pipeline defect detection, it still faces problems such as the need to improve the accuracy and efficiency of the detection algorithm, poor detection of defects with low feature discrimination, and insufficient learning ability for small sample class imbalanced data. In addition, the lack of large public data sets and the scarcity of labeled data still seriously restrict the improvement of algorithm performance. Therefore, designing efficient algorithms to enhance the recognition ability of defects with limited feature discrimination is of great significance to promote the further development of technology in this field. Summary of the invention
[0008] The purpose of this application is to provide a pipeline fault intelligent detection method based on multi-model fusion, which adopts multiple model strategies to make up for the low prediction accuracy of a single model, reduce the false detection rate, and enhance the stability and reliability of the entire prediction model.
[0009] To achieve the above object, the present application provides the following technical solutions:
[0010] In a first aspect, the present application proposes an intelligent pipeline fault detection method based on multi-model fusion. Among them, the method includes the following steps:
[0011] Determine the defect feature categories based on the acquired internal pipeline images, collect the internal pipeline images with the same defect feature categories, and construct multiple types of training data sets;
[0012] Input each type of training data set into the corresponding type of detection model for training to obtain multiple expert detection models;
[0013] Input the real-time acquired internal pipeline images into the expert detection models for prediction respectively, splice the prediction results of each expert detection model in the column direction, align them row by row to form a new tensor, and obtain the preliminary pipeline detection results;
[0014] Based on the non-maximum suppression algorithm, remove redundant prediction boxes to obtain the pipeline detection results.
[0015] As a specific solution in the technical solution of the present application, the determining the defect feature categories based on the acquired internal pipeline images, collecting the internal pipeline images with the same defect feature categories, and constructing multiple types of training data sets includes:
[0016] Select the internal pipeline images with defective feature categories;
[0017] Label the defect feature categories of the acquired internal pipeline images;
[0018] Divide and classify the internal pipeline images with the same defect feature categories into one category to construct multiple types of training data sets;
[0019] Among them, the defect feature categories include: misalignment, residual wall, corrosion, rupture, tree root, leakage, shedding, and obstacle.
[0020] As a specific solution in the technical solution of the present application, the detection model uses the YOLOv5 object detection model, including a backbone network, a feature fusion network, and a detection head.
[0021] As a specific solution in the technical solution of the present application, the inputting each type of training data set into the corresponding type of detection model for training to obtain multiple expert detection models includes:
[0022] Preprocess the training data set images input into the detection model;
[0023] Extract defect features from the preprocessed image to generate first feature maps of multiple sizes;
[0024] Concatenate the first feature maps along the channel dimension to obtain a second feature map, and perform two-way fusion of the second feature map in a top-down and bottom-up manner through the C3 module and the cross-layer feature aggregation mechanism of PANet;
[0025] Based on the second feature map after two-way fusion and the preset anchor boxes, the detection model performs forward propagation to predict the bounding box positions, confidence levels, and class probabilities on the feature map;
[0026] By calculating the loss values of the bounding box regression loss function, object confidence loss function, and classification loss function, the loss values are backpropagated from the output layer to each level of the detection model network, and the parameters of the detection model network are updated using the gradient descent algorithm. After multiple repeated trainings and updates, an expert detection model is obtained.
[0027] As a specific solution in the technical solution of this application, before inputting each type of training data set into the corresponding type of detection model for training, it further includes initializing the YOLOv5 model using a pre-trained model trained on the COCO data set.
[0028] As a specific solution in the technical solution of this application, the preprocessing of the training data set images input into the detection model includes:
[0029] Adjust the size of the training data set images input into the detection model to 640×640;
[0030] Perform data augmentation on the images through, including but not limited to, random flipping, color gamut transformation, HSV adjustment, MixUp, and Mosaic.
[0031] As a specific solution in the technical solution of this application, the extracting defect features from the preprocessed image to generate first feature maps of multiple sizes includes:
[0032] Downsample the image input into the detection model by 4 times through two convolutional layers to extract low-level defect features;
[0033] Use 3 convolutional layers, 4 C3 modules, and 1 SPPF module for multiple downsamplings and defect feature extractions;
[0034] Generate three multi-scale feature maps P3(80×80), P4(40×40), and P5(20×20).
[0035] As a specific solution in the technical solution of this application, the performing two-way fusion of the second feature map in a top-down and bottom-up manner through the C3 module and the cross-layer feature aggregation mechanism of PANet includes:
[0036] Top - down fusion starts from the feature map with lower resolution, passes information upward layer by layer through upsampling operations, and performs fusion operations with the high - resolution feature map;
[0037] Bottom - up fusion starts from the high - resolution feature map, passes information downward layer by layer through downsampling operations, and performs fusion operations with the low - resolution feature map.
[0038] As a specific solution in the technical solution of this application, based on the second feature map after bidirectional fusion and the preset anchor boxes, the detection model predicts the bounding box positions, confidence levels, and class probabilities of the defect feature classes on the feature map through forward propagation, including:
[0039] The initial coordinates of the anchor box are preset as (x a , y a , w a , h a ), where x a , y a are the center coordinates of the anchor box, and w a , h a are the width and height of the anchor box);
[0040] The detection model extracts features from the input second feature map through a convolutional neural network and outputs a quadruple (Δx, Δy, Δw, Δh) at each position of the second feature map, which is the predicted offset relative to the anchor box;
[0041] According to the predicted offset and the initial coordinates of the anchor box, the coordinates of the predicted bounding box (x, y, w, h) are calculated (where x, y are the center coordinates of the box, and w, h are the width and height of the box),
[0042]
[0043] The detection model outputs a confidence score at each anchor box position;
[0044] The detection model uses the Softmax function to calculate the probability of each defect feature class.
[0045] As a specific solution in the technical solution of this application, calculating the loss value by calculating the bounding box regression loss function, the object confidence loss function, and the classification loss function includes:
[0046] (1)
[0047] Among them, IoU represents the intersection - over - union ratio, which measures the overlap degree between the predicted box and the ground - truth box; ρ 2The Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box; c 2 The diagonal length of the minimum bounding box of the predicted bounding box and the ground truth bounding box; α is the weight factor; v is the difference in the width-to-height ratio between the predicted bounding box and the ground truth bounding box.
[0048] (2)
[0049] Among them, p obj represents the predicted object confidence; t obj represents the ground truth object confidence (1 if the object exists, otherwise 0); BCE is the binary cross-entropy, which can be expressed as:
[0050]
[0051] Among them, N is the total number of samples; y i represents the ground truth label of the i-th sample; represents the predicted value (model output) of the i-th sample.
[0052] (3)
[0053] Among them, C is the number of classes; p cls,c is the predicted probability of the c-th class; t cls,c is the ground truth label of the c-th class
[0054] The total loss function is defined as and a weight factor is introduced for each loss term:
[0055]
[0056] Among them, λ bbox = 0.05, λ obj = 1.0, λ cls = 0.5.
[0057] As a specific solution in the technical solution of this application, removing redundant predicted bounding boxes based on the non-maximum suppression algorithm to obtain the pipeline detection result includes:
[0058] Setting the confidence threshold to 0.3 and the intersection over union threshold to 0.25, screening out the predicted bounding boxes with object confidence greater than the confidence threshold, generating a candidate box mask, and removing the candidate boxes with object confidence lower than the confidence threshold;
[0059] Fusing the object confidence and the class confidence, and the confidence conf final (c) of each class is equal to the product of the object confidence conf obj and the class confidence conf cls (c). It can be expressed as:
[0060] conffinal (c) = conf obj ·conf cls (c);
[0061] Sort the candidate boxes in descending order of confidence and retain the top 30k candidate boxes;
[0062] Remove redundant boxes using IoU. IoU (Intersection over Union) formula:
[0063]
[0064] Remove boxes with excessive overlap and only retain boxes with IoU less than iou_thres;
[0065] Output the maximum number of detection boxes, and the number of detection boxes does not exceed 300.
[0066] As a specific solution in the technical solution of this application, the pipeline detection result includes the bounding box position, confidence, and class label, and the output format is [x1, y1, x2, y2, conf, cls], where x1, y1, x2, y2 are the bounding box coordinates, conf is the confidence, and cls is the class.
[0067] Compared with the prior art, the beneficial effects of this application are as follows: This application successfully breaks the limitation of the "preference" of a single model for specific defects, effectively avoids the omission or misjudgment of other defects caused thereby, and realizes the comprehensive coverage and accuracy of detection. By introducing a mechanism of multi-model collaborative work, it can mutually complement the prediction shortboards of a single model in special or complex situations, significantly reduce the error prediction rate, and greatly enhance the robustness and stability of the overall detection system. It also effectively reduces the risk of overfitting of the detection model to noise features. Even in the case of insufficient light and complex structure inside the pipeline, it can ensure that the detection model maintains high accuracy, thereby greatly enhancing the reliability of the detection results. In addition, this application demonstrates high flexibility and scalability, can seamlessly dock with various detection models, and flexibly adapt to diverse model architectures and task requirements, providing a wider selection space and infinite possibilities for the actual application field. Description of the Drawings
[0068] Figure 1 It is a flowchart of the intelligent pipeline fault detection method based on multi-model fusion proposed in the embodiment of this application;
[0069] Figure 2 It is a structure diagram of the YOLOv5 model proposed in the embodiment of this application;
[0070] Figure 3 It is the average loss and accuracy curve of multi-model prediction provided in the example of this application;
[0071] Figure 4 The defect classification criteria provided for the examples of this application;
[0072] Figure 5 Quantitative comparison of the experimental results provided by the examples of this application. Specific implementation manners
[0073] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be further described clearly and completely below. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] To improve the problems such as low accuracy of the detection algorithm, low feature discrimination in defect detection, and insufficient learning ability for small-sample class-imbalanced data proposed in the background art, this embodiment provides an intelligent pipeline fault detection method based on multi-model fusion. The method includes the following steps:
[0075] As Figures 1-3 shown, this embodiment provides an intelligent pipeline fault detection method based on multi-model fusion. The method includes the following steps:
[0076] Step S100: Obtain the internal image of the pipeline.
[0077] Step S200: Perform image preprocessing and construct a data set.
[0078] Step S300: Use different data sets to train expert detection models for different defects respectively.
[0079] Step S400: Use multiple trained expert detection models to predict the pipeline pictures.
[0080] Step S500: Horizontally (in the column direction) splice the prediction results of multiple detection models, align them row by row, and integrate the results from different models into a tensor.
[0081] Step S600: Use the non-maximum suppression (NMS) algorithm to remove redundant prediction boxes.
[0082] Step S700: Output the pipeline detection result.
[0083] The specific step description is as follows:
[0084] Step S100: Obtain the internal image of the pipeline. In this embodiment, a total of 3,463 pipeline images containing defects are collected.
[0085] Step S200: Image preprocessing to construct a training dataset.
[0086] Step S201: Select and delete the internal pipeline images that are defect-free, unclear, or have unrecognizable image content.
[0087] Step S202: Manually annotate them according to the Figure 4 standards. There are a total of 8 types of image defect feature categories, namely: misalignment (972 images), residual wall (16 images), corrosion (438 images), rupture (548 images), tree root (319 images), leakage (240 images), peeling (494 images), and obstacle (443 images). These defect feature categories are represented by the codes CK, CQ, FS, PL, SG, SL, TL, and ZW respectively, as shown in Figure 4 the figure.
[0088] Step S203: Divide the internal pipeline images with the same defect feature category into one class and make a dataset. A total of 8 training datasets are made in this embodiment.
[0089] Step S300: Input the training datasets into the corresponding detection models for training respectively.
[0090] Step S301: Take the misalignment defect type as an example in this embodiment to train and obtain an expert detection model. As shown in Figure 5As shown, YOLOv5s achieved 140 FPS, meaning it can process images very quickly and is highly suitable for real-time applications or scenarios that require quick responses. With the constructed model, the latency is only 7.1 milliseconds, the lowest among all the listed models. Although the mAP@[Iou=0.25:0.75] of YOLOv5s is 36.4 and the mAP@[Iou=0.25] is 56.3, which are not optimal, it achieves an effective balance between accuracy and speed. Compared with other models, YOLOv5s has significant advantages in terms of speed and latency. Therefore, the model of this application is constructed based on YOLOv5s, and its architecture mainly consists of three parts: the backbone network, the feature fusion network, and the detection head. Before model training, we first uniformly adjusted the size of all images in the training dataset to 640×640 pixels, and then used a series of data augmentation techniques, including random flipping, color gamut transformation (such as HSV adjustment), MixUp, and Mosaic, etc., to process the input images to improve the generalization ability of the model. To initialize the YOLOv5 model, we selected a model pre-trained on the COCO dataset. In the training process, the input image first undergoes spatial downsampling processing, and two convolutional layers are used to reduce the image size by 4 times to extract low-level features. Subsequently, the model generates three feature maps of different scales: P3 (80×80), P4 (40×40), and P5 (20×20) through a combination of 3 convolutional layers, 4 C3 modules, and 1 SPPF module, after multiple downsampling and feature extraction steps. These feature maps are respectively suitable for detecting small, medium, and large targets. Because these feature maps contain multi-level information of the image, capturing details and context information at different resolutions. Next, these feature maps are concatenated along the channel dimension to obtain a fused feature map. In this way, the context information from different scales is effectively combined, enhancing the model's perception ability for targets of different sizes.
[0091] After that, the C3 module is used for the bidirectional fusion process, and this module implements the cross-layer feature aggregation mechanism in PANet. Specifically, the second feature map is bidirectionally fused in two aspects: top-down and bottom-up through the C3 module and the cross-layer feature aggregation mechanism of PANet. Top-down fusion starts from the feature map with lower resolution, and the information is passed upward layer by layer through upsampling operations and fused with the high-resolution feature map; bottom-up fusion starts from the high-resolution feature map, and the information is passed downward layer by layer through downsampling operations and fused with the low-resolution feature map. This further strengthens the flow of information, enabling better transfer of features between different levels. This cross-layer feature fusion not only enhances the expression of context information but also improves the model's performance in multi-scale object detection.
[0092] Finally, the detection model predicts the position, confidence, and class of the bounding box based on the preset anchor boxes and the extracted feature maps. Specifically, the initial coordinates of the anchor boxes are preset as (x a , y a , w a , h a ), where x a , y a are the center coordinates of the anchor box, and w a , h a are the width and height of the anchor box. Each anchor box predicts the offset (change amount) relative to its position, and these offsets are used to adjust the anchor box to more accurately match the true bounding box of the target object;
[0093] The detection model extracts features from the input second feature map through a convolutional neural network and outputs a quadruple (Δx, Δy, Δw, Δh) at each position of the second feature map, which is the predicted offset relative to the anchor box;
[0094] Based on the predicted offsets and the initial coordinates of the anchor boxes, the coordinates (x, y, w, h) of the predicted bounding box are calculated (where x, y are the center coordinates of the box, and w, h are the width and height of the box),
[0095]
[0096] The detection model outputs a confidence score (the product of the probability of the predicted defect occurrence and the intersection over union of the predicted box and the true bounding box) at each anchor box position, indicating whether there is an object at that position and the overlap degree between the predicted box and the true box.
[0097] The class prediction is based on the class label of the object predicted by each anchor box. For each anchor, the model predicts the probability of each class and uses the Softmax function to calculate the probability of each class.
[0098] During the entire training process, the loss value is calculated by the bounding box regression, object confidence, and classification loss functions to optimize the detection model. Among them, the bounding box regression loss function improves the accuracy of the model in object localization by minimizing the coordinate difference between the predicted bounding box and the ground truth bounding box, enabling the model to accurately regress the position of the target. The object confidence loss function optimizes the model's judgment on the presence or absence of an object by minimizing the difference between the confidence of the predicted box and the confidence of the ground truth box, thereby improving the accuracy of the object presence judgment. The classification loss function enables the model to accurately classify each detected target by minimizing the difference between the predicted class probability and the ground truth class label. Specifically, the optimization process of the detection model is as follows: First, the input image is propagated forward through the neural network to generate prediction results. Using the chain rule, the gradient of the loss function is backpropagated to each layer of the network. The gradient descent algorithm is used to update the network parameters to reduce the loss. During the entire training process, the above steps are continuously repeated: input data -> forward propagation -> calculate loss -> backpropagation -> update weights. Through multiple rounds of training, the model will be continuously optimized, gradually reducing the loss and improving the accuracy of defect prediction.
[0099] The loss values calculated by the bounding box regression loss function, object confidence loss function, and classification loss function are respectively expressed as:
[0100] (1) Among them, IoU represents the intersection over union, which measures the overlap degree between the predicted box and the ground truth box; ρ 2 is the Euclidean distance between the centers of the predicted box and the ground truth box; c 2 is the length of the diagonal of the minimum bounding box of the predicted box and the ground truth box; α is the weight factor; v is the difference in the aspect ratio between the predicted box and the ground truth box.
[0101] (2) Among them, p obj represents the predicted object confidence; t obj represents the ground truth object confidence (1 for the presence of the object, otherwise 0); BCE is the binary cross-entropy, which can be expressed as:
[0102]
[0103] Among them, N is the total number of samples; y i represents the ground truth label of the i-th sample; represents the predicted value (model output) of the i-th sample.
[0104] (3) Among them, C is the number of classes; p cls,c is the probability of the c-th class predicted; t cls,c is the ground truth label of the c-th class. The total loss function is defined as And a weight factor is introduced for each loss:
[0105]
[0106] In the present invention, λ bbox = 0.05, λ obj = 1.0, λ cls = 0.5.
[0107] Step S302: Repeat step S301 to train the remaining 7 expert models (broken wall, corrosion, crack, root, leakage, peeling, obstacle).
[0108] Step S400: Use multiple trained expert detection models to predict the real-time acquired internal pipeline image input. The tensor shape of the detection result output by each model is [1, 20160, 13].
[0109] Step S500: Concatenate the prediction results of multiple expert models in dimension 1 (column direction) of the tensor output by each model. The tensor shape of the output is [1, 161280, 13].
[0110] Specifically, the real-time acquired internal pipeline image is respectively input into multiple expert detection models for prediction to obtain the original prediction results of each expert model. The prediction output of each expert model is a tensor with a shape of [1, 20160, 13], which contains bounding box coordinates, confidence levels, and class information. Then, the prediction results of each expert model are concatenated in dimension 1 (i.e., column direction) through horizontal concatenation to form a new tensor. For example, if there are N expert models, the shape of the concatenated tensor is [1, 20160×N, 13], that is, the number of columns becomes N times the original. Then, the rows of the concatenated tensor are aligned row by row to ensure that the data of each prediction box is correctly corresponding in the row to maintain the validity of the prediction results of each expert at the same position. In this embodiment, there are 8 expert detection models. Finally, all the concatenated and aligned data are integrated into a unified tensor with a shape of [1, 161280, 13] as the preliminary result of pipeline detection.
[0111] Step S600: Use the non-maximum suppression (NMS) algorithm to remove redundant prediction boxes.
[0112] Step S601: Set the confidence threshold (conf_thres) to 0.3 and the intersection over union threshold (iou_thres) to 0.25. Filter out the boxes with the target confidence greater than conf_thres, generate a candidate box mask, and remove the candidate boxes with the target confidence lower than the threshold to reduce the computational amount.
[0113] Step S602: Combine the object confidence and the class confidence. The confidence conf final (c) of each class is equal to the product of the object confidence conf obj and the class confidence conf cls (c). It can be expressed as:
[0114] con final (c) = conf obj · conf cls (c)
[0115] Step S603: Convert the bounding box in the center point form to the top - left and bottom - right form.
[0116] Step S604: Sort the candidate bounding boxes in descending order of confidence and retain the top 30k candidate bounding boxes.
[0117] Step S605: Use IoU to remove redundant bounding boxes. The IoU (Intersection over Union) formula:
[0118]
[0119] Remove the bounding boxes with too much overlap and only retain the bounding boxes with IoU less than iou_thres.
[0120] Step S606: Output the maximum number of detected bounding boxes. The final number of detected bounding boxes does not exceed 300.
[0121] Step S700: Output the pipeline detection results, showing the bounding boxes, confidences, and class labels. The output format for each image is: [x1, y1, x2, y2, conf, cls], where x1, y1, x2, y2 are the coordinates of the bounding box; conf is the confidence; cls is the class index.
[0122] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A pipeline fault intelligent detection method based on multi-model fusion, characterized in that: The method comprises the following steps: Determine the defect feature category based on the acquired internal images of the pipeline, group the internal images of the pipeline with the same defect feature category, and construct multiple types of training data sets; Input each type of training data set into the corresponding type of detection model for training, and obtain multiple expert detection models; The real-time internal images of the pipeline are input into the expert detection model for prediction, and the prediction results of each expert detection model are spliced in the column direction and aligned row by row to form a new tensor to obtain the preliminary results of pipeline detection; Based on the non-maximum suppression algorithm, redundant prediction boxes are removed to obtain pipeline detection results.
2. The method according to claim 1, characterized in that The defect feature category is determined based on the acquired internal image of the pipeline, and the internal images of the pipeline with the same defect feature category are grouped to construct multiple types of training data sets, including: Select images of the interior of the pipe with defective feature classes; Annotate the defect feature categories of the acquired pipeline internal image; The internal images of pipelines with the same defect feature category are divided into one category to construct multiple types of training data sets; Among them, the defect feature categories include: dislocation, residual wall, corrosion, crack, tree root, leakage, falling off, and obstacle.
3. The method according to claim 1, characterized in that: The detection model adopts the YOLOv5 target detection model, including a backbone network, a feature fusion network and a detection head.
4. The method according to claim 1, characterized in that: Each type of training data set is input into the corresponding type of detection model for training to obtain multiple expert detection models, including: Preprocess the training dataset images for input detection model; Extract defect features based on the preprocessed image and generate first feature maps of multiple sizes; The first feature map is concatenated along the channel dimension to obtain the second feature map, which is then bidirectionally fused from top to bottom and from bottom to top through the C3 module and the cross-layer feature aggregation mechanism of PANet. Based on the second feature map after bidirectional fusion and the preset anchor box, the detection model predicts the bounding box position, confidence and category probability on the feature map through forward propagation; By calculating the loss values of the bounding box regression loss function, the target confidence loss function and the classification loss function, the loss values are back-propagated from the output layer to each level of the detection model network, and the gradient descent algorithm is used to update the parameters of the detection model network. After repeated training and updating, an expert detection model is obtained.
5. The method according to claim 4, characterized in that The preprocessing of the training data set image input to the detection model includes: The image size of the training dataset for the input detection model is adjusted to 640×640; Data enhancement is performed on images through methods including but not limited to random flipping, color gamut transformation, HSV adjustment, MixUp, and Mosaic.
6. The method according to claim 4, characterized in that The step of extracting defect features based on the preprocessed image and generating first feature maps of multiple sizes includes: The image input to the detection model is downsampled by 4 times through two convolutional layers to extract low-level defect features; Three convolutional layers, four C3 modules, and one SPPF module are used for multiple downsampling and defect feature extraction; Generate three multi-scale feature maps P3 (80×80), P4 (40×40), and P5 (20×20).
7. The method according to claim 4, characterized in that The top-down and bottom-up bidirectional fusion of the second feature map is performed through the cross-layer feature aggregation mechanism of the C3 module and PANet, including: Top-down fusion starts from a lower-resolution feature map, passes information upward layer by layer through upsampling operations, and performs fusion operations with high-resolution feature maps; Bottom-up fusion starts from a high-resolution feature map, passes information downward layer by layer through downsampling operations, and performs fusion operations with low-resolution feature maps.
8. The method according to claim 4, characterized in that Based on the second feature map after bidirectional fusion and the preset anchor point box, the detection model predicts the bounding box position, confidence and category probability of the defect feature category on the feature map through forward propagation, including: The initial coordinates of the anchor box are preset as , is the center point coordinate of the anchor box, is the width and height of the anchor box); The detection model extracts features from the input second feature map through a convolutional neural network and outputs a quadruple (Δx, Δy, Δw, Δh) at each position in the second feature map, which is the predicted offset relative to the anchor box; Based on the predicted offset and the initial coordinates of the anchor box, the coordinates of the predicted bounding box (x, y, w, h) are calculated (where x, y are the coordinates of the center point of the box, and w, h are the width and height of the box). The detection model outputs a confidence score at each anchor box location; The detection model uses the Softmax function to calculate the probability of each defect feature category.
9. The method according to claim 4, characterized in that The loss value is calculated by calculating the bounding box regression loss function, the target confidence loss function and the classification loss function, including: (1) in, , It represents the intersection-over-union ratio, which measures the overlap between the predicted box and the true box. The Euclidean distance between the center points of the predicted box and the true box; The minimum diagonal length of the bounding box between the predicted box and the real box; is the weight factor; is the difference in aspect ratio between the predicted box and the real box; (2) Among them, represents the target confidence of the prediction; represents the true target confidence (1 if the target is included, otherwise 0); BCE is the binary cross entropy, which can be expressed as: Where N is the total number of samples; represents the true label of the i-th sample; represents the predicted value of the i-th sample (model output); (3) Where C is the number of categories; is the predicted probability of the cth class; is the true c-th class label The overall loss function is defined as , and introduces a weight factor into each loss: , in, , , .
10. The method according to claim 1, characterized in that The method of removing redundant prediction frames based on the non-maximum suppression algorithm to obtain pipeline detection results includes: Set the confidence threshold to 0.3 and the intersection-over-union threshold to 0.25, filter out the prediction boxes whose target confidence is greater than the confidence threshold, generate candidate box masks, and remove the candidate boxes whose target confidence is lower than the confidence threshold; Fusion of target confidence and category confidence, the confidence of each category Equal to target confidence and category confidence The product of It can be expressed as: ; Sort candidate boxes by confidence from high to low, and keep the top 30k candidate boxes; Use IoU to remove redundant boxes; IoU (Intersection over Union) formula: Remove boxes with too much overlap and only keep boxes with IoU less than iou_thres; The maximum number of detection frames to be output. The number of detection frames shall not exceed 300.
Citation Information
Cited By
Rolling bearing failure mode identification method based on dynamic gradient correction mechanism
CN121190406A