Lightweight drainage pipeline defect detection method based on improved YOLOv8

By improving the YOLOv8 lightweight network model, combining GhostConv and CSPPC modules, the MLLA attention module and CIoU loss function are introduced, which solves the problem of large amount of parameters and insufficient detection stability in underwater pipeline defect detection, and achieves efficient and real-time pipeline defect detection.

CN120355997APending Publication Date: 2025-07-22QINGDAO PENGPAI OCEAN EXPLORATION TECH CO LTD

Patent Information

Application Number
CN202510460114.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-22

Smart Images

  • Figure CN120355997A_ABST
    Figure CN120355997A_ABST
Patent Text Reader

Abstract

The invention relates to the field of underwater pipeline defect detection, in particular to a lightweight drainage pipeline defect detection method based on improved YOLOv8. Comprising the following steps: S1, collecting a drainage pipeline defect data image, and preprocessing and enhancing the image; s2, inputting the underwater pipeline defect image obtained in the step S1 into an improved YOLOv8 lightweight network model, synchronously generating detection results of three scales, and finally outputting a three-dimensional tensor containing target frame coordinates, confidence and classification probability; and S3, carrying out training and hyper-parameter adjustment on the improved YOLOv8 lightweight network model, and evaluating the detection performance of the model. The provided model realizes the real-time reasoning speed of 336 FPS under the 4.7 MB ultra-small volume, the detection accuracy of a small target is improved, and pipeline defect detection in a complex underwater environment can meet the requirements of detection precision, detection speed and detection cost at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of underwater pipeline defect detection, and in particular to a lightweight drainage pipeline defect detection method based on improved YOLOv8. Background Art

[0002] In the field of urban underwater pipeline defect detection, traditional pipeline defect detection methods usually adopt manual inspection methods and physical detection methods. The method based on manual inspection is inefficient and subjective, and the physical detection methods based on magnetic flux leakage / ultrasound are vulnerable to environmental interference and have limited applicable scenarios. Traditional pipeline defect detection methods are inefficient, costly, and vulnerable to environmental interference. In addition, there are problems such as low clarity and difficulty in feature extraction due to small defect volume, and they have gradually been replaced by object detection technologies based on deep learning.

[0003] The YOLO series of algorithms has become a research hotspot in the field of underwater pipeline defect detection due to the real-time advantage of its single-stage detection architecture. Among them, YOLOv5s, as a classic model, adopts the CSPDarknet53 backbone network and the PAN-FPN feature pyramid structure, reduces computational redundancy through the cross-stage partial connection (CSP) module, and expands the receptive field with the spatial pyramid pooling fast (SPPF), achieving an average precision of 36.7% on the COCO dataset. However, this model has significant defects: its parameter quantity of 18.5MB and computational complexity of 23.8 GFLOPs make it difficult to deploy on edge devices. The actual detection shows that the missed detection rate of tiny defects (<32×32 pixels) on the self-built underwater pipeline dataset is as high as 27.1%, and the standard convolutional layer accounts for 58.3% of the total parameter quantity, resulting in low feature extraction efficiency.

[0004] YOLOv10 released in 2024 innovates for lightweight requirements, adopts an anchor-free detection mechanism to eliminate the calculation of preset anchor points, replaces the standard convolutional layer with depthwise separable convolutions, compresses the model volume to 5.38MB (a decrease of 70.9%), reduces the computational amount to 8.2 GFLOPs, and introduces a dynamic sparse attention (DSA) mechanism to generate a spatial mask with 1×1 convolutions to suppress background noise. However, this solution exposes obvious shortcomings in practical applications: its attention module depends on accurate position prediction. In scenarios where the image is blurred due to reflection on the inner wall of the pipeline and interference from suspended matter, the position prediction error causes the attention weight to fail, and the actual detection standard deviation in a turbid water environment reaches ±2.3%; at the same time, the inference speed of depthwise separable convolutions on edge devices such as Jetson Nano only increases by 18.3%, and the high-precision detection ability (mAP95 = 0.561) of the model on the self-built dataset is significantly weaker than that of traditional models.

[0005] In summary, the existing deep learning models generally face the following problems when completing underwater pipeline defect detection: Traditional models such as YOLOv5s have high detection accuracy, but the huge number of parameters (9.11 million) and computing requirements are difficult to meet the deployment of mobile robot embedded platforms; Although new lightweight solutions such as YOLOv10 reduce resource consumption, due to the design defects of the attention mechanism, the detection stability in complex environments is insufficient. In particular, the problems of low contrast features and loss of small target spatial information in underwater images have not been effectively solved. Summary of the Invention

[0006] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art, and propose a lightweight drainage pipeline defect detection method based on improved YOLOv8. The model proposed by this method achieves a real-time inference speed of 336 FPS with an ultra-small volume of 4.7MB, improves the detection accuracy of small targets, and enables the pipeline defect detection in complex underwater environments to meet the requirements of detection accuracy, detection speed, and detection cost at the same time.

[0007] The technical solution of the present invention is: A lightweight drainage pipeline defect detection method based on improved YOLOv8, which includes the following steps: S1. Collect defect data images of drainage pipelines, and preprocess and enhance the images; S2. Input the underwater pipeline defect images obtained in step S1 into the improved YOLOv8 lightweight network model, synchronously generate detection results at three scales, and finally output a three-dimensional tensor containing the target box coordinates, confidence, and classification probability; S3. Train the improved YOLOv8 lightweight network model and adjust the hyperparameters, and evaluate the detection performance of the model.

[0008] In the present invention, in step S1, the Labelme software is used to annotate and classify the collected defect pictures, and the obtained data set is subjected to image enhancement. The enhanced data is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1.

[0009] The improved YOLOv8 lightweight network model includes a backbone network, a neck network, and a detection head; The backbone network includes a preliminary feature extraction part and a four-stage downsampling structure. The preliminary feature extraction part includes a GhostConv module, and the four-stage downsampling structure includes a first-stage downsampling structure, a second-stage downsampling structure, a third-stage downsampling structure, and a fourth downsampling structure connected in series in sequence; The neck network includes an upsampling part and a downsampling part. The feature maps output by the backbone network are concatenated twice in the upsampling part. The output of the first upsampling stage of the neck network and the output of the last layer of the backbone network are concatenated in the downsampling part respectively. Through this structure, features of different scales are fused to improve the detection accuracy. The detection head part includes multiple output modules of different scales, and the front end of the detection head with the smallest scale integrates the MLLA attention module.

[0010] The first-stage downsampling structure includes a second GhostConv module and a first CSPPC module connected in series in sequence, and outputs a feature map of 160×160×128. The second-stage downsampling structure includes a third GhostConv module and a second CSPPC module connected in series in sequence, and outputs a feature map of 80×80×256. The third-stage downsampling structure includes a fourth GhostConv module and a third CSPPC module connected in series in sequence, and outputs a feature map of 40×40×512. The fourth-stage downsampling structure includes a fifth GhostConv module, a fourth CSPPC module and a first SPPF module connected in series in sequence, and outputs a feature map of 20×20×512.

[0011] The upsampling part includes a first UpSample module, a first Concat module, a fifth CSPPC module, a second UpSample module, a second Concat module, and a sixth CSPPC module connected in series in sequence. The 20×20×512 feature map output by the fourth-stage downsampling structure is input into the first UpSample module for bilinear interpolation, and the first UpSample module outputs a feature map of 40×40×512. In the first Concat module, the 40×40×512 feature map output by the first UpSample module and the 40×40×512 feature map output by the third-stage downsampling structure are concatenated. The feature map output by the first Concat module passes through the fifth CSPPC module and the second UpSample module in sequence, and the resolution is restored to 80×80×256. In the second Concat module, the 80×80×256 feature map output by the second UpSample module is concatenated with the 80×80×256 feature map output by the second-stage downsampling structure, and is input into the sixth CSPPC module. Cross-scale feature fusion is performed through the sixth CSPPC module, and a feature map of 80×80×256 is output. The feature map output by the sixth CSPPC module is input into the detection head and the downsampling part at the same time.

[0012] The downsampling part includes a sixth GhostConv module, a third Concat module, a seventh CSPPC module, a seventh GhostConv module, a fourth Concat module, and an eighth CSPPC module connected in sequence; The 40×40×256 feature map obtained after processing the feature map output by the upsampling part through the sixth GhostConv module is concatenated with the output of the fifth CSPPC module in the third Concat module and input into the seventh CSPPC module. After being processed by the seventh CSPPC module, the feature map is compressed into a feature map with a resolution of 40×40×512 and input into the detection head; At the same time, the 40×40×512 feature map passes through the seventh GhostConv module and the fourth Concat module in sequence, and in the fourth Concat module, it is feature concatenated with the 20×20×512 feature map output by the backbone network, and then input into the eighth CSPPC module for processing, and a 20×20×512 feature map is output and input into the detection head.

[0013] Multi-object joint optimization and output are implemented in the detection head. The CIoU loss function is used to optimize the positioning accuracy, and joint constraints are imposed on the distance between the center points of the prediction boxes, the aspect ratio, and the overlapping area; the Focal Loss is combined to solve the problem of class imbalance; the output layer synchronously generates detection results at three scales and finally outputs the coordinates of the target boxes and the classification probabilities; The large-scale feature map is used to detect the overall structural anomalies of the pipeline; the medium-scale feature map is used to identify medium defects; the small-scale feature map is used to identify small defects, and real-time detection is achieved through weighted fusion.

[0014] In step S3, the accuracy evaluation formula of the improved YOLOv8 lightweight network model is: ; ; ; ; In the formula, TP represents that the prediction is a positive class and the real sample is a positive class; FP represents that the prediction is a positive class and the real sample is a negative class; FN represents that the prediction result is a negative class and the real sample is a positive class; Precision represents the number of correctly predicted positive class samples / the total number of all samples predicted as positive classes; Recall represents the number of correctly predicted positive class samples / the total number of all positive class samples; AP represents the area enclosed by the P-R curve and the coordinate axes; mAP represents the average recognition accuracy mean of all categories.

[0015] The beneficial effects of the present invention are: (1) The backbone network is reconstructed using the CSPPC module. By leveraging the redundancy characteristics of feature maps, the GFLOPs are reduced by 21.67%. Combining with GhostConv lightweight convolution, while reducing the number of parameters by 31.61%, the computational efficiency is increased by 21.67%, reducing the model size to 4.7MB, meeting the real-time detection requirements of embedded devices. (2) By introducing the MLLA attention module at the front end of the small-scale detection head, the spatial perception of defects is enhanced through position encoding, while maintaining the parallel computing efficiency, the recall rate of small targets is increased by 5.14%. (3) Through the multi-module joint optimization strategy, on the basis of lightweight, the detection accuracy of mAP50 0.759 is maintained, which is 3.97% higher than that of the similar lightweight model YOLOv10. Through the improved YOLOv8 lightweight network model constructed in this application, in the typical defect detection scenario of municipal pipelines, a coordinated improvement of lightweight is achieved with an accuracy of 0.888, an efficiency of 336 FPS, and a model size reduced to 4.7MB, providing a reliable technical path for infrastructure safety monitoring in complex environments. Description of the Drawings

[0016] Figure 1 It is a schematic diagram of a pipeline defect being a hole; Figure 2 It is a schematic diagram of a pipeline defect being a crack; Figure 3 It is a schematic diagram of a pipeline defect being a protrusion; Figure 4 It is a schematic diagram of a pipeline defect being an offset; Figure 5 It is a schematic diagram of the improved YOLOv8 lightweight network model proposed in this application; Figure 6 It is a structural diagram of the GhostConv module; Figure 7 It is a structural diagram of the CSPPC module; Figure 8 It is a schematic diagram of some results output by the method described in the present invention. Detailed Embodiments

[0017] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the accompanying drawings.

[0018] In the following description, specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0019] A lightweight drainage pipeline defect detection method based on improved YOLOv8 according to the present invention includes the following steps.

[0020] In the first step, collect drainage pipeline defect data images, preprocess the images, perform data augmentation on the collected images, including data augmentation techniques such as image transformation, scaling, left - right flipping, and mosaic, and divide the augmented image data into a training set, a validation set, and a test set.

[0021] In this embodiment, the collection of drainage pipeline defect data images is completed by an underwater mobile robot of Qingdao Pengpai Ocean Exploration Technology Co., Ltd. in the sewage pipeline. The robot walks in the pipeline to take pipeline defect images. Pipeline defects usually cover four categories, namely holes, cracks, protrusions, and misalignments, as Figures 1 to 4 shown.

[0022] Use Labelme software to annotate and classify the collected defect pictures, and finally perform image augmentation on the obtained dataset. Image augmentation uses means such as mosaic augmentation (probability 1.0), left - right flipping (probability 0.5), and random scaling (0.5 - 1.5 times) to enrich the number of defect images. In this embodiment, a total of 10,000 pictures are collected, and the augmented data is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0023] In the second step, input the underwater pipeline defect images obtained in the first step into the improved YOLOv8 lightweight network model, and the model can synchronously generate detection results at three scales.

[0024] As Figure 5 shown, the improved YOLOv8 lightweight network model in this application includes a backbone network, a neck network, and a detection head.

[0025] The backbone network includes a preliminary feature extraction part and a four - level downsampling structure. The preliminary feature extraction part is used to perform preliminary feature extraction, and feature extraction is performed through the four - level downsampling structure to obtain features at different scales, facilitating subsequent network for feature fusion at different scales.

[0026] The preliminary feature extraction part includes a first GhostConv module. Input the underwater pipeline defect images, and perform preliminary feature extraction through the first GhostConv module.

[0027] Specifically, an underwater pipeline defect image with an input resolution of 640×640×3 is subjected to preliminary feature extraction through the first GhostConv module. The first GhostConv module includes the first layer of Ghost, the second layer of Ghost, the third layer of Ghost, and an Add operation layer. In the input stage of the first GhostConv module, the number of channels is adjusted through convolution. In this embodiment, the input adjusts the number of channels through a 1×1 convolution, maps 3 channels to 64 channels, and obtains the initial residual path feature of 640×640×64.

[0028] The data processing of the first layer of Ghost all includes core feature generation, redundant feature generation, feature concatenation, batch normalization processing, and ReLU activation. In the core feature generation process, the input channels are compressed through a conventional convolution. In this embodiment, the input channels are compressed from 64 to 32, and 640×640×32 is output. In the redundant feature generation process, depthwise separable convolution is performed on the channel features to generate complementary channel features. In this embodiment, 5×5 depthwise separable convolution is performed on the 32-channel features to generate complementary 32-channel features, and 640×640×32 is output. In the feature concatenation process, the core features and redundant features are concatenated. In this embodiment, 640×640×64 is obtained after concatenation. In the batch normalization process, the concatenated features are normalized. In this embodiment, the dimension 640×640×64 is maintained. In the ReLU activation process, non-linearity is introduced. The output in this embodiment is still 640×640×64.

[0029] The data processing of the second layer of Ghost and the third layer of Ghost all includes core feature generation, redundant feature generation, feature concatenation, and batch normalization processing. The processing process is the same as that of the first layer of Ghost and will not be elaborated here.

[0030] The Add operation layer performs a path merging operation, adds the output of the third layer of Ghost to the initial residual path element by element, fuses shallow details and deep features, and finally outputs.

[0031] The GhostConv module uses a 1×1 convolution kernel to generate an initial feature map. Subsequently, depthwise separable convolution with a 3×3 kernel is performed on 20% of the channels, and the remaining 80% of the channels generate "ghost features" through a cheap linear transformation, expanding the original 3 channels to 64 channels, while reducing the number of parameters and retaining key spatial information. The size of the feature map output by the GhostConv module is 320×320×64, providing high-resolution shallow features for subsequent processing.

[0032] The four-level downsampling structure includes a first-level downsampling structure, a second-level downsampling structure, a third-level downsampling structure, and a fourth-level downsampling structure connected in series in sequence.

[0033] Among them, the first - level downsampling structure includes a second GhostConv module and a first CSPPC module connected in series in sequence. In this embodiment, the first - level downsampling structure downsamples the 320×320×64 feature map output by the preliminary feature extraction part and outputs a 160×160×128 feature map.

[0034] The second - level downsampling structure includes a third GhostConv module and a second CSPPC module connected in series in sequence. In this embodiment, the second - level downsampling structure uses partial convolution on the feature map output by the first - level downsampling structure, processes 25% of the input channels, and outputs an 80×80×256 feature map.

[0035] The third - level downsampling structure includes a fourth GhostConv module and a third CSPPC module connected in series in sequence. In this embodiment, the third - level downsampling structure combines skip connections with the feature map output by the second - level downsampling structure to fuse features of different levels and outputs a 40×40×512 feature map.

[0036] The fourth - level downsampling structure includes a fifth GhostConv module, a fourth CSPPC module, and a first SPPF module connected in series in sequence, and outputs a 20×20×512 feature map.

[0037] Inside each CSPPC module, grouped convolution and channel shuffle strategies are adopted to reduce the number of parameters while maintaining the feature expression ability, and to focus on enhancing the edge feature response of small targets such as pipeline cracks and holes.

[0038] The structure of the CSPPC module is as Figure 7 shown. Its data - processing flow is as follows: first convolution, the number of channels is converted to C_out; channel equalization, the feature map is evenly divided along the channel dimension into a direct channel and a processing channel; cyclic feature refinement, performing N point - wise convolutions on the processing channel; multi - scale feature fusion, concatenating the original direct channel, the cyclic initial channel, and the results of N cycles in the channel dimension, and finally the channel dimension is 0.5C_out×(N + 2); feature integration output, compressing the expanded 0.5C_out×(N + 2) channels back to C_out channels through the final convolutional layer.

[0039] On the basis of maintaining the dimensional compatibility of the input and output channels, this module adopts a dual-branch processing structure: in the PartialConv branch, 3×3 convolution operations are only performed on 25% of the input channels to extract spatial features, and the remaining 75% of the channels retain the original information through identity mapping. Finally, feature fusion is achieved through channel concatenation. This design of "sparse calculation + feature inheritance" combines the gradient shunting advantage of the CSP structure and the redundant feature compression ability of PartialConv. It not only reduces the FLOPs to 1 / 16 of the original module (when k = 0.25), but also effectively suppresses the interference information in underwater blurred images by focusing on the feature extraction of the core channels, enabling the improved model to significantly reduce the model size, the number of model parameters, GFLOPs, and improve the inference speed while maintaining accuracy.

[0040] The neck network includes an upsampling part and a downsampling part. The feature map output by the fourth-level downsampling structure is input into the upsampling part. After the upsampling part processes the feature map, it outputs a large-scale feature map to the detection head. At the same time, the feature map output by the upsampling part is input into the downsampling part. After the downsampling part processes the feature map, it can respectively obtain medium-scale and small-scale feature maps and output them to the detection head.

[0041] The upsampling part includes a first UpSample module, a first Concat module, a fifth CSPPC module, a second UpSample module, a second Concat module, and a sixth CSPPC module connected in series in sequence.

[0042] The 20×20×512 feature map output by the fourth-level downsampling structure is input into the first UpSample module. After bilinear interpolation by the first UpSample module, a 40×40×512 feature map is output.

[0043] The output of the third CSPPC module is connected to the input of the first Concat module. That is to say, in the first Concat module, the 40×40×512 feature map output by the first UpSample module and the 40×40×512 feature map output by the third CSPPC module are concatenated.

[0044] The feature map output by the first Concat module passes through the fifth CSPPC module and the second UpSample module in sequence, and the resolution is restored to 80×80×512. Then, it is concatenated with the 80×80×256 feature map output by the second-level downsampling structure of the backbone network in the second Concat module, and finally input into the sixth CSPPC module. Through cross-scale feature fusion by the sixth CSPPC module, an 80×80×256 feature map is output. The feature map output by the sixth CSPPC module is input into the detection head and the downsampling part at the same time.

[0045] The downsampling part includes a sixth GhostConv module, a third Concat module, a seventh CSPPC module, a seventh GhostConv module, a fourth Concat module, and an eighth CSPPC module connected in sequence.

[0046] The sixth GhostConv module uses a GhostConv module with a stride of 2. The feature map output by the sixth CSPPC module is concatenated with the 40×40×512 feature map from the fifth GhostConv module in the upper sampling part in the third Concat module and input into the seventh CSPPC module. After being processed by the seventh CSPPC module, the feature map is compressed into a 40×40×512 feature map, and the 40×40×512 feature map is input into the detection head.

[0047] Meanwhile, the 40×40×512 feature map passes through the seventh GhostConv module and the fourth Concat module in sequence, and is concatenated with the 20×20×512 feature map from the SPPF module in the backbone network, and then input into the eighth CSPPC module for processing, and is compressed into a 20×20×512 feature map; after being processed in the eighth CSPPC module, a 20×20×512 feature map is output and input into the detection head.

[0048] The detection head includes multiple different-scale output modules. Through the multiple-scale output modules, feature maps of three scales of 80×80, 40×40, and 20×20 are retained. The feature maps of the above three scales are used to detect pipeline defect targets of different sizes, and the smallest-scale feature map is specifically used to capture tiny defects of 20×20 pixels.

[0049] Specifically, the large-scale feature map with a resolution of 80×80 is used to detect the overall structural abnormality of the pipeline; the medium-scale feature map with a resolution of 40×40 is used to identify medium defects such as corrosion and misalignment; the small-scale feature map with a resolution of 20×20 is used to focus on fine defects such as crack tips and micropores, and finally real-time detection is achieved through weighted fusion.

[0050] The front end of the smallest-scale output module is embedded with an MLLA (Mamba-Like Linear Attention) attention module.

[0051] The MLLA attention module realizes feature enhancement through three-stage reconstruction on the premise of maintaining the consistency of input and output dimensions: first, RoPE (Rotary Position Embedding) is used to perform rotational position encoding on the input features to establish the relative position relationship between pixels in the frequency domain space; then, the global context correlation is modeled through the linear attention mechanism, and its computational complexity is from the traditional attention of O (N 2 ) reduced to O ( N ); Finally, gated convolution is introduced to achieve local feature compensation, forming a "global-local" collaborative feature optimization mechanism.

[0052] The MLLA attention module improves the detection accuracy by 2.42% with only a 1.3% increase in the number of parameters through linear attention reconstruction and position encoding optimization, breaking through the accuracy-speed trade-off bottleneck of lightweight models and enhancing the feature extraction ability in complex environments.

[0053] Joint optimization and output of multiple targets. The CIoU loss function is used to optimize the positioning accuracy, and joint constraints are imposed on the distance between the center points of the predicted boxes, the aspect ratio, and the overlapping area. Combining Focal Loss to solve the problem of class imbalance, the loss weight of difficult samples is increased by 4.2 times. The output layer synchronously generates detection results at three scales, and the final output is a three-dimensional tensor containing the target box coordinates, confidence, and classification probability, as Figure 8 shown.

[0054] In the third step, the improved YOLOv8 lightweight network model is trained and hyperparameters are adjusted. After determining the optimal parameters, the detection performance of the model is evaluated using the test set.

[0055] The constructed improved YOLOv8 lightweight network model is evaluated. Training and hyperparameter adjustment are carried out. After determining the optimal parameters, the performance of the model is tested using the test set. The main evaluation metrics include mean average precision (mAP), precision, recall, floating point operations (GFLOPs), number of parameters (Params), and model size.

[0056] Before training the constructed model, the hyperparameters are selected as AdamW, the initial learning rate is 0.01, the momentum is 0.937, the BatchSize is set to 16, and the iterative training is carried out for 150 rounds. The training set is input into the improved YOLOv8 lightweight network model for training, and the best weight file is obtained. Then, the prediction results of the improved YOLOv8 lightweight network model are evaluated.

[0057] mAP, Precision, Recall, GFLOPs, Params, and model size are used as the evaluation criteria for the improved network model. Among them, mAP represents the average recognition accuracy of the ratio of the intersection to the union of the predicted box and the ground truth box at a certain threshold. mAP comprehensively analyzes the precision and recall rate to evaluate the model performance more comprehensively. In addition, the number of parameters, computational volume, and model size of the model all affect the model in real-time detection.

[0058] The evaluation formula of the model is as follows: ; ; ; ; Wherein, TP represents that the prediction is a positive class and the true sample is a positive class; FP represents that the prediction is a positive class and the true sample is a negative class; FN represents that the prediction result is a negative class and the true sample is a positive class; Precision represents the number of correctly predicted positive class samples / the total number of all samples predicted as positive classes; Recall represents the number of correctly predicted positive class samples / the total number of all positive class samples; AP represents the area enclosed by the P-R curve and the coordinate axes; mAP represents the average recognition accuracy mean value of all categories.

[0059] According to different evaluation criteria (precision, recall rate, mAP50, mAP95, GFLOPS, model size, and FPS), the evaluation results are shown in Table 1. Table 1 outlines the impacts of different methods: the baseline model (A), the introduced MLLA attention mechanism (M), the introduced CSPPC module proposed based on PartialConv (C), and the GhostConv module (G). Table 1 defines the baseline model A, the improved models A+M, A+M+C, and A+M+C+G in sequence, and quantitatively discusses the changes in five evaluation metrics in these models.

[0060] Table 1. Impacts of Different Modules on Evaluation Criteria Model Precision Recall mAP50 mAP95 GFLOPS Model Size Parameters FPS(Task / s) A 0.867 0.719 0.770 0.581 12.0 6.8 MB 3258844 329 A+M 0.841 0.756 0.785 0.585 12.1 7.1 MB 3392988 318 A+M+C 0.876 0.710 0.765 0.556 9.90 5.3 MB 2508252 337 A+M+C+G 0.888 0.685 0.759 0.548 9.40 4.7 MB 2228636 336 .

[0061] In the ablation experiment, the baseline model A has certain performance. Its precision is 0.867, recall is 0.719, mAP50 is 0.770, mAP95 is 0.581, GFLOPS is 12.0, model size is 6.8MB, number of parameters is 3,258,844, and FPS is 329. When the MLLA attention mechanism (A + M) is introduced, the precision slightly drops to 0.841, the recall increases to 0.756, mAP50 increases to 0.785, mAP95 increases to 0.585, GFLOPS slightly increases to 12.1, the model size increases to 7.1MB, the number of parameters increases to 3,392,988, and FPS drops to 318. This indicates that although MLLA improves the recall and mAP metrics, it increases the computational resource overhead. Then the CSPPC module (A + M + C) is introduced. The precision increases to 0.876, the recall drops to 0.710, mAP50 drops to 0.765, mAP95 drops to 0.556, GFLOPS significantly drops to 9.90, the model size significantly decreases to 5.3MB, the number of parameters decreases to 2,508,252, and FPS increases to 337. This shows that the CSPPC module achieves lightweight and improves the processing speed while maintaining a high precision. Finally, the GhostConv module (A + M + C + G) is added. The precision further increases to 0.888, the recall drops to 0.685, mAP50 drops to 0.759, mAP95 drops to 0.548, GFLOPS drops to 9.40, the model size decreases to 4.7MB, the number of parameters decreases to 2,228,636, and FPS slightly drops to 336. It can be seen that the GhostConv module further optimizes the feature extraction ability and further reduces the model size and resource occupancy on the premise of maintaining the precision. Generally speaking, each improvement measure enables the model to achieve a good balance in precision, lightweight and real-time performance, and can better meet the requirements of municipal pipeline defect detection.

[0062] Furthermore, the improved YOLOv8 lightweight network model proposed in this application is compared with other models, and the results of detecting the drainage pipeline dataset are shown in Table 2. In Table 2, the improved model proposed in this application is defined as MCSG-YOLOv8.

[0063] Table 2 Parameter comparison between the improved YOLOv8 lightweight network model described in this application and existing models Model Precision Recall mAP50 mAP95 GFLOPS Model Size Parameters YOLOv5s 0.882 0.739 0.788 0.609 23.8 18.5 MB 9113084 YOLOv8s 0.899 0.728 0.791 0.608 42.4 23.9 MB 11781148 YOLOv10 0.826 0.670 0.730 0.561 8.20 5.38MB 2695976 MCSG-YOLOv8 0.888 0.685 0.759 0.548 9.40 4.7 MB 2228636 。

[0064] In terms of model performance comparison, YOLOv8s performs best in terms of accuracy, with an accuracy value of 0.899, indicating a high proportion of true positive samples among the results predicted as positive samples and a low false positive rate; the accuracy of MCSG-YOLOv8 is 0.888, which is close to 0.882 of YOLOv5s and at a relatively high level. These two models perform well in accurately identifying positive samples, while the accuracy of YOLOv10 is relatively low at 0.826, and its ability to judge positive samples is slightly worse. In terms of recall rate, YOLOv5s is the highest at 0.739, which can better detect all actual positive samples and has a low false negative rate; the recall rates of MCSG-YOLOv8 and YOLOv8s are 0.685 and 0.728 respectively, which are relatively close and moderate, and there is room for improvement in the integrity of detecting positive samples. The recall rate of YOLOv10 is the lowest at 0.670, and its detection ability is weak. For mAP50 and mAP95, YOLOv8s performs outstandingly in both metrics. The mAP50 is 0.791 and the mAP95 is 0.608, with a high average detection accuracy under different IOU thresholds and an obvious advantage under high IOU thresholds; MCSG-YOLOv8 and YOLOv5s are close and perform well in mAP50. The mAP95 of MCSG-YOLOv8 is 0.548, which is lower than 0.609 of YOLOv5s, and there is room for improvement in the detection accuracy of MCSG-YOLOv8 under high IOU thresholds. YOLOv10 is relatively low in these two metrics. In terms of GFLOPS, YOLOv5s and YOLOv8s have high computational resource requirements, which are 23.8 and 42.4 respectively. MCSG-YOLOv8 and YOLOv10 have low requirements, which are 9.40 and 8.20 respectively. The latter two can operate efficiently and have good deployability in environments with limited resources. In terms of model size, MCSG-YOLOv8 is the smallest at only 4.7MB, suitable for resource-constrained scenarios. YOLOv10 is 5.38MB, followed by YOLOv5s and YOLOv8s, which are relatively large at 18.5MB and 23.9MB respectively. In terms of the number of parameters, YOLOv8s has the largest number of parameters, reaching 11,781,148, with a long training time, easy overfitting, and slow inference speed. YOLOv5s comes second. MCSG-YOLOv8 and YOLOv10 have fewer parameters, which are 2,228,636 and 2,695,976 respectively. The models are lightweight and have relatively high training and inference efficiency.

[0065] The improved YOLOv8 lightweight network model proposed by the present invention is close to and has high accuracy compared with YOLOv5s, has a moderate recall rate, has acceptable mAP50 performance but there is room for improvement in mAP95. It has significant advantages in terms of GFLOPS, model size, and number of parameters, with low computational resource requirements, small model size, and few parameters. It is very suitable for the municipal pipeline defect detection task in resource-constrained environments and can achieve fast and efficient detection while ensuring a certain detection accuracy. Although YOLOv8s performs well in terms of accuracy and mAP metrics, it has high computational resource requirements, a large model, and many parameters, and is suitable for scenarios with sufficient computational resources and extremely high accuracy requirements. YOLOv5s has good performance in terms of recall rate and mAP metrics, but also has problems such as a relatively large model and high computational resource requirements. YOLOv10 is relatively weak in various metrics and only has certain advantages in terms of model size and computational resource requirements, and its overall performance is relatively poor compared to the other three models. Some of the results output by the method described in this application are as Figure 8 shown.

[0066] The above results show that the improved YOLOv8 lightweight network model proposed in this application can reduce the computational amount and the number of model parameters while basically maintaining the recognition accuracy, can meet the requirements of real-time detection, and at the same time has stronger deployment capabilities, and can provide a better basis for the drainage pipeline defect detection work.

[0067] The above has introduced in detail the lightweight drainage pipeline defect detection method based on the improved YOLOv8 provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown in this article, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A lightweight drainage pipeline defect detection method based on improved YOLOv8, characterized in that, It includes the following steps: S1. Collect defect data images of drainage pipes, and preprocess and enhance the images; S2. Input the underwater pipe defect images obtained in step S1 into the improved YOLOv8 lightweight network model, synchronously generate detection results at three scales, and finally output a three-dimensional tensor containing the coordinates of the target box, confidence, and classification probability; S3. Train the improved YOLOv8 lightweight network model and adjust hyperparameters, and evaluate the detection performance of the model.

2. The lightweight drainage pipeline defect detection method based on the improved YOLOv8 according to claim 1, characterized in that In step S1, the Labelme software is used to label and classify the collected defect pictures, and the obtained dataset is image-enhanced. The enhanced data is divided into a training set, a validation set, and a test set according to the ratio of 8:1:

1.

3. The lightweight drainage pipeline defect detection method based on the improved YOLOv8 according to claim 1, characterized in that The improved YOLOv8 lightweight network model includes a backbone network, a neck network, and a detection head; The backbone network includes a preliminary feature extraction part and a four-stage downsampling structure. The preliminary feature extraction part includes a GhostConv module, and the four-stage downsampling structure includes a first-stage downsampling structure, a second-stage downsampling structure, a third-stage downsampling structure, and a fourth-stage downsampling structure connected in series in sequence; The neck network includes an upsampling part and a downsampling part. The feature maps output by the backbone network are spliced twice in the upsampling part; the output of the first upsampling stage of the neck network and the output of the last layer of the backbone network are respectively spliced in the downsampling part; The detection head part includes multiple output modules of different scales, and the front end of the smallest-scale detection head integrates an MLLA attention module.

4. The lightweight drainage pipeline defect detection method based on improved YOLOv8 according to claim 3, characterized in that The first-stage downsampling structure includes a second GhostConv module and a first CSPPC module connected in series in sequence, and outputs a 160×160×128 feature map; The second-stage downsampling structure includes a third GhostConv module and a second CSPPC module connected in series in sequence, and outputs an 80×80×256 feature map; The third-stage downsampling structure includes a fourth GhostConv module and a third CSPPC module connected in series in sequence, and outputs a 40×40×512 feature map; The fourth-stage downsampling structure includes a fifth GhostConv module, a fourth CSPPC module, and a first SPPF module connected in series in sequence, and outputs a 20×20×512 feature map.

5. The lightweight drainage pipeline defect detection method based on the improved YOLOv8 according to claim 3, characterized in that, The upsampling part includes a first UpSample module, a first Concat module, a fifth CSPPC module, a second UpSample module, a second Concat module, and a sixth CSPPC module connected in series in sequence; The 20×20×512 feature map output by the fourth-stage downsampling structure is input into the first UpSample module for bilinear interpolation, and the first UpSample module outputs a 40×40×512 feature map; In the first Concat module, the 40×40×512 feature map output by the first UpSample module and the 40×40×512 feature map output by the third-stage downsampling structure are spliced; The feature map output by the first Concat module passes through the fifth CSPPC module and the second UpSample module in sequence, and the resolution is restored to 80×80×256; In the second Concat module, the 80×80×256 feature map output by the second UpSample module is concatenated with the 80×80×256 feature map output by the second-level downsampling structure and input into the sixth CSPPC module; Cross-scale feature fusion is performed through the sixth CSPPC module to output an 80×80×256 feature map. The feature map output by the sixth CSPPC module is input into the detection head and the downsampling part at the same time.

6. The lightweight drainage pipeline defect detection method based on improved YOLOv8 according to claim 4, wherein The downsampling part includes a sixth GhostConv module, a third Concat module, a seventh CSPPC module, a seventh GhostConv module, a fourth Concat module, and an eighth CSPPC module connected in sequence; The 40×40×256 feature map obtained by processing the feature map output by the upsampling part through the sixth GhostConv module is concatenated with the output of the fifth CSPPC module in the third Concat module and input into the seventh CSPPC module. After being processed by the seventh CSPPC module, the feature map is compressed into a feature map with a resolution of 40×40×512 and input into the detection head; At the same time, the 40×40×512 feature map passes through the seventh GhostConv module and the fourth Concat module in sequence, and in the fourth Concat module, it is feature concatenated with the 20×20×512 feature map output by the backbone network, and then input into the eighth CSPPC module for processing, and an output of 20×20×512 feature map is input into the detection head.

7. The lightweight drainage pipeline defect detection method based on the improved YOLOv8 according to claim 3, characterized in that, Multi-object joint optimization and output are implemented in the detection head. The CIoU loss function is used to optimize the positioning accuracy, and the distance between the center points of the prediction boxes, the aspect ratio, and the overlapping area are jointly constrained; the Focal Loss is combined to solve the problem of class imbalance; the output layer synchronously generates detection results at three scales, and finally outputs the target box coordinates, confidence, and classification probability; The large-scale feature map is used to detect the overall structural abnormality of the pipeline; The medium-scale feature map is used to identify medium defects; the small-scale feature map is used to identify small defects, and real-time detection is achieved through weighted fusion.

8. The lightweight drainage pipeline defect detection method based on the improved YOLOv8 according to claim 1, characterized in that, In step S3, the accuracy evaluation formula of the improved YOLOv8 lightweight network model is: ; ; ; ; In the formula, TP represents that the prediction is a positive class and the real sample is a positive class; FP represents that the prediction is a positive class and the real sample is a negative class; FN represents that the prediction result is a negative class and the real sample is a positive class; Precision represents the number of correctly predicted positive class samples / the total number of all samples predicted as positive classes; Recall represents the number of correctly predicted positive class samples / the total number of all positive class samples; AP represents the area enclosed by the P-R curve and the coordinate axis; mAP represents the average recognition accuracy mean of all categories.

Citation Information

Patent Citations

  • Blind person travel auxiliary equipment based on deep learning

    CN118942065A

  • High-voltage transmission line insulator defect detection method based on LightWeight-YOLOv8n

    CN119027367A

  • Drainage pipeline defect detection method based on improved YOLOv8s

    CN119169378A

Cited By

  • Underwater target detection method and system, storage medium and equipment

    CN121527607A

  • Target detection method, device and equipment for chemical experimental instrument and storage medium

    CN121582892A

  • Deep learning-based drainage pipeline defect detection method and system

    CN122265159A