Sewage pipeline defect detection method and system based on improved YOLO11 network

By improving the YOLO11 network, using the MSGA module, AIFI module and Wise-Inner-IoU loss function, the difficulty of detecting small objects and complex background interference in the detection of sewer pipeline defects is solved, achieving higher detection accuracy and better detection effect.

CN120125568APending Publication Date: 2025-06-10JINLING INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510308795.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

YOLO11n faces problems such as difficulty in detecting small objects, complex background interference and data set imbalance in sewer pipeline defect detection.

Method used

By improving the YOLO11 network, the MSGA module, AIFI module and Wise-Inner-IoU loss function are adopted to enhance the model's ability to capture information on target structures in complex contexts and sensitivity to defects.

Benefits of technology

On the basis of maintaining the model's lightweight architecture and not increasing the amount of parameters, the detection accuracy is improved, the detection effect is achieved, and the huge deployment potential is demonstrated on real-time detection edge devices such as pipeline robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125568A_ABST
    Figure CN120125568A_ABST
Patent Text Reader

Abstract

The invention discloses a sewage pipeline defect detection method and system based on an improved YOLO11 network. The sewage pipeline defect detection method comprises the following steps that a pipeline CCTV robot is used for collecting a defect image sample in a sewage pipeline; the YOLO11n basic network is improved by using an MSGA module, an AIFI module and a positioning loss function Wise-Inner-IoU so as to construct a sewage pipeline defect model; inputting the acquired defect image sample in the sewage pipeline into a sewage pipeline defect model for training, and performing data enhancement on the acquired defect image sample in real time by adopting a Python library Albumentations in the training process; and inputting the defect target image in the to-be-detected sewage pipeline into the trained sewage pipeline defect model to obtain a defect detection result in the to-be-detected sewage pipeline. According to the method, the detection precision is improved, the capturing capability of target structure information under a complex background and the sensitivity to defects are enhanced, and a better detection effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of sewage pipeline defect detection, and particularly relates to a sewage pipeline defect detection method and system based on an improved YOLO11 network. Background Art

[0002] The detection of sewer pipeline defects is a key task in urban infrastructure management, involving the automatic detection and evaluation of problems such as cracks, corrosion, and blockages inside the pipeline. Traditional methods rely on manual inspection or rules based on image processing, but these methods often have limitations such as long time consumption, poor accuracy, and being greatly affected by the environment.

[0003] With the development of deep learning technology, methods based on convolutional neural networks (CNNs), especially the YOLO (You Only Look Once) series of models, have achieved remarkable success in object detection tasks, and are particularly suitable for sewer pipeline defect detection. As the latest model in the YOLO series, YOLO11n is optimized for efficient object detection and real-time application scenarios and has multiple advantages.

[0004] Firstly, YOLO11n can maintain a good balance between processing speed and detection accuracy, support real-time detection, and is particularly suitable for sewer pipeline detection that requires rapid processing of a large number of images. Secondly, with its strong convolutional feature extraction ability, YOLO11n can accurately identify defects with small sizes and irregular shapes, and process targets of various sizes through multi-scale detection, ensuring the effective identification of defects of different scales. It also has strong anti-noise ability, can handle problems such as noise, blur, and uneven illumination in sewer pipeline images, and simplifies the development process through end-to-end training. In addition, YOLO11n has good scalability and can run on a variety of hardware platforms to meet the needs of different scenarios.

[0005] However, YOLO11n also faces some challenges in sewer pipeline defect detection, such as difficulties in detecting small objects, complex background interference, and dataset imbalance. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a sewage pipeline defect detection method and system based on an improved YOLO11 network. On the basis of maintaining the lightweight architecture of the model and not increasing the number of parameters, the detection accuracy is successfully improved, and at the same time, the ability to capture target structure information in complex backgrounds and the sensitivity to defects are enhanced, thereby achieving a better detection effect. This not only highlights the significant optimization of the model's accuracy while maintaining high efficiency, but also demonstrates its great deployment potential on real-time detection edge devices such as pipeline robots, providing strong support for the expansion of related application scenarios.

[0007] To achieve the above object, the present invention is implemented by the following technical solutions:

[0008] As Figure 1 shown, an embodiment of the present invention provides a sewage pipeline defect detection method based on an improved YOLO11 network, including the following steps:

[0009] Collect defect image samples in the sewage pipeline using a pipeline CCTV robot;

[0010] Improve the YOLO11n basic network using the MSGA module, AIFI module and positioning loss function Wise-Inner-IoU to construct a sewage pipeline defect model;

[0011] Input the collected defect image samples in the sewage pipeline into the sewage pipeline defect model for training, and use the Python library Albumentations to perform real-time data augmentation on the collected defect image samples during the training process;

[0012] Input the defect target image in the sewage pipeline to be detected into the trained sewage pipeline defect model to obtain the defect detection result in the sewage pipeline to be detected.

[0013] Furthermore, in the process of obtaining defect image samples and defect target images in the pipeline to be detected, an image sensor, a position sensor and an environmental parameter sensor are installed on the CCTV detection robot at the same time, and various types of sensing data are obtained through multiple synchronous acquisitions at a preset sampling frequency in different time periods and different pipeline types;

[0014] According to the layout and detection requirements of the sewage pipeline, the preset path of the CCTV detection robot during the acquisition process is:

[0015] For straight pipelines, a straight-ahead path may be adopted;

[0016] For pipelines with branches, a path that traverses all branches will be planned.

[0017] Furthermore, improving the YOLO11n basic network to construct a sewage pipeline defect model includes the following steps:

[0018] Use the MSGA module to improve the neck network of YOLO11n, and through the introduction of a multi-scale gated attention mechanism, perform adaptive adjustment and selective weighting between feature maps of different scales;

[0019] Use the AIFI module to replace the SPPF module in the backbone network of YOLO11n, and use the self-attention mechanism to enhance the attention to important regions in the image and the representation of pipeline defect features;

[0020] Combine two loss functions, Wise-IoU and Inner-IoU, and their related strategies to construct the loss function Wise-Inner-IoU. Dynamically adjust the gradient gain according to the quality of the anchor box and introduce an auxiliary bounding box mechanism to adjust the size of the auxiliary bounding box and the core area close to the target.

[0021] Furthermore, the MSGA module focuses on small-scale defect regions through multi-scale feature fusion by weighting small-scale features and an adaptive attention mechanism, including the following four stages: multi-scale fusion, spatial selection, spatial interaction and cross modulation, and recalibration.

[0022] Furthermore, in the multi-scale fusion stage, local context extraction widens the spatial range of X through depth convolution and dilated convolution, and global context extraction captures extensive context information through channel pooling in G, and is integrated to form a comprehensive fused feature map U, which is expressed by the formula:

[0023]

[0024] where DW and DW-D represent depth convolution and dilated convolution, Conv 1x1 is a pointwise convolution for local context extraction, P Max and P Avg represent max pooling and average pooling for global context extraction respectively, and the symbol [;] represents concatenation;

[0025] In the spatial selection stage, the fused feature map U is projected onto two channels and aligned with the inputs X and G. The spatial selection weights calculated by channel-wise softmax refine the attention between X and G, generate versions X' and G', use the residual connections of X' and G' to improve gradient flow and feature utilization, and optimize the segmentation effect through context-aware attention, which is expressed by the formula:

[0026]

[0027] where SW i ∈R 2×H×W is obtained through channel-wise softmax (S), ensuring that the sum of the weights of the two channels (i ∈ [1,2]) at each spatial position is 1, representing the relative contribution according to the content input;

[0028] In the spatial interaction and cross modulation layer stage, X′ generates X″ by sigmoid-weighted combination with the global context from G′, while G′ absorbs the detailed context of X′ and evolves into G″. Mutual enhancement ensures that X″ and G″ fuse local and global contexts, and are fused by multiplication, which is expressed by the formula:

[0029]

[0030] In the recalibration stage, an attention map is generated through pointwise convolution and Sigmoid activation to recalibrate the input X from the encoder; after X is multiplied by the attention map, it is further processed through another pointwise convolution layer to enhance its adaptive multi-scale receptive field, thereby providing precise and context-aware features for the decoder, which is expressed by the formula:

[0031]

[0032] Furthermore, the AIFI module is used to process high-level semantic information, focusing on internal scale feature interaction at the S5 layer / high-level feature layer and avoiding interaction of low-level features, including the following steps:

[0033] Linearize the features of the S5 layer using feature linearization and the multi-head self-attention mechanism, and arrange them into rows and concatenate them into a long vector;

[0034] Process this vector using the multi-head self-attention mechanism to capture the connections between concept entities in the image and extract pixel-level semantic features of animals;

[0035] The input high-level features After slicing, flattening, and positional encoding, the input feature X undergoes a linear transformation to obtain:

[0036] Q = W q X

[0037] K = W k X

[0038] V = W v X;

[0039] Obtain the input multi-head self-attention mechanism to obtain the intra-scale interaction features, and the formula is expressed as follows:

[0040]

[0041] F5 = Reshape(Att(Q, K, V))

[0042] Reshape(X) = X i ′ 1,i2,im ;

[0043] Among them, Q is the query, representing the relevance of the query; K is the key, representing self-features; V is the value, representing the utilization of values; W q 、W k 、W v are the mapping matrices for the query, key, and value respectively.

[0044] Further, by combining the two loss functions, Wise-IoU and Inner-IoU, and their related strategies, a loss function Wise-Inner-IoU is constructed, which is expressed by the formula:

[0045] L Inner-WIoUIv3 = L WIoUIv3 + IoU - IoU inner .

[0046] In a second aspect, the present invention provides a sewage pipeline defect detection system based on an improved YOLO11 network, including the following modules:

[0047] A sample collection module, which is used to collect defect image samples in the sewage pipeline by using a pipeline CCTV robot;

[0048] A model construction module, which is used to improve the YOLO11n basic network by using the MSGA module, the AIFI module, and the positioning loss function Wise-Inner-IoU to construct a sewage pipeline defect model;

[0049] A model training module, which is used to input the collected defect image samples in the sewage pipeline into the sewage pipeline defect model for training, and in the training process, use the Python library Albumentations to perform real-time data augmentation on the collected defect image samples;

[0050] A defect detection module, which is used to input the defect target image in the sewage pipeline to be detected into the trained sewage pipeline defect model to obtain the defect detection result in the sewage pipeline to be detected.

[0051] In a third aspect, the present invention provides an electronic system, including:

[0052] A processor;

[0053] A memory for storing executable instructions of the processor;

[0054] Wherein, the processor is configured to execute the instructions to implement the sewage pipeline defect detection method based on the improved YOLO11 network as described in any item of the first aspect.

[0055] In a fourth aspect, the present invention provides a computer-readable storage medium, when the instructions in the storage medium are executed by the processor of an electronic device, enabling the electronic device to execute the sewage pipeline defect detection method based on the improved YOLO11 network as described in any item of the first aspect.

[0056] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The sewage pipeline defect detection method and system based on the improved YOLO11 network provided by the present invention have successfully improved the detection accuracy while maintaining the lightweight architecture of the model and without increasing the number of parameters. At the same time, it enhances the ability to capture the target structure information under complex backgrounds and the sensitivity to defects, thereby achieving a better detection effect. This not only highlights the significant optimization of the model's accuracy while maintaining high efficiency but also demonstrates its great deployment potential on real-time detection edge devices such as pipeline robots, providing strong support for the expansion of related application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 FIG. is a flowchart of the sewage pipeline defect detection method based on the improved YOLO11 network provided by an embodiment of the present invention.

[0058] Figure 2 FIG. is a block diagram of the improved YOLO11 object detection model provided by an embodiment of the present invention.

[0059] Figure 3 FIG. is a block diagram of the internal model of the MSGA module provided by an embodiment of the present invention.

[0060] Figure 4 FIG. is a flowchart of the internal process of the AIFI module provided by an embodiment of the present invention.

[0061] Figure 5 FIG. is a block diagram of the YOLO11 object detection benchmark model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.

[0063] Embodiment

[0064] As Figure 1 shown, an embodiment of the present invention provides a sewage pipeline defect detection method based on the improved YOLO11 network, including the following steps:

[0065] Step S1: Use a pipeline CCTV robot to collect defect image samples in the sewage pipeline;

[0066] Step S2: Use the MSGA module, the AIFI module, and the positioning loss function Wise-Inner-IoU to improve the YOLO11n basic network to construct a sewage pipeline defect model;

[0067] Step S3: Input the collected defect image samples in the sewage pipeline into the sewage pipeline defect model for training, and during the training process, use the Python library Albumentations to perform real-time data augmentation on the collected defect image samples;

[0068] Step S4: Input the defect target image in the sewage pipeline to be detected into the trained sewage pipeline defect model to obtain the defect detection result in the sewage pipeline to be detected.

[0069] In this embodiment, the image samples are the basis for defect detection and are used to train and test the model. The in-pipeline position data can help locate defects and is crucial for subsequent repair work. Environmental parameter data, such as temperature, humidity, light intensity, etc., will affect the image quality and the feature performance of defects. This application uses a multi-sensor fusion method for data collection, and installs an image sensor, a position sensor, and an environmental parameter sensor on the CCTV detection robot at the same time to ensure the synchronous collection of various types of data. To ensure the comprehensiveness and accuracy of the data, multiple collections will be carried out at different time periods and for different pipeline types to cover more actual working conditions.

[0070] The preset path of the CCTV detection robot during the collection process will be designed according to the layout of the sewage pipeline and the detection requirements. For straight pipelines, a straight-forward path may be adopted; for pipelines with branches, a path that traverses all branches will be planned. The setting of the collection frequency is related to the condition of the pipeline and the detection accuracy requirements. In areas where defects may exist, such as pipeline joints and turns, the collection frequency will be increased; while in relatively normal areas, the collection frequency will be appropriately reduced to balance the data volume and detection efficiency.

[0071] The collection rules of the CCTV detection robot are that common collection rules include collection at fixed intervals, collection triggered according to specific events, etc. For example, when the robot detects a sudden change in the environmental parameters in the pipeline, image collection is immediately triggered.

[0072] Preprocess and perform data augmentation on the collected images to improve the image quality and the accuracy of feature extraction. To increase the diversity and quantity of training data and improve the generalization ability of the model, this application uses data augmentation technology. The main data augmentation methods include image flipping, rotation, scaling, cropping, etc. By performing enhancement operations and transformations on the original images, a large number of new images can be generated, thereby expanding the training dataset. To reduce the interference of environmental parameters, especially the influence of humidity and temperature in the sewage pipeline, this application performs normalization processing in the data preprocessing stage.

[0073] Among them, the main steps of the enhancement process fusion include: original image --> standardization processing --> geometric transformation --> noise injection --> blur processing --> normalization output.

[0074] Considering the compression ratio, storage cost, and compatibility of the data. In order to reduce the pressure of data storage and transmission in this application, the JPEG format with a relatively high compression ratio is selected, and the resolution of the images is 640*640, reducing the data volume and accelerating the model processing speed. In terms of color mode, this application adopts the RGB mode to provide rich color information, which helps to distinguish different types of defects.

[0075] In step S2, the MSGA module, AIFI module, and the positioning loss function Wise-Inner-IoU are used to improve the YOLO11n basic network to construct a sewage pipeline defect model, including the following steps:

[0076] Use the MSGA module to improve the neck network of YOLO11n. By introducing a multi-scale gated attention mechanism, adaptive adjustment and selective weighting are performed between feature maps of different scales;

[0077] Use the AIFI module to replace the SPPF module in the backbone network of YOLO11n, and utilize the self-attention mechanism to enhance the attention to important regions in the image and the representation of pipeline defect features;

[0078] Combine the two loss functions Wise-IoU and Inner-IoU and their related strategies to construct the loss function Wise-Inner-IoU, and dynamically adjust the gradient gain according to the quality of the anchor box and introduce an auxiliary bounding box mechanism to adjust the size of the auxiliary bounding box and the core area close to the target.

[0079] Defects in sewer pipes (such as cracks, corrosion, blockages, etc.) are usually small in size. Aiming at the problem that the YOLO11n model is prone to missing detections or inaccurate positioning when dealing with small objects.

[0080] As Figure 3 shown, the MSGA module can weight small-scale features through multi-scale feature fusion, thereby improving the detection accuracy of small objects. Through the adaptive attention mechanism, MSGA enables the model to better focus on small-scale defect areas and avoid missing key details due to interference from background information. The MSGA module includes the following four stages: multi-scale fusion, spatial selection, spatial interaction and cross modulation, and recalibration.

[0081] In the multi-scale fusion stage, local context extraction widens the spatial range of X through depth convolution and dilated convolution, and global context extraction captures extensive context information through channel pooling G and integrates it to form a comprehensive fused feature map U, which is expressed by the formula:

[0082]

[0083] Among them, DW and DW-D represent depthwise convolution and dilated convolution, and Conv 1x1 is a pointwise convolution for local context extraction, and P Max and P Avg represent max pooling and average pooling for global context extraction respectively. The symbol [;] represents concatenation;

[0084] In the spatial selection stage, the fused feature map U is projected onto two channels and aligned with the inputs X and G. The spatial selection weights calculated through channel-wise softmax refine the attention between X and G, generating the versions X' and G', emphasizing the key regions and suppressing the irrelevant regions, and optimizing the receptive field selection. In addition, the residual connections of X' and G' improve the gradient flow and feature utilization, and further optimize the segmentation effect through context-aware attention, which is expressed by the formula:

[0085]

[0086] where SW i ∈R 2×H×W is obtained through channel-wise softmax (S), ensuring that the sum of the weights of the two channels (i ∈ [1, 2]) at each spatial position is 1, representing the relative contribution according to the content input;

[0087] In the spatial interaction and cross-modulation layer stage, X′ generates X″ by sigmoid-weighted combination with the global context from G′, while G′ absorbs the detailed context of X′ and evolves into G″. This mutual enhancement ensures that X″ and G″ fuse local and global contexts, and through multiplicative fusion, combines fine details and global views, effectively segmenting complex structural changes. The formula is expressed as:

[0088]

[0089] In the recalibration stage, an attention map is generated through pointwise convolution and Sigmoid activation for recalibrating the input X from the encoder; after X is multiplied by the attention map, it is further processed through another pointwise convolution layer to enhance its adaptive multi-scale receptive field, thereby providing accurate, context-aware features for the decoder and promoting accurate segmentation, which is expressed by the formula:

[0090]

[0091] The SPPF (Spatial Pyramid Pooling) of YOLO11n is processed independently during feature extraction and lacks a direct feature interaction mechanism. Therefore, the AIFI module is introduced to process high-level semantic information, focusing on internal scale feature interaction at the S5 layer (high-level feature layer) and avoiding the interaction of low-level features, as Figure 4 shown, including the following steps:

[0092] Linearize the features of layer S5 using feature linearization and the multi-head self-attention mechanism, arrange them in rows, and concatenate them into a long vector;

[0093] Process this vector using the multi-head self-attention mechanism to capture the connections between concept entities in the image and extract pixel-level semantic features of the animal;

[0094] The input high-level features After slicing, flattening, and positional encoding, the input feature X undergoes a linear transformation to obtain:

[0095] Q = W q X

[0096] K = W k X

[0097] V = W v X;

[0098] Obtain the input multi-head self-attention mechanism to obtain the intra-scale interaction features, and the formula is as follows:

[0099]

[0100] F5 = Reshape(Att(Q, K, V))

[0101] Reshape(X) = X i ′ 1,i2,im ;

[0102] Among them, Q is the query Query, representing the relevance of the query; K is the key Key, representing the self-feature; V is the value Value, representing the utilization value; W q , W k , W v are the mapping matrices for the query, key, and value respectively.

[0103] As Figure 5 shown, combine the two loss functions Wise-IoU and Inner-IoU and their related strategies to construct a loss function Wise-Inner-IoU, and the algorithm is as follows:

[0104] a) Wise-IoU

[0105] L WIoUv3 = rR WIoU L IoU ;

[0106]

[0107] b) Inner-IoU

[0108]

[0109]

[0110] union=(w gt *h gt )*(ratio) 2 +(w*h)*(ratio) 2 -inter;

[0111]

[0112] wherein, the GT box and the predicted box are respectively denoted as b gt and b. denotes the center point of the GT box, where (x c , y c ) denotes the center point of the anchor. The widths and heights of the GT box and the anchor are denoted by w gt , h gt and w, h respectively. The variable ratio corresponds to the scaling factor, usually within the numerical range of [0.5, 1.5].

[0113] Therefore, combining the two can obtain the Wise’-Inner-IoU loss function, which is expressed by the formula:

[0114] L Inner-WIoUIv3 =L WIoUIv3 +IoU - IoU inner .

[0115] The dynamic allocation strategy of the Wise-IoU gradient gain can effectively reduce the adverse effects of low-quality samples, thereby improving the detection accuracy and effect of the network. Inner-IoU introduces a scaled-down auxiliary bounding box to optimize the loss function in complex scenarios of sewage pipelines, which include small targets, overlapping and occluded targets. It focuses on the overlapping area and can improve the accuracy in the case of high overlap. Fine-tuning its scale can make Inner-IoU focus on the core of small targets, reduce background noise and enhance the target detection effect. At the same time, Inner-IoU combines the dynamic focusing mechanism of Wise-IoU. This loss function can intelligently adjust its focus, enabling the model to pay more attention to difficult-to-distinguish targets or regions during training, such as highly similar or severely occluded targets.

[0116] Step S3: Input the collected defect image samples of the sewage pipeline into the sewage pipeline defect model for training, and during the training process, use the Python library Albumentations to perform real-time data augmentation on the collected defect image samples, specifically including operations such as geometric transformation, blurring, and noise addition.

[0117] In step S3, first, a series of preprocessing operations are performed on the collected defect image samples in the sewage pipeline, including normalizing the images to eliminate the influence of environmental factors such as light intensity, humidity, and temperature on the image quality. At the same time, the images are cropped and scaled to a unified resolution (e.g., 640×640) to adapt to the input requirements of the model. In addition, to ensure that the model can learn the accurate location and category information of the defect targets, the images also need to be annotated to generate corresponding defect target bounding boxes and category labels for supervised learning.

[0118] After the preprocessing is completed, the Python library Albumentations is used to perform data augmentation on the image samples to expand the diversity and scale of the dataset. The specific augmentation methods include random rotation (±30°), horizontal flipping, vertical flipping, random cropping (retaining 80%-100% of the area), color adjustment (random change range of brightness, contrast, and saturation is ±20%), noise injection (Gaussian noise and salt-and-pepper noise), and blurring (Gaussian blur, kernel size of 3×3 or 5×5), etc. The above data augmentation means can simulate the image features under different working conditions and enhance the model's adaptability and generalization ability to complex environments.

[0119] Step S4: Input the defect target image in the sewage pipeline to be detected into the trained sewage pipeline defect model to obtain the defect detection result in the sewage pipeline to be detected.

[0120] In step S4, after the preprocessing operation, ensure that the format, resolution, and color mode of the image to be detected are consistent with the training data, and input the image into the trained sewage pipeline defect detection model. The model obtains the detection result through forward propagation calculation, and the output includes the bounding box coordinates (x, y, w, h) of the defect target, the confidence score, and the corresponding category label.

[0121] To improve the accuracy and reliability of the detection result, post-processing is performed on the detection result output by the model.

[0122] First, set a confidence threshold (e.g., 0.5), and only retain the detection results with a confidence higher than this threshold to filter out misdetected targets with low confidence.

[0123] Secondly, the non-maximum suppression (NMS) algorithm is used to process the overlapping detection boxes. Specifically, the detection boxes are sorted from high to low according to the confidence, the detection box with the highest confidence is retained, and other detection boxes with an overlap degree (IoU) greater than a certain threshold (e.g., 0.5) with it are removed, thus eliminating the problem of repeated detection.

[0124] In addition, according to the requirements of the actual application scenario, the detected bounding boxes are appropriately adjusted and optimized. For example, the range of the bounding boxes is expanded to ensure the integrity of the defect area, or the bounding boxes are cropped to remove the invalid edge parts.

[0125] Finally, the detection results are presented in a visual way. The detected defect target bounding boxes are drawn on the original image, and the corresponding class labels and confidence levels are marked to intuitively display the detection results. At the same time, the detection results are output in a structured data format (CSV file), including information such as the class, location coordinates, and confidence level of the defect targets, which is convenient for subsequent analysis and processing.

[0126] In addition, for the detected defect targets, classification statistics and alarm prompts can also be carried out according to their classes and severity levels, providing decision-making support for pipeline maintenance and repair work.

[0127] To comprehensively and systematically evaluate the roles of the various improved modules in the improved YOLO11n model, ablation experiments were carried out for three key improvement points under the same experimental environment and dataset. Seven different module combinations were designed, and the performance of each improvement point was tested separately and synergistically, and compared with the baseline model YOLO11n. Table 1 below shows the experimental results of these combinations on the test set.

[0128] Table 1

[0129]

[0130] As can be seen from Table 1 above, when the MSGA module is introduced to replace the basic structure in the YOLO11n backbone and neck networks, the model accuracy is improved from 75.8% to 81.3% (↑5.5%), mAP@0.5 reaches 83.2% (↑0.3%), but the recall rate slightly decreases by 0.2%. As the harmonic mean of precision and recall, the F1 score drops to 0.77 (↓0.04). This indicates that the MSGA module significantly optimizes the extraction efficiency of the model for high-resolution detailed features by enhancing the multi-scale feature aggregation ability, thereby improving the accuracy of classification confidence. However, it may over-focus on local significant targets, resulting in the missed detection of some low-confidence targets and affecting the balance between recall rate and F1.

[0131] After enabling the AIFI module alone, the recall rate is improved from 79.5% to 82.3% (↑2.8%), mAP@0.5 increases to 84.4% (↑1.5%), and the F1 score stabilizes at 0.8. This is because the AIFI module strengthens the model's ability to associate and model multi-scale targets in complex scenarios through an adaptive feature interaction mechanism. Especially in occlusion or small-target scenarios, it can capture more context information of low-confidence targets, thus significantly improving the recall rate and localization robustness.

[0132] After introducing Wise-Inner-IoU to replace the traditional IoU calculation, the mAP@0.5 increased from 82.9% of the baseline to 86.4% (↑3.5%), the recall rate increased by 3.0% (79.5% → 82.5%), but the F1 score fluctuated slightly due to the decrease in precision (75.8% → 78.6%). Wise-Inner-IoU effectively suppresses the interference of low-quality samples on the training process by dynamically adjusting the weight distribution of bounding box regression, while optimizing the geometric alignment accuracy of bounding boxes, thus showing significant advantages in the localization task.

[0133] After synergistically integrating the three modules, the improved model reached its peak in terms of precision, F1, and mAP@0.5, which were 84.4% (↑8.6%), 0.85 (↑0.04), and 87.2% (↑4.3%) respectively. Although the recall rate did not increase significantly compared to the baseline model (79.5% → 80.0%), the detail enhancement of MSGA, the context modeling of AIFI, and the localization optimization of Wise-Inner-IoU complemented each other:

[0134] 1) MSGA focuses on the accurate classification of high-confidence targets, laying the foundation for precision;

[0135] 2) AIFI captures the ignored low-confidence targets through global feature interaction, alleviating the decline in recall rate;

[0136] 3) Wise-Inner-IoU optimizes the localization deviation from the level of the loss function, driving a significant increase in mAP.

[0137] Finally, the module collaboration strategy achieved improvements in all core metrics. In particular, the 4.3% increase in mAP@0.5 verified the high adaptability of the improved solution to complex detection tasks. If further optimization of the recall rate is required, the deep combination of the AIFI module and the dynamic label assignment strategy can be explored to balance the marginal effect between precision and recall.

[0138] In the embodiments of the present invention, a sewage pipeline defect detection system based on the improved YOLO11 network is also provided, including the following modules:

[0139] A sample acquisition module for collecting defect image samples in the sewage pipeline using a pipeline CCTV robot;

[0140] A model construction module for improving the YOLO11n basic network using the MSGA module, the AIFI module, and the localization loss function Wise-Inner-IoU to construct a sewage pipeline defect model;

[0141] A model training module for inputting the collected defect image samples in the sewage pipeline into the sewage pipeline defect model for training, and using the Python library Albumentations to perform real-time data augmentation on the collected defect image samples during the training process;

[0142] A defect detection module for inputting the defect target image in the sewage pipeline to be detected into the trained sewage pipeline defect model to obtain the defect detection result in the sewage pipeline to be detected.

[0143] In an embodiment of the present invention, an electronic system is further provided, including:

[0144] A processor;

[0145] A memory for storing executable instructions of the processor;

[0146] Wherein, the processor is configured to execute the instructions to implement the sewage pipeline defect detection method based on the improved YOLO11 network as described in any one of the above.

[0147] In an embodiment of the present invention, a computer-readable storage medium is further provided. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the sewage pipeline defect detection method based on the improved YOLO11 network as described in any one of the above.

[0148] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. A sewage pipe defect detection method based on an improved YOLO11 network is characterized in that: The following steps are involved: Use pipeline CCTV robots to collect defect image samples in sewage pipelines; The MSGA module, AIFI module and positioning loss function Wise-Inner-IoU are used to improve the YOLO11n basic network to build a sewage pipe defect model; The collected defect image samples in the sewage pipe are input into the sewage pipe defect model for training, and during the training process, the Python library Albumentations is used to perform real-time data enhancement on the collected defect image samples; The defect target image in the sewage pipe to be detected is input into the trained sewage pipe defect model to obtain the defect detection result in the sewage pipe to be detected.

2. The sewage pipe defect detection method based on the improved YOLO11 network according to claim 1 is characterized in that: In the process of acquiring defect image samples and defect target images in the pipeline to be inspected, image sensors, position sensors and environmental parameter sensors are installed on the CCTV inspection robot at the same time, and multiple synchronous acquisitions are performed at preset sampling frequencies in different time periods and different pipeline types to obtain various sensor data; According to the layout and detection requirements of the sewage pipeline, the preset path of the CCTV detection robot during the collection process is: For straight pipelines, a straight forward path may be adopted; For pipelines with branches, a path that traverses all branches is planned.

3. The sewage pipe defect detection method based on the improved YOLO11 network according to claim 1 is characterized in that: Improvements to the YOLO11n base network to build a sewage pipe defect model include the following steps: The MSGA module is used to improve the neck network of YOLO11n, and a multi-scale gated attention mechanism is introduced to perform adaptive adjustment and selective weighting between feature maps of different scales; Use AIFI module to replace SPPF module in YOLO11n backbone network, and use self-attention mechanism to enhance attention to important areas in the image and pipeline defect feature representation; The two loss functions Wise-IoU and Inner-IoU and their related strategies are combined to construct the loss function Wise-Inner-IoU. The gradient gain is dynamically adjusted according to the quality of the anchor box, and the auxiliary bounding box mechanism is introduced to adjust the size of the auxiliary bounding box and the core area close to the target.

4. The sewage pipe defect detection method based on the improved YOLO11 network according to claim 1 or 3, characterized in that: The MSGA module weights small-scale features through multi-scale feature fusion and focuses on small-scale defect areas through an adaptive attention mechanism, which includes the following four stages: multi-scale fusion, spatial selection, spatial interaction and cross-modulation, and recalibration.

5. The sewage pipe defect detection method based on the improved YOLO11 network according to claim 4 is characterized in that: In the multi-scale fusion stage, local context extraction broadens the spatial range of X through deep convolution and dilated convolution, and global context extraction captures extensive context information through channel pooling G, which is integrated to form a comprehensive fusion feature map U, which is expressed as: Among them, DW and DW-D represent deep convolution and dilated convolution, Conv 1x1 is a pointwise convolution for local context extraction, P Max and P Avg They represent the maximum pooling and average pooling for global context extraction, respectively. The symbol [;] indicates a connection; In the spatial selection stage, the fused feature map U is projected into two channels and aligned with the inputs X and G. The spatial selection weights calculated by channel-level softmax are used to refine the attention between X and G, generate versions X' and G', and use the residual connection of X' and G' to improve gradient flow and feature utilization, as well as optimize the segmentation effect through context-aware attention. The formula is expressed as: in, It is obtained through channel-level softmax(S), ensuring that the sum of the weights of the two channels (i∈[1,2]) at each spatial position is 1, indicating the relative contribution of the input according to the content; In the spatial interaction and cross-modulation layer stage, X′ is combined with the global context from G′ through sigmoid weighting to generate X″, while G′ absorbs the detail context of X′ and evolves into G″. Mutual enhancement ensures that X″ and G″ fuse the local and global contexts and are fused through multiplication. The formula is expressed as: In the recalibration phase, an attention map is generated by point-wise convolution and sigmoid activation to recalibrate the input X from the encoder; after X is multiplied by the attention map, it is further processed by another point-wise convolution layer to enhance its adaptive multi-scale receptive field, thereby providing accurate and context-aware features for the decoder. The formula is expressed as:

6. The sewage pipe defect detection method based on the improved YOLO11 network according to claim 4 is characterized in that: The AIFI module is used to process high-level semantic information, focusing on the interaction of internal scale features and avoiding the interaction of low-level features at the S5 layer / high-level feature layer, and includes the following steps: The features of the S5 layer are linearized using feature linearization and multi-head self-attention mechanism, and are arranged into rows and concatenated into long vectors. The vector is processed using a multi-head self-attention mechanism to capture the connections between conceptual entities in the image and extract pixel-level semantic features of the animal. High-level features of the input After slicing, flattening and position encoding, the input feature X is linearly transformed to obtain: Q=W q X K=W k X V=W v X; The input multi-head self-attention mechanism is obtained to obtain the intra-scale interaction features. The formula is as follows: F5=Reshape(Att(Q,K,V)) Reshape(X)=X′ i1,i2,im ; Among them, Q is the query, indicating the relevance of the query; K is the key, indicating the self-feature; V is the value, indicating the utilization value; W q , W k , W v are mapping matrices for query, key, and value respectively.

7. The sewage pipe defect detection method based on the improved YOLO11 network according to claim 4 is characterized in that: The two loss functions Wise-IoU and Inner-IoU and their related strategies are combined to construct a loss function Wise-Inner-IoU, which is expressed as follows: L Inner-WIoUIv3 =L WIoUIv3 +IoU-IoU inner .

8. The sewage pipe defect detection system based on the improved YOLO11 network is characterized by: Includes the following modules: The sample collection module is used to collect defect image samples in the sewage pipe using the pipeline CCTV robot; Model building module, which is used to improve the YOLO11n basic network using the MSGA module, AIFI module and positioning loss function Wise-Inner-IoU to build a sewage pipe defect model; The model training module is used to input the collected defect image samples in the sewage pipe into the sewage pipe defect model for training, and during the training process, the Python library Albumentations is used to perform real-time data enhancement on the collected defect image samples; The defect detection module is used to input the defect target image in the sewage pipe to be detected into the trained sewage pipe defect model to obtain the defect detection result in the sewage pipe to be detected.

9. An electronic system, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the sewage pipe defect detection method based on the improved YOLO11 network as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the sewage pipe defect detection method based on the improved YOLO11 network as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Bird identification and tracking method, device and system for transformer substation, and storage medium

    CN121170849A

  • Photovoltaic cell defect detection method fusing multi-scale features and re-parameterization strategy

    CN121329931A

  • Method and device for detecting small target defects of underground pipeline based on CMS-RTDETR model

    CN122612610A