Sewage pipeline defect detection method and device based on multi-scale image feature fusion
Through the lightweight pipeline defect detection model based on multi-scale image feature fusion, the problem of traditional manual interpretation methods is solved, and efficient and accurate sewage pipeline defect detection is achieved, and the accuracy of the detection results is improved.
Patent Information
- Application Number
- CN202510188523.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The traditional manual interpretation method consumes time and effort in sewage pipeline defect detection, and is easily affected by the experience and subjective judgment of the operator, resulting in deviations in the identification results and it is difficult to meet the needs of efficient and accurate modern inspections.
A lightweight pipeline defect detection model based on multi-scale image feature fusion is adopted, and the improved YOLOv8 model, including the C2f-FAM module, the HS-BiFPN module and the DySample module, real-time detection and analysis of the internal images of the sewage pipeline are realized.
It improves the efficiency and accuracy of sewage pipeline defect detection, reduces the dependence of manual interpretation, enhances the processing ability of small targets and complex backgrounds, and significantly improves the accuracy of the detection results.
Smart Images

Figure CN120163772A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pipeline detection, and particularly relates to a sewage pipeline defect detection method and device based on multi-scale image feature fusion. Background Art
[0002] As an important part of urban infrastructure, sewage pipelines are mainly responsible for collecting and transporting domestic sewage, industrial wastewater, and rainwater. Their operating conditions directly affect the development level, public health, and environmental quality of cities. Therefore, ensuring the efficient operation of the sewage pipeline system is crucial for the sustainable development of cities. With the continuous growth of the economy and the acceleration of urbanization, significant progress has been made in the construction and development of urban sewage pipe networks in China. However, in recent years, there has been a common problem of "emphasizing construction while neglecting management and maintenance" in sewage pipeline management, resulting in the gradual emergence of defects such as pipeline structure damage, corrosion, and blockage. With the aging of pipelines and the continuous increase in sewage volume, these problems have become more serious, not only affecting daily life but also potentially causing safety hazards such as road collapses and urban waterlogging. Therefore, it is particularly important to regularly inspect the internal conditions of sewage pipelines in order to timely assess the type and location of defects and take appropriate countermeasures, such as pipeline maintenance, repair, or replacement of severely damaged parts.
[0003] In the inspection and maintenance of sewage pipelines, the CCTV (Closed Circuit Television) system has become a widely used detection method. Generally, a robot equipped with a camera device and a lighting device moves inside the pipeline to record videos in real time for evaluating the structural condition of the pipeline. This technology has been widely recognized for its advantages such as simple operation and low cost. Through CCTV videos, technicians can observe the defect conditions on the inner surface of the pipeline, such as cracks, collapses, and sediments. Although CCTV video analysis has been widely applied, traditional manual interpretation methods are not only time-consuming and laborious but also easily affected by the experience and subjective judgment of operators, resulting in deviation in recognition results. For example, cracks may be misjudged as fractures, and the accuracy of manual detection results is easily affected by the experience of the inspectors. Although technicians can refer to standards such as the Pipeline Assessment and Certification Program (PACP) as auxiliary tools, this human-based inspection method is difficult to meet the modern detection requirements of high efficiency and accuracy. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a sewage pipeline defect detection method and device based on multi-scale image feature fusion.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A sewage pipeline defect detection method based on multi-scale image feature fusion, comprising:
[0007] Step S1: Obtain the historical sewage pipeline defect image dataset;
[0008] Step S2: Preprocess and perform data augmentation on the historical pipeline defect image dataset;
[0009] Step S3: Train an improved YOLOv8 model based on the preprocessed and data-augmented historical sewage pipeline defect image dataset to obtain a lightweight pipeline defect detection model;
[0010] Step S4: Input the real-time collected internal image of the sewage pipeline into the lightweight pipeline defect detection model for pipeline defect detection.
[0011] Preferably, in Step S2, the preprocessing includes grayscale conversion, contrast enhancement, and denoising, and the data augmentation includes: randomly cropping, horizontally, vertically, and randomly rotating the historical sewage pipeline defect image dataset.
[0012] Preferably, the improved YOLOv8 model includes: C2f-FAM module, HS-BiFPN module, and Upsample module; among them, the C2f-FAM module introduces EMSConvP multi-scale convolution for efficient extraction of features at different scales; the HS-BiFPN module replaces the original neck of the YOLOv8 model, and through cross-level feature fusion operations, realizes flexible fusion of feature maps at different scales; through the upsampling of the DySample module, the scale of the feature map is dynamically optimized according to the size of the target and the complexity of the scene.
[0013] The present invention also provides a sewage pipeline defect detection device based on multi-scale image feature fusion, including:
[0014] An acquisition module for acquiring the historical sewage pipeline defect image dataset;
[0015] A processing module for preprocessing and performing data augmentation on the historical pipeline defect image dataset;
[0016] A training module for training an improved YOLOv8 model based on the preprocessed and data-augmented historical sewage pipeline defect image dataset to obtain a lightweight pipeline defect detection model;
[0017] A detection module for inputting the real-time collected internal image of the sewage pipeline into the lightweight pipeline defect detection model for pipeline defect detection.
[0018] Preferably, the preprocessing includes grayscale conversion, contrast enhancement, and denoising, and the data augmentation includes: randomly cropping, horizontally, vertically, and randomly rotating the historical sewage pipeline defect image dataset.
[0019] Preferably, the improved YOLOv8 model includes: a C2f-FAM module, an HS-BiFPN module, and an Upsample module; among them, the C2f-FAM module introduces EMSConvP multi-scale convolution for efficient extraction of features at different scales; the HS-BiFPN module replaces the original neck of the YOLOv8 model, and through cross-level feature fusion operations, realizes flexible fusion of feature maps at different scales; through the upsampling of the DySample module, the scale of the feature map is dynamically optimized according to the size of the target and the complexity of the scene.
[0020] Based on YOLOv8, the present invention proposes a lightweight pipeline defect detection model based on multi-scale feature fusion. First, the C2f module in the backbone network is replaced with a C2f-FAM module to enhance the ability of multi-scale feature extraction. Secondly, the HS-BiFPN is used to replace the original neck structure, enhancing the model's ability to recognize multi-scale targets and enrich feature representations. Finally, DySample is introduced to optimize the upsampling operation to improve the model's target capture ability in complex environments. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0022] Figure 1 It is a flowchart of the sewage pipeline defect detection method in the embodiment of the present invention;
[0023] Figure 2 It is a schematic diagram of the improved YOLOv8 model in the embodiment of the present invention;
[0024] Figure 3 It is a schematic diagram of EMSConvP in the embodiment of the present invention;
[0025] Figure 4 It is a schematic diagram of HS-BiFPN in the embodiment of the present invention;
[0026] Figure 5 It is a schematic diagram of Cross-Level Feature Fusion in the embodiment of the present invention;
[0027] Figure 6 It is a schematic diagram of DySample in the embodiment of the present invention; among them, (a) is a schematic diagram of the DySample structure; (b) is a schematic diagram of the static range factor version and the dynamic range factor version. Detailed Embodiments
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0029] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.
[0030] Embodiment 1:
[0031] As Figure 1 shown, a sewage pipeline defect detection method based on multi-scale image feature fusion in an embodiment of the present invention includes:
[0032] Step S1, obtaining a historical sewage pipeline defect image dataset;
[0033] Step S2, performing preprocessing and data augmentation processing on the historical pipeline defect image dataset;
[0034] Step S3, training an improved YOLOv8 model according to the preprocessed and data-augmented historical sewage pipeline defect image dataset to obtain a lightweight pipeline defect detection model;
[0035] Step S4, inputting the internally collected real-time sewage pipeline image into the lightweight pipeline defect detection model for pipeline defect detection.
[0036] As an implementation manner of the embodiment of the present invention, in step S2, the preprocessing includes grayscale conversion, contrast enhancement, and denoising processing, and the data augmentation includes: randomly cropping, horizontally, vertically, and randomly rotating the historical sewage pipeline defect image dataset.
[0037] As an implementation manner of the embodiment of the present invention, as Figure 2As shown in the figure, the improved YOLOv8 model includes: C2f-FAM module, HS-BiFPN module and Upsample module; among them, the C2f-FAM module introduces EMSConvP multi-scale convolution for efficient extraction of features at different scales; the HS-BiFPN module replaces the original neck of the YOLOv8 model, and through cross-level feature fusion operations, realizes flexible fusion of feature maps at different scales, so that the feature maps at each scale can make more full use of the context information of other scales, thereby generating fine-grained feature expressions, and further improving the multi-scale object detection ability of the model; through the upsampling of the DySample module, the scale of the feature map is dynamically optimized according to the size of the target and the complexity of the scene. This adaptive adjustment mechanism enables the model to capture details more effectively when facing targets of various sizes and different backgrounds, so as to improve the ability of small target detection and complex background processing.
[0038] Furthermore, the working process of the C2f-FAM module is as follows:
[0039] In the sewage pipeline detection task, since the dataset usually contains a large number of small target objects with low resolution, although the traditional convolutional layer expands the receptive field through downsampling operations, it is easy to cause the loss of key feature information, thus affecting the detection accuracy. Inspired by EMCAD
[29] , the embodiment of the present invention proposes an improved module C2f-FAM, which enhances the ability to extract features at different scales by introducing EMSConvP multi-scale convolution. In this module, C2f-FAM replaces the standard convolution (Conv) at the second position in the traditional Bottleneck structure with EMSConvP.
[0040] The EMSConvP module consists of two parts: multi-scale convolution and channel fusion layer. First, features with different receptive fields are extracted through multi-scale convolution kernels to enhance the expression ability of multi-scale targets; subsequently, 1×1 convolution is used to adjust the number of channels and effectively fuse multi-scale features.
[0041] The multi-scale convolution uses convolution kernels of different sizes (1×1, 3×3, 5×5 and 7×7), and the input feature map is processed by channel grouping. The number of channels in each group is dynamically calculated based on the number of input channels and the number of convolution kernels, and by restricting the number of channels in each group to be not less than 16 to ensure the effectiveness of the convolution operation. Therefore, the embodiment of the present invention chooses to replace only the third and fourth C2f modules in the backbone network. During the feature extraction process, the feature maps of each group perform convolution operations independently, and the structure is as Figure 3 shown.
[0042] After completing the multi-scale convolution, the feature maps are merged through a concatenation operation, and the number of channels is adjusted and the features are fused through 1×1 convolution. This design realizes the fusion of multi-scale features while retaining the global semantic information, thus enhancing the model's sensitivity to features of different scales. Compared with the traditional full-channel convolution, this module reduces the computational complexity and the number of parameters through group convolution, thereby reducing the computational overhead and memory occupancy, and improving the operation efficiency while maintaining the performance.
[0043] The C2f-FAM module enhances the multi-scale feature extraction ability in pipeline detection by introducing EMSConvP, especially improving the detection performance for small targets. This module effectively reduces the computational complexity and the number of parameters through group convolution, thus optimizing the computational efficiency and memory occupancy while maintaining high detection accuracy, and significantly enhancing the detection ability of the model.
[0044] Furthermore, the working process of the HS-BiFPN module is as follows:
[0045] To address challenges such as variable sizes, noise interference, and environmental changes in pipeline defect detection, inspired by the existing multi-scale feature fusion method HS-FPN
[23] , a lightweight neck structure - HS-BiFPN is proposed. Compared with the original neck structure of YOLOv8, HS-BiFPN effectively improves the detection accuracy while reducing the computational amount and the number of parameters. The structure is as Figure 4 shown.
[0046] HS-BiFPN mainly consists of two parts: a feature selection module and a cross-stage feature fusion module. The specific process is as follows: First, the different-scale feature maps generated by the backbone network will undergo effective feature screening through the feature selection module. Subsequently, these different-scale feature maps will be fused through the cross-stage feature fusion module to generate features with richer semantic information. This fusion helps to more precisely capture the subtle features in the pipeline image, thereby enhancing the detection ability of the model.
[0047] Feature Selection Module: In the CA module, first, global average pooling and global max pooling operations are performed on the input feature map to calculate the average and maximum values of each channel respectively, so as to extract global features. Global average pooling helps to uniformly obtain information from the feature map, while global max pooling focuses on extracting the most representative data from each channel, thus minimizing information loss. Through this pooling method, the CA module can effectively capture the important features of each channel and reduce the interference of redundant information. Then, the Sigmoid activation function is used to generate the weight values of each channel, and each channel is weighted with these values to strengthen the expression of key features. Subsequently, the calculated weights are multiplied with the original feature map channel by channel to generate a weighted feature map. This process enables the model to focus on more discriminative features, suppress redundant or irrelevant features, and thus improve the expression ability of the feature map. In addition, to ensure the smooth progress of cross-level feature fusion and achieve effective fusion between feature maps of different scales, 1×1 convolution operations are adopted to match the dimensions of feature maps of different scales, and the number of channels of each feature map is adjusted to a unified 256, so that in the subsequent feature fusion stage, feature maps can be fused within the same dimensional space, thus avoiding information loss or inconsistency caused by channel mismatch.
[0048] Cross-Level Fusion Module: In a deep neural network, the multi-scale feature maps generated by the backbone network usually contain semantic information at different levels. High-level features usually come from the deeper layers of the network and have high semantic richness, capable of capturing global context information and abstract target features. However, due to the low spatial resolution of high-level feature maps, their ability to accurately locate targets is relatively weak, especially in the detection of small targets, showing certain limitations. In contrast, low-level features usually come from the shallower layers of the network, have high spatial resolution, and can accurately locate the position of the target and capture detailed information. However, low-level features are relatively limited in semantic expression and are difficult to effectively distinguish different targets in complex scenes. Therefore, relying solely on high-level or low-level features has its own deficiencies, and the fusion of the two can effectively make up for their respective shortcomings. By combining the advantages of high-level and low-level features, the expression ability of the feature map can be enhanced, thereby improving the accuracy and robustness of target detection.
[0049] Therefore, the embodiment of the present invention proposes a Cross-Level Feature Fusion (CLF) module, as Figure 5As shown. In the top-down path, first, the high-level features are used as weights to perform feature fusion on the low-level features to enhance the semantic information of the low-level features. Specifically, given a high-level feature map and a low-level feature map as inputs, the CLF module first upsamples the high-level features using a 3×3 convolutional kernel and a transposed convolution with a stride of 2. Subsequently, the CA module processes the high-level features to generate the weight values for each channel and uses these weights to filter the redundant information in the low-level features. After this filtering process, the low-level features are fused with the high-level features to generate new feature maps N3, N4, and N5. Next, these fused feature maps enter the bottom-up path for cross-level feature fusion again. In this process, the low-level features are downsampled using a 3×3 convolutional kernel and a convolution with a stride of 2, and the CA module generates corresponding weights to filter the high-level features. The filtered high-level features are re-fused with the low-level features. Through this series of top-down and bottom-up cross-level fusion operations, not only is the information loss effectively reduced, but also the expressive ability of the feature maps is enhanced, and finally, feature maps with higher semantic information and spatial resolution are generated. In this way, the model can better capture the details of multi-scale targets, thereby improving the detection accuracy in complex scenarios.
[0050] Furthermore, the working process of the Upsample module is as follows:
[0051] Upsampling is a technique widely used in image processing and deep learning, aiming to convert low-resolution images or feature maps into high-resolution outputs. By enlarging the image size, upsampling can effectively improve the resolution of the image or feature map, thereby enhancing the model's ability to capture detailed information. Its core goal is to restore a more refined feature representation and provide richer feature support for subsequent processing and analysis. In YOLOv8, the traditional upsampling method uses nearest-neighbor interpolation. This method is widely used because of its fast calculation speed and simple implementation, but its limitation is that it only expands by copying the values of the nearest pixels and cannot generate new pixel information. Therefore, nearest-neighbor interpolation is weak in detail preservation and feature expression. Especially in complex pipeline defect detection tasks, noise interference and environmental complexity make it difficult to effectively capture subtle features and cannot meet the actual needs. To overcome these problems, the embodiments of the present invention adopt DySample as the upsampling operator.
[0052] The input feature map, upsampled feature map, generated offset, and original sampling grid of DySample are respectively denoted as X, X', O, and G. As Figure 6 shown in (a) of, DySample first constructs a sampling set S from the input feature map X through a point sampling generator, and then re-samples the sampling set using a grid sampling function to finally generate the upsampled feature map X'. This process can be represented by the following formula (1):
[0053]
[0054] DySample provides two versions of the generator: the static range factor version and the dynamic range factor version, as shown in (b) of Figure 6 As shown. In the static range factor version, the offset O is generated through a linear layer and a Pixel Shuffle operation, and combined with the original sampling grid G to finally obtain the sampling set S. This process is shown in Formulas (2) and (3). In the dynamic range factor version, a range factor is first generated, and then this factor is used to modulate the offset O. The generation of the dynamic offset depends on the Sigmoid function (σ). By adding the adjusted offset O to the sampling grid G, the sampling set S is generated. This method enables the dynamic range factor version to have stronger adaptability during the upsampling process, being able to flexibly adjust the sampling position according to the local features and context information of the input image, thereby improving the ability to retain details in the upsampling result.
[0055]
[0056] The key innovation of DySample lies in its dynamic sampling mechanism. By combining the generated offset with the original grid position, it dynamically adjusts the sampling position, thereby significantly improving the quality of the upsampled image. Especially in the processing of detail retention and complex textures, DySample demonstrates more excellent performance compared to traditional upsampling methods. By introducing the modulation of static and dynamic range factors, DySample can flexibly adjust the sampling process according to the local information of the input features, and thus achieve more accurate image reconstruction.
[0057] The embodiment of the present invention adopts an improved pipeline defect detection model; first, the C2f-FAM module is proposed, and by introducing the multi-scale convolution module EMSConvP, the ability of multi-scale feature extraction and the detection ability of small targets are enhanced. Then, the HS-BiFPN is designed, and by effectively fusing the feature maps from different scales, the complementary fusion between the low-resolution and high-resolution feature maps is realized, so as to more comprehensively capture the features of targets at different scales. In addition, the DySample dynamic upsampling module is introduced, which can adjust the resolution of different feature maps as needed to improve the adaptability of the model to targets at different scales, and thus improve the ability of small target detection and complex background processing.
[0058] Embodiment 2:
[0059] The embodiment of the present invention also provides a sewage pipeline defect detection device based on multi-scale image feature fusion, including:
[0060] An acquisition module, configured to acquire a historical sewage pipeline defect image data set;
[0061] A processing module for preprocessing and data augmentation processing of the historical pipeline defect image dataset;
[0062] A training module for training an improved YOLOv8 model based on the preprocessed and data-augmented historical sewage pipeline defect image dataset to obtain a lightweight pipeline defect detection model;
[0063] A detection module for inputting the internally captured image of the sewage pipeline in real time into the lightweight pipeline defect detection model for pipeline defect detection.
[0064] As an implementation manner of an embodiment of the present invention, the preprocessing includes grayscale conversion, contrast enhancement, and denoising processing, and the data augmentation includes: randomly cropping, horizontally, vertically, and randomly rotating the historical sewage pipeline defect image dataset.
[0065] As an implementation manner of an embodiment of the present invention, the improved YOLOv8 model includes: a C2f-FAM module, an HS-BiFPN module, and an Upsample module; wherein, the C2f-FAM module introduces EMSConvP multi-scale convolution for efficient extraction of features of different scales; the HS-BiFPN module replaces the original neck of the YOLOv8 model, and realizes flexible fusion of feature maps of different scales through cross-level feature fusion operations; through the upsampling of the DySample module, the scale of the feature map is dynamically optimized according to the size of the target and the complexity of the scene.
[0066] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A sewage pipe defect detection method based on multi-scale image feature fusion, characterized in that: include: Step S1, obtaining a historical sewage pipe defect image dataset; Step S2, preprocessing and data expansion processing of the historical pipeline defect image dataset; Step S3, training an improved YOLOv8 model based on the historical sewage pipe defect image dataset after preprocessing and data expansion to obtain a lightweight pipeline defect detection model; Step S4: input the real-time collected internal image of the sewage pipe into the lightweight pipe defect detection model to perform pipe defect detection.
2. The sewage pipe defect detection method based on multi-scale image feature fusion according to claim 1 is characterized in that: In step S2, preprocessing includes graying, contrast enhancement and denoising, and data expansion includes: random cropping, horizontal, vertical and random rotation of the historical sewage pipe defect image dataset.
3. The sewage pipe defect detection method based on multi-scale image feature fusion according to claim 2 is characterized in that: The improved YOLOv8 model includes: C2f-FAM module, HS-BiFPN module and Upsample module; among them, the C2f-FAM module introduces EMSConvP multi-scale convolution for efficient extraction of features of different scales; the HS-BiFPN module replaces the original neck of the YOLOv8 model, and realizes the flexible fusion of feature maps of different scales through cross-level feature fusion operation; through the upsampling of the DySample module, the scale of the feature map is dynamically optimized according to the size of the target and the complexity of the scene.
4. A sewage pipe defect detection device based on multi-scale image feature fusion, characterized in that: include: An acquisition module, used to acquire a historical sewage pipe defect image dataset; A processing module is used to preprocess and expand the historical pipeline defect image data set; A training module is used to train an improved YOLOv8 model based on the historical sewage pipe defect image dataset after preprocessing and data augmentation to obtain a lightweight pipeline defect detection model; The detection module is used to input the real-time collected internal images of the sewage pipe into the lightweight pipeline defect detection model for pipeline defect detection.
5. The sewage pipe defect detection device based on multi-scale image feature fusion according to claim 4 is characterized in that: Preprocessing includes grayscale, contrast enhancement and denoising, and data expansion includes random cropping, horizontal, vertical and random rotation of the historical sewage pipe defect image dataset.
6. The sewage pipe defect detection device based on multi-scale image feature fusion according to claim 5 is characterized in that: The improved YOLOv8 model includes: C2f-FAM module, HS-BiFPN module and Upsample module; among them, the C2f-FAM module introduces EMSConvP multi-scale convolution for efficient extraction of features of different scales; the HS-BiFPN module replaces the original neck of the YOLOv8 model, and realizes the flexible fusion of feature maps of different scales through cross-level feature fusion operation; through the upsampling of the DySample module, the scale of the feature map is dynamically optimized according to the size of the target and the complexity of the scene.
Citation Information
Patent Citations
YOLOv8 target detection method based on attention mechanism and multi-scale feature fusion
CN116883801A
Drainage pipeline defect detection method based on improved YOLOv8s
CN119169378A
Improved YOLOv8 chip surface defect lightweight target detection model and training method and application thereof
CN119399143A