CAM generation method for sewer line defect detection
The CAFAL-CAM algorithm, which combines class-aware attention fusion and adversarial learning to generate CAMs, solves the problems of missing global semantic information and inaccurate boundary localization in sewer pipe defect detection. It achieves high-precision semantic segmentation and boundary localization, thereby improving the overall performance of the CAM generation algorithm.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-17
AI Technical Summary
Existing CAM generation algorithms suffer from problems such as missing global semantic information, insufficient segmentation accuracy, and inaccurate boundary positioning in sewer pipe defect detection. They are particularly difficult to meet engineering requirements under complex lighting conditions, high foreground-background similarity, and significant differences in defect scale.
We propose a CAM generation algorithm CAFAL-CAM based on class-aware attention fusion and adversarial learning. By combining the class-aware attention fusion CAM generation algorithm CAF-CAM and the adversarial learning-based CAM collaborative optimization mechanism ALCO-CAM with an efficient loss function, we enhance the decoupling ability of defect foreground and background and the segmentation ability of small-scale defects. Furthermore, we improve the boundary localization accuracy through adversarial learning.
It significantly improves the semantic segmentation accuracy and boundary localization accuracy of CAM. The experimental results show that the mIoU index on the sewer-ML dataset of sewer pipe defects reaches 72.8%, which is 2.2% higher than the existing algorithm, and the training time is reduced by 30%.
Smart Images

Figure CN121883785A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of weakly supervised semantic segmentation, and more particularly to CAM generation, specifically a CAM generation method for detecting defects in sewer pipes. Background Technology
[0002] With the widespread application of computer vision and deep learning-based defect detection technologies in the field of sewer pipe defect detection, semantic segmentation methods have demonstrated significant advantages due to their pixel-level resolution capabilities. They can not only simultaneously identify and locate multiple types of defects, but also extract geometric feature parameters of defect areas, providing data support for subsequent evaluation. Compared to traditional fully supervised semantic segmentation methods, which are time-consuming and costly in the annotation process, existing weakly supervised detection methods, while significantly reducing annotation costs, are limited in practical applications by technical bottlenecks such as insufficient segmentation accuracy and difficulty in coordinating computational efficiency optimization. In the initial stage of weakly supervised semantic segmentation, namely the CAM generation stage, traditional CAM generation algorithms, due to the local receptive field characteristics and high-level feature abstraction mechanism of convolutional neural networks, cause class-aware attention to over-focus on the most discriminative local regions of the target, making it difficult to effectively capture the global semantic information of the defect. In the sewer pipe scenario, the limitations of traditional CAM generation algorithms are further amplified. Due to the complex lighting conditions, high similarity between foreground and background, significant differences in defect scale, and blurred defect contours in the sewer pipe scenario, traditional CAM generation algorithms struggle to meet engineering requirements in terms of defect area coverage completeness and boundary localization accuracy. Therefore, designing a CAM generation algorithm that combines high-precision segmentation, accurate boundary localization, and efficient computational performance for complex scenarios such as sewer pipes has become a key bottleneck in improving the practical value of weakly supervised semantic segmentation defect detection methods.
[0003] To address the issues of missing global semantic information and insufficient segmentation accuracy in traditional CAM generation algorithms, existing solutions explore different technical paths to improve CAM quality. Jiang et al. proposed an online attention accumulation strategy, expanding target coverage by integrating class-aware attention maps from different training stages. However, this method has limited adaptability to multi-scale defects in pipe environments and fails to effectively solve the problems of region omission due to scale changes and insufficient segmentation accuracy for small-scale defects. Sun et al. proposed introducing out-of-distribution data (ODD) to construct a cross-language image matching framework, combining a channel-space dual-dimensional attention mechanism to alleviate the problem of target over-focusing. However, their out-of-distribution data training paradigm struggles to model the unique background noise features of the sewer pipe scene due to the high similarity between target defects and the background, and its sensitivity to small-scale defects remains insufficient. Li et al. proposed a progressive patch learning method, effectively improving the target coverage integrity of CAM by strengthening the classification network's perception of local textures in stages. Kweon et al. utilized the structural prior knowledge of the SAM model to constrain the CAM generation process, effectively improving CAM quality. However, while these two methods can improve the ability to capture local details, they still suffer from semantic confusion and missegmentation of boundary regions when dealing with the low-contrast defects and blurred backgrounds in sewer pipe inspection, exposing the bottleneck of the generalization ability of existing CAM generation algorithms to complex industrial scenes. Summary of the Invention
[0004] The purpose of this invention is to propose a CAM generation algorithm based on class-aware attention fusion and adversarial learning to solve the problems of insufficient semantic segmentation accuracy and boundary localization accuracy in sewer pipe defect detection.
[0005] The objective of this invention is achieved as follows:
[0006] To address the shortcomings of existing CAM generation algorithms in foreground-foreground decoupling and small-scale defect segmentation, this invention first proposes a class-aware attention fusion CAM generation algorithm, CAF-CAM. By extracting and fusing foreground-foreground and multi-scale class-aware attention, it highlights salient region features while avoiding excessive suppression of insignificant semantic regions, enhancing the model's perception of local details. The collaborative optimization of the two class-aware attention modules achieves class-aware attention fusion and semantic information complementarity, enhancing the global representation capability of CAM and effectively improving CAM semantic segmentation accuracy. Subsequently, to address the insufficient defect boundary segmentation capability of existing CAM generation algorithms, this invention designs an adversarial learning-based CAM collaborative optimization mechanism, ALCO-CAM. Through dynamic adversarial game and positive feedback mechanisms between the classifier and reconstructor, it enhances the boundary segmentation capability of the classification network, further improving the boundary localization accuracy of CAM. Finally, this invention designs efficient and targeted loss functions for CAF-CAM and ALCO-CAM, respectively, and balances the semantic integrity of CAF-CAM with the boundary accuracy of ALCO-CAM through a joint loss function, providing guidance for the model while achieving collaborative optimization of overall performance.
[0007] The specific method is as follows:
[0008] A CAM generation method for sewer pipe defect detection, CAFAL-CAM, includes the following steps:
[0009] Step 1: Data preprocessing. The original sewer pipe image is preprocessed with illumination enhancement and scale transformation. Enhanced and weakened images are generated by adjusting contrast and brightness to improve the identification of defect areas. Small-scale, original-scale, and large-scale images are generated to adapt to multi-scale defect features.
[0010] Step 2: Feature Enhancement. The Class-Aware Attention Fusion (CAF-CAM) generation algorithm employs a dual-channel parallel architecture to achieve class-aware attention fusion and semantic information complementarity: The image enhancement-based class-aware attention fusion module guides the model to focus on the complete defect region through image enhancement strategies, achieving class-aware attention decoupling and re-fusion between the defect foreground and background. This highlights salient region features while avoiding excessive suppression of non-salient semantic regions, thereby obtaining accurate and sufficient semantic information. The self-supervised multi-scale class-aware attention fusion module uses a multi-scale self-supervised training mechanism to enhance the model's ability to perceive local details through cross-scale class-aware attention fusion, improving the accuracy of small-scale defect segmentation. The two modules run in parallel and interact with features through a class-aware attention collaborative optimization strategy to achieve class-aware attention fusion and semantic information complementarity.
[0011] Step 3: Adversarial Training Mechanism. The ALCO-CAM collaborative optimization mechanism, based on adversarial learning, constructs an adversarial game framework between the classifier and the reconstructor. In this mechanism, the classification network in CAF-CAM is used as the classifier, and an adversarial game loop is built by introducing a reconstructor: the reconstructor continuously improves its cross-segment feature reconstruction capability to optimize its performance by minimizing the difference loss between the reconstructed image segment and the original image segment; the classifier, on the other hand, forces itself to generate a CAM with a clear boundary response by maximizing the difference loss between the reconstructed image segment and the original image segment, using an adversarial gradient backpropagation mechanism. During this process, the classifier and reconstructor use an alternating iterative optimization strategy to update parameters: the reconstructor updates parameters first in each training round, continuously learning cross-segment feature reconstruction capability by minimizing feature reconstruction error, and iteratively enhancing the spatial perception accuracy of defect boundaries; subsequently, the classifier adjusts the network parameters based on the adversarial objective function, gradually improving the boundary localization accuracy.
[0012] Step 4: Quantify the impact of scene differences through joint loss function, determine the optimal value of parameters based on parameter sensitivity experiments, balance the impact of illumination and scale differences on the model, and adaptively adjust the model's response intensity to different pipeline environments (such as material and defect type).
[0013] Step 5: Adaptive fine-tuning for new scenarios. Train CAFAL-CAM on the Sewer-ML dataset to obtain basic model parameters. Use a small amount of data from the new scenario to update the classifier parameters in the adversarial learning mechanism to improve the model's adaptability to new scenarios. By dynamically balancing semantic integrity and boundary accuracy through joint loss, the model's generalization ability in different pipeline environments can be improved.
[0014] The positive effects of this invention are:
[0015] This invention addresses the shortcomings of existing CAM generation algorithms in terms of semantic segmentation accuracy and boundary localization accuracy for pipeline defects in sewer pipe scenarios. It proposes a CAM generation algorithm, CAFAL-CAM, based on class-aware attention fusion and adversarial learning. First, to address the deficiencies in foreground-background decoupling and small-scale defect segmentation capabilities of existing CAM generation algorithms, the class-aware attention fusion CAM generation algorithm CAFAL-CAM is proposed. This algorithm improves foreground-background decoupling and small-scale defect segmentation capabilities by extracting and fusing foreground-background and multi-scale class-aware attention. Then, through the collaborative optimization of the two modules, class-aware attention fusion and semantic information complementarity are achieved, enhancing the global representation capability of CAM and thus effectively improving the semantic segmentation accuracy of CAM. Subsequently, to address the insufficient defect boundary segmentation capability of existing CAM generation algorithms, an adversarial learning-based CAM collaborative optimization mechanism, ALCO-CAM, is proposed. This mechanism enhances the boundary segmentation capability of the classification network through dynamic adversarial game and positive feedback mechanisms between the classifier and the reconstructor, thereby further improving the boundary localization accuracy of CAM. Finally, this invention designs efficient and targeted loss functions for CAF-CAM and ALCO-CAM respectively, and balances the semantic integrity of CAF-CAM and the boundary accuracy of ALCO-CAM through a joint loss function, providing guidance for the model while achieving synergistic optimization of overall performance. Experimental verification shows that the CAFAL-CAM algorithm achieves an mIoU index of 72.8% on the Sewer-ML sewer defect dataset, achieving the best performance compared to existing representative algorithms. Specifically, CAFAL-CAM not only surpasses the best existing representative algorithm with a 2.2% mIoU accuracy advantage, but also reduces training time by 30%, achieving a better balance between accuracy and training efficiency. Attached Figure Description
[0016] Figure 1 It is the overall architecture of the class-aware attention fusion module based on image enhancement.
[0017] Figure 2 It is the overall architecture of the self-supervised multi-scale class perception attention fusion module.
[0018] Figure 3 This is a step in the collaborative optimization strategy of perception and attention.
[0019] Figure 4 It is the overall architecture of the CAM collaborative optimization mechanism based on adversarial learning.
[0020] Figure 5 It is λ s and λ bs Results of parameter sensitivity experiments.
[0021] Figure 6 It is λms Results of parameter sensitivity experiments.
[0022] Figure 7 It is λ fuse Results of parameter sensitivity experiments.
[0023] Figure 8 yes and Results of parameter sensitivity experiments.
[0024] Figure 9 This is the overall CAFAL-CAM process.
[0025] Figure 10 This is a comparison of the time consumption of different CAM generation algorithms.
[0026] Figure 11 This is a performance comparison of different CAM generation algorithms. Detailed Implementation
[0027] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0028] Step 1: Data preprocessing. The original sewer pipe image undergoes illumination enhancement and scale transformation preprocessing. Enhanced images are generated by adjusting contrast and brightness to improve the identification of defect areas. Small-scale, original-scale, and large-scale images are generated to adapt to multi-scale defect features. For a single pixel (i,j) in the image, the adjustment formula for the new pixel value g′(i,j) is as follows:
[0029] g'(i,j)=αg(i,j)+β
[0030] Enhanced images are generated using parameters α = 1.5 and β = 300, while weakened images are generated using parameters α = 0.8 and β = 150.
[0031] Step 2: Feature enhancement. The class-aware attention fusion (CAF-CAM) generation algorithm proposed in this invention is used to achieve class-aware attention fusion and semantic information complementarity.
[0032] Figure 1 The overall architecture of the image-enhanced class-aware attention fusion module is demonstrated.
[0033] The image-enhanced class-aware attention fusion module guides the model to focus on the complete defect region through image enhancement strategies. This achieves class-aware attention decoupling and re-fusion between the defect foreground and background, highlighting salient features while avoiding excessive suppression of insignificant semantic regions, thus obtaining accurate and sufficient semantic information. After obtaining the preprocessed pipeline image with enhanced illumination, a classification network generates CAMMo, Me, and Md corresponding to the original image, enhanced image, and weakened image, respectively. Enhancing the original image makes the distinction between instance regions more obvious, strengthens the classification network's ability to perceive morphological edges, and allows the network to learn target region and boundary information more clearly. Weakening the original image suppresses the activation intensity of dominant salient regions and activates suppressed insignificant semantic regions. Therefore, enhanced CAMMe has a high activation score in the most identifiable region and boundary of the target, while weakened CAMMd activates other regions important for semantic segmentation. The combination of these two CAMs can fully obtain the semantic information required for semantic segmentation tasks.
[0034] To fuse the enhanced CAMMe and weakened CAMMd, the confidence scores of the points in the fused CAMMf need to be recalculated. For a CAM generated by a classification network, the formula for calculating the confidence score P(x) of point x(i,j) belonging to class C can be expressed as:
[0035] P(χ)=∑ k w k f k (χ)
[0036] In the formula: f k (χ) represents the activation value of channel k at position x in the last convolutional layer; w k This indicates that f is obtained through global average pooling and the SoftMax function. k (χ) The weights generated.
[0037] This leads to the determination of increased confidence in the midpoints of CAMMe and decreased confidence in CAMMd. Me enhances the weight of high-confidence regions, highlighting image features that significantly contribute positively to classification decisions; while Md weakens high-confidence regions to avoid excessive suppression of non-significant semantic regions, activating other regions important for semantic segmentation. In this process, the confidence score directly quantifies the importance of local image regions to the model's classification: regions with higher confidence scores indicate a greater contribution of their contained visual patterns to the current category prediction.
[0038] To calculate the confidence scores of points in the fused CAMMf, a dynamic confidence fusion strategy is proposed, fully utilizing the spatial complementarity of Me and Md. Me tends to highlight salient regions of model classification dependence, while Md reveals the distribution of neglected latent features through inverse suppression. Based on this, the confidence scores of points in Mf are calculated using the confidence scores of points in Me and Md. The calculation process of the final confidence score Pf(x) for each spatial location x(i,j) in Mf belonging to class C needs to consider the following four probability scenarios:
[0039] (1) Enhance CAM points and points that reduce CAM Both belong to class C. At this point, point x is a high-response region of the dual CAM system. The maximum activation value should be retained to enhance the target core. Therefore, the confidence level P for point x corresponding to the fused CAM system belonging to class C is... f The formula for calculating (χ) can be expressed as:
[0040] P f (χ)=max(P e (χ),P d (χ))
[0041] In the formula: P e (χ) represents the confidence level that point x belongs to class C in the enhanced CAM; P d (χ) represents the confidence level of point x belonging to class C in the weakened CAM.
[0042] (2) Enhance CAM points Or weaken the CAM point It belongs to class C, and P e (χ) or P d When the maximum value of (χ) is greater than 0.5, point x is a significant activation region of a single CAM. Using the trust propagation mechanism, the confidence P of point x corresponding to the fused CAM belonging to class C is given. f (χ) The calculation formula is the same as that in equation (1-3), taking P e (χ) and P d The maximum value in (χ).
[0043] (3) Enhance CAM points Or weaken the CAM point It belongs to category C, but P e (χ) and P d The values of (χ) are all less than or equal to 0.5. In this case, it is necessary to suppress the semantic noise caused by low-confidence spurious activations. Therefore, the confidence P of the CAM corresponding point x belonging to class C is fused. f (χ) is set to 0.
[0044] (4) Enhance CAM points and points that reduce CAM Neither of them belongs to class C. At this point, point x is the background region of the dual CAM consensus, and it needs to be strictly zeroed to avoid misjudgment. Therefore, the confidence P of point x corresponding to the fused CAM belonging to class C is... f (χ) is set to 0.
[0045] This dynamic confidence fusion strategy effectively filters out low-confidence noise responses and suppresses false activation interference from single CAMs. Under the confidence-guided fusion strategy, the fused CAM serves as an adaptive spatial supervision signal for the classification network. This not only drives the network to model the target structure and edge features at a higher spatial resolution but also activates potential discriminative regions suppressed by conventional CAMs through a backpropagation compensation mechanism. This supervision mode doubly optimizes the feature learning process: on the one hand, it reduces the model's overfitting tendency to local high-response regions through region confidence comparison fusion; on the other hand, it forces the network to mine discriminative semantic cues across layers by utilizing the different class-aware attention distributions of Me and Md. Even without pixel-level labels as semantic segmentation supervision, more accurate and comprehensive semantic information can be obtained.
[0046] The loss function designed for the image enhancement-based class-aware attention fusion module is as follows:
[0047] (1) Classification Loss: Based on a weakly supervised learning framework using image-level labels, this invention selects ResNet101 as the classification backbone network. For supervised classification tasks, binary cross-entropy loss is used for training. The input data includes the original image and the image after image enhancement. The classification loss L... cls It can be represented as:
[0048]
[0049] In the formula: N represents the total number of samples; C represents the total number of categories; y i,c This represents the true label value of the i-th sample in the c-th category, where a value of 1 indicates the true category and a value of 0 indicates a false category; z i,c Let z represent the model's original output value for the i-th sample in the c-th class; σ(·) represents the Sigmoid function, which modifies the model's original output value z. i,c Mapping to the [0,1] interval yields the probability of the i-th sample in the c-th class predicted by the model.
[0050] (2) Similarity Loss: Adjusting image contrast and brightness does not change the semantic information of the image. Therefore, when projecting onto the feature space, the CAM Mo generated from the original image, the CAM Me generated from the enhanced image, and the CAM Md generated from the weakened image should maintain consistency in the feature space. To constrain feature alignment and improve the robustness of the model to illumination changes, this invention proposes a similarity loss L. sThe network is trained using L1 loss, and the similarity loss is L... s The definition is as follows:
[0051]
[0052] In the formula: / · / 1 represents L1 loss.
[0053] (3) Background Similarity Loss: A CAM is generated by a classification network and iteratively optimized based on the classification loss of the foreground object. However, since training is based on image-level labels, the semantic information of the background region is suppressed due to the lack of pixel-level supervision, inevitably ignoring the semantic information of many pixels, which are treated as background pixels and their activation values are set to zero. This makes it impossible to learn feature representations by generating gradients through backpropagation, resulting in the neglect of information provided by a large number of background pixels. Therefore, to mine the latent semantics of the background region and alleviate the gradient vanishing problem, this invention proposes a background similarity loss L... bs Background segmentation is performed on the CAMMo generated from the original image and the fused CAMMf, and background activation maps M are extracted from each. ob and M fb The L1 loss is used for background similarity loss, where L is the background similarity loss. bs The definition is as follows:
[0054] L bs = / M ob -M fb / 1
[0055] In the formula: / · / 1 represents L1 loss.
[0056] By jointly optimizing the classification loss, similarity loss, and background similarity loss, the overall loss L used to optimize network weights in the image-enhanced class-aware attention fusion module is optimized. IECAF The definition is as follows:
[0057] L IECAF =L cls +λ s L s +λ bs L bs
[0058] In the formula: λ s and λ bs The hyperparameter representing the overall loss balance; L cls The driving model locates the target region; L s Used to constrain feature alignment and improve the model's robustness to illumination changes; L bs Uncover overlooked contextual information by leveraging the consistency of background features.
[0059] The self-supervised multi-scale class-aware attention fusion module proposed in this invention employs a multi-scale self-supervised training mechanism. It enhances the model's ability to perceive local details and improves the accuracy of small-scale defect segmentation through cross-scale class-aware attention fusion. The two modules run in parallel and interact with each other through a class-aware attention collaborative optimization strategy, achieving class-aware attention fusion and semantic information complementarity.
[0060] Figure 2 The overall architecture of the self-supervised multi-scale class perception attention fusion module is demonstrated.
[0061] In the model initialization phase, a two-stage training strategy is adopted. First, the student branch is pre-trained. Considering the performance and training efficiency of existing weakly supervised semantic segmentation methods based on image-level labels, this invention introduces the EPS method to initialize and pre-train the student branch using image-level labels. Then, the teacher branch is initialized by copying the pre-trained student branch parameters to the teacher branch, enabling the teacher branch to have preliminary CAM generation capabilities.
[0062] In the self-supervised multi-scale class-aware attention fusion module, the teacher branch acts as a supervisory signal generator, providing guidance to the student branch. Its core task is to provide the student branch with robust CAM as pseudo-labels for learning and training through multi-scale class-aware attention fusion and channel-aware recalibration strategies. In the teacher branch, the original image is first preprocessed by scale transformation to obtain a small-scale image I. s Original Image I o and large-scale images I l CAMMs of corresponding scales are generated through a classification backbone network. s M o and M l Then, multi-scale CAM is fused to integrate complementary information. Multi-scale fusion of the k-th channel. The calculation formula is as follows:
[0063]
[0064] In the formula: This represents the small-scale CAM of the k-th channel; The original scale (CAM) of the k-th channel is represented. This represents the large-scale CAM of the k-th channel.
[0065] Since CAMs vary in size at different scales, the CAMs at different scales are adjusted to match the original scale CAMs before fusion. o The same size. To constrain the range of attention scores to the [0,1] interval, this invention... The maximum value is used to normalize the k-th channel to eliminate the influence of scale differences on activation intensity and ensure consistency of numerical range. The normalized k-th channel is then used for multi-scale fusion CAM. The calculation formula is as follows:
[0066]
[0067] In the formula: C represents the number of output channels, and its value is equal to the total number of categories.
[0068] However, if certain regions are activated only in a single-scale CAM, then local inactive regions may exist in the multi-scale fused CAM, which is detrimental to the training of student branches. To address this issue, this invention designs a channel-aware recalibration strategy to refine the multi-scale fused CAM, combining image-level labels for inter-channel denoising. When category k does not appear in label y, the value of the corresponding channel in the multi-scale fused CAM is forcibly set to zero, then the activation value in the background channel is set to a threshold of 0.2, and finally, the non-zero channels are enhanced. In the k-th channel, the value of pixel i after processing by the channel-aware recalibration strategy is... The calculation formula is as follows:
[0069]
[0070] In the formula: M represents the k-th channel. f_norm original activation value of middle pixel i
[0071] Thus, a multi-scale fusion CAMM' optimized by the channel-aware recalibration strategy was obtained. f_norm It integrates semantic information from different scales. Finally, it utilizes M' f_norm This is used to supervise the network's response to single-scale images, helping the model overcome biases in dealing with single-scale images, enhance the collaborative perception of local details and global structure, and thus capture more complete semantic regions.
[0072] For the self-supervised multi-scale class-aware attention fusion module, the loss function designed in this invention is as follows:
[0073] (1) Classification Loss: The module uses ResNet101 as the classification backbone network and employs binary cross-entropy loss for training. The input data includes the original image and the image after scale transformation. Classification loss L cls It can be represented as:
[0074]
[0075] In the formula: N represents the total number of samples; C represents the total number of categories; y i,cThis represents the true label value of the i-th sample in the c-th category, where a value of 1 indicates the true category and a value of 0 indicates a false category; z i,c Let z represent the model's original output value for the i-th sample in the c-th class; σ(·) represents the Sigmoid function, which modifies the model's original output value z. i,c Mapping to the [0,1] interval yields the probability of the i-th sample in the c-th class predicted by the model.
[0076] (2) Multi-scale attentional consistency loss: Multi-scale fusion CAMM' obtained from the teacher branch f_norm It contains complementary information from CAMs at different scales, utilizing M' f_norm Using pseudo-labels to supervise the network's response to single-scale images can help the model overcome biases related to single-scale images, enhance the co-perception of local details and global structure, and thus capture more complete semantic regions. To constrain single-scale CAM and multi-scale fusion CAMM... f_norm To ensure consistency in pixel-level spatial distribution, this invention employs mean squared error loss and defines a multi-scale attention consistency loss L. ms as follows:
[0077]
[0078] In the formula: N represents the total number of samples; C represents the total number of categories; M' represents the i-th sample, k-th class channel, spatial location (h, w). f_norm The value; This represents the CAM value at spatial location (h, w) of the i-th sample, k-th category channel, output by the student branch.
[0079] By jointly optimizing the classification task and aligning with multi-scale features, the overall loss L used to optimize network weights in the self-supervised multi-scale class-aware attention fusion module is optimized. MSACAF The definition is as follows:
[0080] L MSCAF =L cls +λ ms L ms
[0081] In the formula: λ ms The hyperparameter representing the overall loss balance is used to adjust the intensity of multi-scale supervision; L cls Used to drive the model to locate the target region; L ms Used to force cross-scale feature alignment and improve the integrity of small object segmentation.
[0082] To fully utilize the complementary semantic information of the image-enhanced class-aware attention fusion module and the self-supervised multi-scale class-aware attention fusion module, this invention proposes a class-aware attention collaborative optimization strategy to achieve class-aware attention fusion, enhance the global representation capability of CAM, and thus effectively improve the semantic segmentation accuracy of CAM.
[0083] Figure 3 The steps of the class-aware attention collaborative optimization strategy are demonstrated.
[0084] In the class-aware attention collaborative optimization strategy, two modules achieve class-aware attention fusion and semantic information complementarity through a parallel collaborative optimization architecture, improving the ability of the generated CAM to capture global semantic information. Furthermore, this invention designs a joint loss function for the class-aware attention collaborative optimization strategy, constraining the consistency of the outputs of the two modules and achieving a balance in feature contributions. This improves the foreground-background decoupling capability while enhancing the segmentation accuracy of small-scale defects, thereby strengthening the global semantic representation capability.
[0085] A joint loss function L is proposed for the class-aware attention fusion CAM generation algorithm. fuseecls The definition is as follows:
[0086]
[0087] In the formula: L IECAF L represents the loss of the class-aware attention fusion module based on image enhancement; MSCAF For the loss of the self-supervised multi-scale class-aware attention fusion module; M IECAF CAM generated by an image enhancement-based class-aware attention fusion module; M MSCAF CAM generated for a self-supervised multi-scale class-aware attention fusion module; λ fuse The hyperparameters controlling the consistency loss weights are used to constrain the alignment of the CAM spatial distributions of the two modules; / · / 1 represents the L1 loss.
[0088] Joint loss function L fusecls By combining the losses of the two modules, the consistency of their outputs is constrained and the feature contributions are balanced, thereby improving the overall performance.
[0089] Step 3: Adversarial training mechanism. The ALCO-CAM collaborative optimization mechanism based on adversarial learning proposed in this invention constructs an adversarial game framework between the classifier and the reconstructor.
[0090] Figure 4 The overall architecture of the CAM collaborative optimization mechanism based on adversarial learning is presented.
[0091] In the CAM collaborative optimization mechanism based on adversarial learning, the classifier and reconstructor are jointly trained through a feature decoupling mechanism, where the reconstructor consists of a feature encoder and a decoder. First, the generated CAM∈R is obtained from the classifier. n ×H×W The extracted spatial features X∈R are obtained from the feature encoder. d×H×W , where d represents the dimension of the feature channels. Then, a target category is randomly selected from the set of semantic categories contained in the image instance, and its corresponding activation mask is used as the target CAM. The target CAM is then used to decouple feature X into target fragment X. t Non-target fragment X nt X t and X nt Its mathematical form is as follows:
[0092] X t =M t ⊙X
[0093] X nt =(1-M) t )⊙X
[0094] Where: M t The target CAM; ⊙ represents element-wise multiplication.
[0095] Due to the differential property of element-wise multiplication, the gradient can be obtained through the target segment X. t Or non-target fragment X nt Backpropagation occurs along two paths: propagation to the CAM parameters to optimize the classifier; and propagation to the feature encoder to update the feature representation. This differentiability establishes a mathematical connection between the parameter space and the features, thus supporting end-to-end joint training. This yields the decomposed target and non-target segments. In subsequent reconstructor and classifier update stages, based on the decomposed X... t and X nt The reconstructor learns cross-segment reconstruction capabilities by minimizing the reconstruction loss, thus optimizing its own performance; while the classifier learns to generate more accurate CAMs by maximizing the reconstruction loss to suppress cross-segment reconstruction.
[0096] During the reconstructor update phase, gradient optimization is performed only on the reconstructor's encoder-decoder network, while the classifier network remains frozen. In this phase, the loss signal is backpropagated to the reconstructor network through spatial features. The fragment X is then... t and X nt The input is fed into the decoder to obtain the reconstruction result. and For the target segment, minimize the reconstruction result within the non-target region. The difference between the input image I and the input image I is constrained using L1 loss. It can be represented as:
[0097]
[0098] In the formula: ⊙ represents the verification mask for non-target regions; ⊙ represents element-wise multiplication.
[0099] For non-target segments, minimize the reconstruction result within the target region. The difference between the input image I and the input image I is constrained using L1 loss. It can be represented as:
[0100]
[0101] In the formula: The symbol represents the verification mask for the target region; ⊙ represents element-wise multiplication.
[0102] The total loss function is achieved by weighted fusion of two types of constraints, and the total loss L for training the reconstructor is... RU for:
[0103]
[0104] In the formula: and This is a hyperparameter used to dynamically adjust the contribution weights of target and non-target segments in the loss function.
[0105] However, further experiments revealed that forcibly minimizing the reconstructor's loss leads to a severe feature decoupling overcompleteness problem, resulting in mode collapse. As the CAM segmentation accuracy improves, the target feature X... t With non-target features X nt The independence of the data is significantly enhanced, causing the residual information between segments to approach zero. In this situation, the reconstructor cannot complete the cross-region reconstruction task using residual cues. Instead, the reconstructor directly generates the original image by remembering the distribution of the training set to minimize the total loss of the reconstructor, rather than completing the reconstruction task based on the residuals. This phenomenon undermines the theoretical basis of the adversarial game mechanism. Even if the classifier outputs a high-precision CAM, the reconstructor can still achieve low reconstruction error through the memory effect, and the model gets trapped in a local optimum with spurious convergence.
[0106] To address the aforementioned issues, this invention proposes a random residual supply strategy. This strategy breaks the overcompleteness of feature decoupling through controlled noise injection, maintaining necessary correlations between segments and forcing the reconstructor to continuously rely on residual cues for cross-segment inference. Unlike features that are completely decomposed into target and non-target features, the random residual supply strategy synthesizes residuals for each segment and provides them to the features of another segment. In this way, the reconstructor can always utilize these residuals to reconstruct segments, thereby maintaining the ability to reconstruct using residuals while suppressing the reconstructor's memory bias.
[0107] The stochastic residual supply strategy defines a random space binary mask g∈[0,1]. H×W Its spatial structure is divided by an s×s grid. Each The cell value is either 0 or 1, independently sampled according to a Bernoulli distribution B(q), where q is the residual retention probability. Cross-segment residual signals are synthesized using a mask g, achieving fine-grained control over noise distribution while avoiding local overfitting. After processing with a random residual feeding strategy, the target features in the RU stage are... It can be represented as follows:
[0108]
[0109] In the formula: g⊙X nt This represents the residual in the non-target region.
[0110] Non-target features in the RU stage after processing by the stochastic residual supply strategy It can be represented as follows:
[0111]
[0112] In the formula: g⊙X t This represents the residual of the target region.
[0113] By employing a stochastic residual feeding strategy, the original decomposition features X are... t and X nt Replace with residual enhancement features and Synchronize the verification mask and The area is expanded to eliminate the influence of the noise injection region. (Expanded version) and It can be represented as:
[0114]
[0115] The stochastic residual feeding strategy maintains the minimum necessary correlation between segments by introducing controllable noise, thus breaking the overcomplete coupling of decoupled features and maintaining the reconstructor's responsiveness to segmentation defects through residual correlation. This dynamic balancing mechanism suppresses the reconstructor's tendency to overfit due to memory, ensuring the model's generalization performance in complex scenarios.
[0116] During the classifier update phase, only the classifier's parameters are optimized, while the reconstructor's weights remain fixed. The loss in this phase is backpropagated to the classifier only through the target CAM Mt, forcing the classifier to optimize M. t To improve accuracy, cross-segment reconstruction is suppressed. Similar to the reconstructor update phase, the reconstruction result is obtained first. and For the target segment, maximize the reconstruction result within the non-target region. The difference between the generated image M and the input image I is used to prevent the classifier from missing target regions. If the classifier generates M... t The target region was not fully covered, and the reconstructor was unable to recover the original image using the target fragment during the CU stage, resulting in... In the missed detection area The difference from I is significant; through backpropagation, the target fragment reconstruction loss forces the classifier to optimize M. t To reduce missed detections, L1 loss is used to implement constraints. It can be represented as:
[0117]
[0118] In the formula: The verification mask represents the non-target region; ⊙ represents element-wise multiplication; ||·||1 represents L1 loss; the negative sign indicates that minimizing the loss leads to maximizing the refactored framework result. The difference between the input image I and the input image I.
[0119] For non-target segments, maximize the reconstruction result within the target region. The difference between the reconstructor and the input image I is considered to prevent non-target fragments from being mixed with target features. If the classifier misclassifies a non-target region as a target region, the non-target fragments generated by the reconstructor will be refactored. In the false detection area Unable to reconstruct the original image, the non-target segment reconstruction loss penalizes this error, driving the classifier to suppress M. t Overactivation. L1 loss is used to implement constraints. It can be represented as:
[0120]
[0121] In the formula: The target region is represented by a verification mask; ⊙ represents element-wise multiplication; ||·||1 represents L1 loss; the negative sign indicates that minimizing the loss leads to maximizing the refactored framework result. The difference between the input image I and the input image I.
[0122] The classification loss L provided by the class-aware attention fusion CAM generation algorithm fusecls The total loss L for training the classifier CU for:
[0123]
[0124] In the formula: and This is a hyperparameter used to adjust the contribution weights of target and non-target segments in the loss function.
[0125] Step 4: Quantify the impact of scene differences through joint loss function, determine the optimal value of parameters based on parameter sensitivity experiments, balance the impact of illumination and scale differences on the model, and adaptively adjust the model's response intensity to different pipeline environments (such as material and defect type).
[0126] To explore the impact mechanism of different hyperparameter adjustments on the model's detection performance and to determine the optimal parameter combination, this invention designed a series of experiments based on the control variable method to systematically evaluate the impact of changes in key hyperparameters in the loss function on the model's performance indicators.
[0127] Figure 5 It showed λ s and λ bs Parameter sensitivity experiment results, Figure 6 It showed λ ms Parameter sensitivity experiment results, Figure 7 It showed λ fuse Parameter sensitivity experiment results, Figure 8 Showing and Results of parameter sensitivity experiments.
[0128] For the hyperparameters balancing the loss in the loss function, the optimal hyperparameter values for the model's segmentation ability are obtained through model parameter sensitivity experiments: In the experiments, λ in the image enhancement-based class-aware attention fusion module is set... s and λ bs The value is set to 0.5, and the λ value in the self-supervised multi-scale class-aware attention fusion module is set to 0.5. ms Set λ to 150 in the perceptual attention collaborative optimization strategy. fuse The value is 0.1, which is the setting value in the CAM collaborative optimization mechanism based on adversarial learning. and It is 0.5. and The values are 0.6 and 0.4.
[0129] Step 5: Adaptive fine-tuning for new scenarios. Train CAFAL-CAM on the Sewer-ML dataset to obtain basic model parameters. Use a small amount of data from the new scenario to update the classifier parameters in the adversarial learning mechanism to improve the model's adaptability to new scenarios. By dynamically balancing semantic integrity and boundary accuracy through joint loss, the model's generalization ability in different pipeline environments can be improved.
[0130] Figure 9 For the overall CAFAL-CAM process, Figure 10 The time consumption of different CAM generation algorithms is compared. Figure 11 The performance comparison of different CAM generation algorithms is shown.
Claims
1. A CAM generation method for detecting defects in sewer pipes, characterized in that, The method includes the following steps: Step 1: Data preprocessing. The original sewer pipe image is preprocessed with illumination enhancement and scale transformation. Enhanced and weakened images are generated by adjusting contrast and brightness to improve the identification of defect areas. Small-scale, original-scale, and large-scale images are generated to adapt to multi-scale defect features. Step 2: Feature Enhancement. The Class-Aware Attention Fusion (CAF-CAM) generation algorithm employs a dual-channel parallel architecture to achieve class-aware attention fusion and semantic information complementarity: The image enhancement-based class-aware attention fusion module guides the model to focus on the complete defect region through image enhancement strategies, realizing the decoupling and re-fusion of class-aware attention between the defect foreground and background. This highlights the features of salient regions while avoiding excessive suppression of non-salient semantic regions, thereby obtaining accurate and sufficient semantic information. The self-supervised multi-scale class-aware attention fusion module adopts a multi-scale self-supervised training mechanism, enhancing the model's ability to perceive local details and improving the accuracy of small-scale defect segmentation through cross-scale class-aware attention fusion. The two modules run in parallel and interact with features through a class-aware attention collaborative optimization strategy to achieve class-aware attention fusion and semantic information complementarity. Step 3: Adversarial Training Mechanism. The ALCO-CAM collaborative optimization mechanism, based on adversarial learning, constructs an adversarial game framework between the classifier and the reconstructor. In this mechanism, the classification network in CAF-CAM is used as the classifier, and an adversarial game loop is built by introducing a reconstructor: the reconstructor continuously improves its cross-segment feature reconstruction capability to optimize its performance by minimizing the difference loss between the reconstructed image segment and the original image segment; the classifier, on the other hand, forces itself to generate a CAM with a clear boundary response by maximizing the difference loss between the reconstructed image segment and the original image segment, using an adversarial gradient backpropagation mechanism. During this process, the classifier and reconstructor use an alternating iterative optimization strategy to update parameters: the reconstructor prioritizes updating parameters in each training round, continuously learning cross-segment feature reconstruction capability by minimizing feature reconstruction error, and iteratively enhancing the spatial perception accuracy of defect boundaries; subsequently, the classifier adjusts the network parameters based on an adversarial objective function, gradually improving the boundary localization accuracy. Step 4: Quantify the impact of scene differences through joint loss function, determine the optimal value of parameters based on parameter sensitivity experiments, balance the impact of illumination and scale differences on the model, and adaptively adjust the model's response intensity to different pipeline environments (such as material and defect type). Step 5: Adaptive fine-tuning for new scenarios. Train CAFAL-CAM on the Sewer-ML dataset to obtain basic model parameters. Use a small amount of data from the new scenario to update the classifier parameters in the adversarial learning mechanism to improve the model's adaptability to new scenarios. By dynamically balancing semantic integrity and boundary accuracy through joint loss, the model's generalization ability in different pipeline environments can be improved.
2. The CAM generation method for sewer pipe defect detection as described in the claim, characterized in that, The image enhancement-based perceptual attention fusion module in step 2 adopts a dynamic confidence fusion strategy to calculate the final confidence P of the fusion CAM Mf for the points x(i,j) in the enhanced CAM Me and the weakened CAM Md according to four probability scenarios f (x): (1) Enhance CAM points and points that reduce CAM Both belong to class C. At this point, point x is a high-response region of the dual CAM system. The maximum activation value should be retained to enhance the target core. Therefore, the confidence level P for point x corresponding to the fused CAM system belonging to class C is... f The formula for calculating (χ) can be expressed as: P f (x) = max(P e (x), P d (x)) where: P e (x) represents the confidence that point x belongs to class C in the enhanced CAM; P d (x) represents the confidence that point x belongs to class C in the reduced CAM (2) Enhance CAM points Or weaken the CAM point It belongs to class C, and P e (χ) or P d When the maximum value of (χ) is greater than 0.5, point x is a significant activation region of a single CAM. Using the trust propagation mechanism, the confidence P of point x corresponding to the fused CAM belonging to class C is given. f (χ) The calculation formula is the same as that in equation (1-3), taking P e (χ) and P d The maximum value in (χ), (3) Enhance CAM points Or weaken the CAM point It belongs to category C, but P e (χ) and P d The values of (χ) are all less than or equal to 0.
5. In this case, it is necessary to suppress the semantic noise caused by low-confidence spurious activations. Therefore, the confidence P of the CAM corresponding point x belonging to class C is fused. f (χ) is set to 0. (4) Enhance CAM points and points that reduce CAM Neither of them belongs to class C. At this point, point x is the background region of the dual CAM consensus, and it needs to be strictly zeroed to avoid misjudgment. Therefore, the confidence P of point x corresponding to the fused CAM belonging to class C is... f (χ) is set to 0.
3. The CAM generation method for sewer pipe defect detection as described in the claim, characterized in that, Step two's self-supervised multi-scale class-based perceptual attention fusion module processes the small-scale, original-scale, and large-scale CAM M generated by the teacher branch. s M o and M l Perform fusion, multi-scale fusion CAM of the k-th channel The calculation formula is as follows: To constrain the attention score to the [0,1] interval, by... The maximum value is used to normalize the k-th channel to eliminate the influence of scale differences on activation intensity and ensure numerical range consistency. The normalized k-th channel then undergoes multi-scale fusion CAM. The calculation formula is as follows: In the formula: C represents the number of output channels, and its value is equal to the total number of categories. To address the issue of localized inactive regions in multi-scale fused CAMs due to activation only in single-scale CAMs, a channel-aware recalibration strategy is employed to refine the multi-scale fused CAM. Image-level labels are used for inter-channel denoising. When category k is not present in label y, the corresponding channel value in the multi-scale fused CAM is forcibly set to zero. Then, the activation value in the background channel is set to a threshold of 0.
2. Finally, non-zero channels are enhanced. In the k-th channel, the value of pixel i after processing by the channel-aware recalibration strategy is... The calculation formula is as follows: In the formula: M represents the k-th channel. f_norm original activation value of middle pixel i Thus, a multi-scale fusion CAM M' is obtained through the channel perception recalibration strategy optimization f_norm which combines semantic information from different scales, and finally uses M' f_norm to supervise the network's response to single-scale images, helping the model overcome the bias to single-scale images and enhance the collaborative perception of local details and global structure, thereby capturing more complete semantic regions.
4. A CAM generation method for detecting defects in sewer pipes as described in the claim, characterized in that, The random residual supply strategy in step three first defines a random space binary mask g∈[0,1]. H×W Its spatial structure is divided into s×s grids, each The cell value is either 0 or 1, independently sampled according to a Bernoulli distribution B(q), where q is the residual retention probability. Cross-segment residual signals are synthesized using a mask g to finely control noise distribution while avoiding local overfitting. After processing with a random residual supply strategy, the target features in the RU stage are... It can be represented as follows: In the formula: g o X nt represents a non-target region residual Non-target features in the RU stage after processing by the stochastic residual supply strategy It can be represented as follows: In the formula: g⊙X t Represents the residual of the target region; synchronously verifies the mask. and The extended verification mask is obtained by expanding the mask to eliminate the influence of noise injection regions.
5. A CAM generation method for detecting defects in sewer pipes as described in the claim, characterized in that, In step four, the hyperparameter values for optimizing the model's segmentation ability are obtained through model parameter sensitivity experiments: λ in the image enhancement-based class-aware attention fusion module is set... s and λ bs The value is set to 0.5, and the λ value in the self-supervised multi-scale class-aware attention fusion module is set to 0.
5. ms Set λ to 150 in the perceptual attention collaborative optimization strategy. fuse The value is 0.1, which is the setting value in the CAM collaborative optimization mechanism based on adversarial learning. and It is 0.
5. and The values are 0.6 and 0.
4.
6. A CAM generation method for detecting defects in sewer pipes as described in the claim, characterized in that, Step four involves designing joint loss functions for each module. The loss function for the image enhancement-based class-aware attention fusion module includes: Classification loss: Similarity loss: Background similarity loss: L bs =|M ob -M fb / 1 Joint loss: L IECAF =L cls +λ s L s +λ bs L bs The loss function of the self-supervised multi-scale class-aware attention fusion module includes: Classification loss: Multi-scale attentional consistency loss: Joint loss: L MSCAF =L cls +λ ms L ms The joint loss function of the class-aware attention fusion CAM generation algorithm CAF-CAM is: In the ALCO-CAM collaborative optimization mechanism based on adversarial learning, the joint loss function of the reconstructor is: in The joint loss function for the classifiers is: in