A method for detecting defects of urban drainage pipes based on CCTV video

CN122067029BActive Publication Date: 2026-09-04ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610510493.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-17
Publication Date
2026-09-04
Estimated Expiration
2046-04-17

AI Technical Summary

Technical Problem

[0004]鉴于背景技术的不足,本发明所要解决的技术问题是提供一种基于CCTV视频的城市排水管道缺陷检测方法,该模型框架提出了更为合理的无监督学习策略,采用位置先验的方法增强模型对缺陷的识别能力和结果的可解释性,有效解决现有CCTV视频检测模型高成本数据标注和长时视频识别能力弱的问题

Benefits of technology

1、本申请的融合了多样化的自监督学习策略、特殊微调策略和两阶段推理模式,提出了更为合理的无监督学习策略,采用位置先验的方法增强模型对缺陷的识别能力和结果的可解释性,有效解决现有CCTV视频检测模型高成本数据标注和长时视频识别能力弱的问题,推动城市道路智慧化运维及检测,为城市道路的有机更新及城市安全提供坚实的理论基础与技术支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122067029B_ABST
    Figure CN122067029B_ABST
Patent Text Reader

Abstract

The application discloses a kind of urban drainage pipeline defect detection methods based on CCTV video, which framework includes: D2-VideoMAE encoder, based on double mask strategy and fog contrast learning strategy;Scan-TM, based on pre-training D2-VideoMAE encoder, a small amount of soft label CCTV video data is used for fine-tuning training;PP-Model, based on pre-training D2-VideoMAE encoder, a small amount of soft label CCTV video data and prompt construction are used for fine-tuning training, the performance and interpretability of the model are enhanced by using defect orientation priori, and finally the type and time range of defects in CCTV video can be output.The application effectively solves the problems of poor interpretability, low video recognition performance and complex use of existing pipeline defect recognition model, and provides an efficient and reliable technical solution for intelligent detection of urban underground drainage pipeline and safe operation and maintenance of urban infrastructure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an intelligent defect detection system for urban drainage pipelines, and more particularly to a defect detection method for urban drainage pipelines based on CCTV video. Background Technology

[0002] With the acceleration of urbanization, the scale and age of urban underground drainage pipes are gradually increasing, making CCTV robots equipped with high-definition cameras the preferred choice for detecting defects in drainage pipes. However, as the requirements for detection range and accuracy increase, the verification of CCTV videos from construction sites requires a significant amount of manpower and resources. Furthermore, the varying experience of different inspectors can lead to subjective results. For example, the invention patent with publication number CN103423596A discloses a method for detecting and evaluating drainage pipes using handheld video and closed-circuit television. This method uses handheld video for preliminary detection and calculates the degree of pipe deformation, pipe breakage / collapse, pipe joint misalignment, pipe protrusion / penetration, and pipe blockage based on the data. According to the calculation results, the pipes are classified according to standards, and corresponding operations are performed based on the classification results. This invention can effectively improve the efficiency of urban drainage pipe detection and evaluation in my country, reduce the workload of drainage pipe cleaning and closed-circuit television inspection, save manpower and resources, and effectively detect serious pipe defects, preventing accidents and ensuring the safe operation of urban drainage networks. However, while this existing drainage pipe detection and evaluation method improves detection efficiency, the improvement is limited, and the accuracy needs significant improvement.

[0003] Therefore, existing video processing and recognition methods for CCTV drainage pipe video recognition have several unresolved issues: First, understanding and recognizing video content requires complex models, which in turn demand a large amount of labeled data, making detailed video labeling extremely costly in terms of manpower and time. Second, existing models are weak in recognizing CCTV drainage pipe videos with a lot of blank information and cannot handle long video content. Clearly, current technologies still have significant shortcomings in terms of efficient CCTV video recognition and practicality. Summary of the Invention

[0004] In view of the shortcomings of the prior art, the technical problem to be solved by the present invention is to provide a method for detecting defects in urban drainage pipes based on CCTV video. The model framework proposes a more reasonable unsupervised learning strategy and adopts a location prior method to enhance the model's ability to identify defects and the interpretability of the results, effectively solving the problems of high cost data annotation and weak long-term video recognition ability of existing CCTV video detection models.

[0005] Therefore, the invention is achieved using the following technical solution: A method for detecting defects in urban drainage pipes based on CCTV video, including S1: The D²-VideoMAE encoder, based on a dual-masking strategy and a fog contrast learning strategy, uses a large amount of CCTV video footage of drainage pipes from multiple cities and countries for unsupervised learning to train a pre-trained encoder capable of perceiving pipe structures and defects in CCTV video. S2: Scan-TM, based on a pre-trained D²-VideoMAE encoder, is fine-tuned using a small amount of soft-labeled CCTV video data. This model employs a peak segmentation strategy to ultimately output suspected defect time segments from long CCTV videos with a large amount of blank background. S3: PP-Model, based on a pre-trained D²-VideoMAE encoder, uses a small amount of soft-labeled CCTV video data and prompt construction for fine-tuning training. It adopts a defect location prior to enhance the model's performance and interpretability, and can finally output the type of defect in CCTV video and the time range of its occurrence.

[0006] Preferably, step S1 includes: S11: Collects a large amount of CCTV video data of drainage pipes from different regions without any processing; S12: Simulate fog and noise to establish original video comparison samples for self-supervised comparison training of the model; S13: Establish encoder and decoder structures, and adopt a dual-mask learning strategy. Set up a large-area mask on the D²-VideoMAE encoder to block video content, while covering a small area on the decoder to restore the original video and the video data of the comparison sample. S14: Compare the real and predicted samples in the masked regions of the original video and the contrast sample. Adjust the model results using MAE loss, contrast loss and an auxiliary class loss to construct a closed-loop self-supervised learning strategy. The obtained pre-trained D²-VideoMAE encoder has the basic function of perceiving the structure and defect morphology of drainage pipes.

[0007] Preferably, step S2 includes: S21: Collect a small amount of CCTV video for soft labeling, only labeling the defect type and the approximate time period of occurrence (i.e., accurate to the second), for fine-tuning training of the model; S22: The Scan-TM model is built based on the pre-trained D²-VideoMAE encoder. The Scan-TM decoder adopts an attention mechanism structure similar to that of the encoder. S23: During the training process, the model freezes all parameters of the pre-trained D²-VideoMAE encoder, trains only the parameters of the Scan-TM encoder structure based on a small amount of soft-label data, outputs the probability of the existence of defects and the probability of their categories, and performs the corresponding loss calculation. S24: During the inference process, the CCTV video to be predicted will be analyzed by the Scan-TM model to predict the defects in each frame. Based on peak detection and segmentation algorithms, segments in the long CCTV video that are suspected of having a certain defect will be obtained for subsequent inference and identification.

[0008] Preferably, step S3 includes: S31: PP-Model is built based on the pre-trained D²-VideoMAE encoder. A Prompt parameter layer is overlaid on the encoder for fine-tuning. The detection head adopts a position prior strategy to enhance the performance and interpretability of the model. S32: During the training process, the model uses the PEFT strategy to freeze the first few layers of the D²-VideoMAE encoder structure. Based on a small amount of soft-label data, only the parameters of the last few layers of the encoder structure, the Prompt parameter, and the detection head are trained. Finally, the type of defect and the time range of occurrence are output, and the corresponding loss and location prior loss are calculated. S33: During the inference process, the suspected segments output by Scan-TM are input into PP-Model to ultimately determine the defect status of the segment. Based on the Laplacian variance sharpness score, a defect report is output, including: defect type, time range of occurrence in the video, and key representative frames.

[0009] Preferably, in step S12, fog and noise are artificially introduced into the original video to generate contrast samples. These samples are then used in the contrastive learning process during D²-VideoMAE pre-training. The fog and noise can be described as follows: (1) (2) in, I orig ( x , y () represents the original image. F ( x , y ) indicates a fog background (simulated as a non-uniform gray background). α Controlling the intensity of the fog, N (0, σ 2 () represents Gaussian noise; during data augmentation, the fog intensity and noise variance are randomly selected within a predefined range to enhance robustness and generalization ability; given an input video Then the comparison learning samples can be represented as .

[0010] Preferably, in step S13, inside the encoder, the encoder mask is constrained to a fixed spatial position to capture local appearance features and spatial continuity; the number of visible masks is used to... It means that among them The invisible mask ratio is 90%, which forces the model to learn a strong global representation from limited visible information; After 3D block input, input x i =[ x 1, x [2] The vectors are converted into token embeddings, category embeddings, and location embeddings. The invisible mask is processed by a decoder composed of ViViT-based multi-head attention (MHA) and multi-layer perceptron (MLP) to generate intermediate vectors. and categories ; In the decoder, a mask for invisible regions is extracted from random spatial locations using a conventional sampling and rolling offset strategy. This decoder, composed of a multi-head self-attention mechanism and a multilayer perceptron module, is specifically designed to reconstruct invisible regions, thereby learning temporal continuity features. The number of masks in the decoder is... The combined TOKEN sequence is Only for the invisible parts Z m Reconstruction will be carried out, and the reconstruction will be carried out as follows: This allows computation to be focused on a small subset of targets, thereby improving training efficiency.

[0011] Preferably, in step S14, the overall training loss of D²-VideoMAE consists of three parts: MAE reconstruction loss, contrastive learning loss, and auxiliary classification loss. The definition of MAE reconstruction loss is: (3) in, Y extra Represents the true value, while P = τ · patch · patch Contrastive learning loss is used to measure the similarity between contrastive pairs. Its purpose is to minimize the loss between positive samples (i.e., the loss between samples with and without fog) while maximizing the loss between negative pairs. (4) in,S It is by z 1 and z 2. The constructed similarity matrix, , diag_labels Then corresponding to S Labels of positive samples on the matrix diagonal; an auxiliary classification loss is introduced by using the encoder's CLS token, thus providing weak geometric supervision to mitigate errors caused by CCTV viewpoint drift: (5) (6) Using weighting coefficients λ = 0.2 and μ = 0.1; After extensive pre-training on unlabeled CCTV videos, the encoder acquired a powerful defect feature extraction capability, which was subsequently used to train Scan-TM and PP-model.

[0012] Preferably, in step S22, the Scan-TM is trained on the basis of a pre-trained D²-VideoMAE encoder, and the encoder parameters are frozen at this stage; using limited soft-labeled data, time-segment data of each defect occurrence is generated; Scan-TM uses a dedicated decoder to simulate spatial relationships based on multiple region labeling features of each frame and to estimate the probability of defect existence; the dedicated decoder consists of a multi-head self-attention mechanism (MHA) and a multilayer perceptron (MLP) module with the same encoder but fewer layers.

[0013] Preferably, in step S23, an input is given. The encoder will generate multiple category token embeddings. These embeddings are stacked into a sequence of categorical features. The sequence is projected and sinusoidal positional encoding is added, then embedded as a query input into the decoder to obtain... A sampling rate of 1 frame per second is used to reduce redundancy and computational overhead; the decoder output can be represented as: (7) The two prediction heads output the probability of the presence of the defect, respectively. and defect category probability ; The loss function for defect presence and coarse classification is defined as follows: (8) (9) In addition, a background (BG) class is added during training (but not during inference) to ensure clear background classification and suppress fluctuations in predictions. (10) This design can alleviate category imbalance and reduce background interference in long CCTV videos.

[0014] Preferably, in step S24, for a given original surveillance video sample, a low-rate temporal scan is first performed using uniformly sampled frames; each sampled frame is encoded using a trained Scan-TM model to generate per-frame probabilities for 16 defect categories; in order to convert noisy probability trajectories into discrete defect events, a category-level peak segmentation strategy is applied to the smooth probability curve of each category, specifically: values ​​below a removal threshold are considered background, peaks above a peak threshold are considered event centers, and adjacent peaks are merged or split based on whether adjacent peak valleys are below a removal threshold; for each coarse interval, an interval-level prior model is constructed using the average of the temporal posterior values ​​within that interval.

[0015] The beneficial effects of adopting the above technical solution are as follows: 1. This application integrates diverse self-supervised learning strategies, special fine-tuning strategies, and two-stage inference models, and proposes a more reasonable unsupervised learning strategy. It adopts a position prior method to enhance the model's ability to identify defects and the interpretability of results, effectively solving the problems of high-cost data annotation and weak long-term video recognition capabilities of existing CCTV video detection models. This promotes the intelligent operation and maintenance and detection of urban roads, and provides a solid theoretical foundation and technical support for the organic renewal of urban roads and urban safety.

[0016] 2. The proposed D²-VideoMAE encoder can obtain a pre-trained model with preliminary perception of the basic structure and defects of drainage pipes by using a large number of unlabeled CCTV videos based on a dual masking strategy and a contrastive learning strategy. At the same time, the contrastive learning method of adding positive fog samples can also help the model cope with more foggy scenarios.

[0017] 3. The proposed Scan-TM model can train the non-pre-trained model part using a small amount of soft-labeled data. Based on the peak segmentation algorithm, it can obtain slices that may produce defects from the original long CCTV video, reducing the recognition pressure of the subsequent model and increasing the recognition accuracy of the model in long CCTV videos.

[0018] 4. The proposed PP-Model has a Prompt structure. During the training phase, only the last 30% of the pre-trained encoder, the Prompt structure, and the detection head are trained. During inference, the specific type of defect and the precise boundary of the time range can be obtained from the slice of suspected defects. Attached Figure Description

[0019] The present invention includes the following figures: Figure 1 This is a schematic diagram of the process of the present invention; Figure 2 This provides the overall reasoning framework for the PipeMAE-TD model. Figure 3 for Figure 2 Detailed diagram of the coarse-to-medium forecasting steps; Figure 4 for Figure 2 Detailed diagram of the intermediate and detailed forecasting steps; Figure 5 This is a diagram illustrating the training principle of D²-VideoMAE. Figure 6 This is a diagram illustrating the training principle of Scan-TM. Figure 7 This is a diagram illustrating the training principle of PP-Model. Figure 8 The training and validation loss of the D²-VideoMAE pre-trained model; Figure 9 The training results of the D²-VideoMAE pre-trained model are shown in the image. Figure 10 The training and validation loss for Scan-TM; Figure 11 The training and validation loss for the PP-Model. Detailed Implementation

[0020] To further illustrate the technical means and effects of the invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structure, features and effects of the invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0021] For reference Figures 1-11 As shown, this application provides a method for detecting defects in urban drainage pipelines based on CCTV video (where CCTV is closed-circuit television), including... S1: The D²-VideoMAE encoder, based on a dual-masking strategy and a fog contrast learning strategy, uses a large amount of CCTV video footage of drainage pipes from multiple cities and countries for unsupervised learning to train a pre-trained encoder capable of perceiving pipe structures and defects in CCTV video. S2: Scan-TM (Scan-Aware Temporal Model), based on a pre-trained D²-VideoMAE encoder, is fine-tuned using a small amount of soft-labeled CCTV video data. This model employs a peak segmentation strategy to ultimately output suspected defect time segments from long CCTV videos with a large amount of blank background. S3: PP-Model (Prior-Guided Prompt-Based Model), based on a pre-trained D²-VideoMAE encoder, uses a small amount of soft-labeled CCTV video data and prompt construction for fine-tuning training. It adopts a defect location prior to enhance the model's performance and interpretability, and can finally output the type of defect in CCTV video and the time range of its occurrence.

[0022] The specific training steps for the D²-VideoMAE encoder in step S1 above are as follows: S11: Collect a large amount of CCTV video footage of drainage pipes from different regions; S12: Simulating fog and noise to establish contrast samples from the original video for self-supervised contrastive training of the model. The specific method is as follows: Due to the complex internal environment of drainage pipes, videos captured by CCTV often contain varying degrees of fog and noise. To address this issue, this invention generates contrast samples by artificially introducing fog and noise into the original video. These samples are then used in the contrastive learning process during D²-VideoMAE pre-training. These controllable disturbances are used as positive samples for contrastive learning, while reducing the adverse effects of environmental interference, thus improving the robustness of the model in a simple and effective way. Fog and noise can be described as: (1) (2) in, I orig ( x , y () represents the original image. F ( x , y ) indicates a fog background (simulated as a non-uniform gray background). α Controlling the intensity of the fog, N (0, σ 2 () represents Gaussian noise. During data augmentation, the fog intensity and noise variance are randomly selected within a predefined range to enhance robustness and generalization ability. Given an input video Then the comparison learning samples can be represented as .

[0023] S13: Establish the encoder and decoder structure, employing a dual-mask learning strategy. A large-area mask is set on the D²-VideoMAE encoder to occlude video content, while a small area is covered on the decoder to recover the original video and comparison sample video data. Specifically, within the encoder, the encoder mask is constrained to a fixed spatial location to capture local appearance features and spatial continuity. The number of visible masks is determined by... It means that among them The proportion of invisible masks can be as high as 90%, forcing the model to learn a strong global representation from limited visible information.

[0024] After 3D block input, input x i =[ x 1, x [2] is converted into token embeddings, category embeddings, and position embeddings. The invisible mask is processed by a decoder composed of ViViT-based multi-head attention (MHA) and multi-layer perceptron (MLP) to generate intermediate vectors. and categories .

[0025] In the decoder, masks of invisible regions are extracted from random spatial locations using a conventional sampling and rolling offset strategy. The decoder consists of a multi-head self-attention (MHA) mechanism and a multilayer perceptron (MLP) module, specifically designed to reconstruct invisible regions, thereby learning temporal continuity features. The number of masks in the decoder is... The combined TOKEN sequence is Only for the invisible parts Z m Reconstruction will be carried out, and the reconstruction will be carried out as follows: This allows computation to be focused on a small subset of targets, thereby improving training efficiency.

[0026] S14: Compare the real and predicted samples in the masked regions of the original video and the contrast sample. Adjust the model results using MAE loss, contrast loss, and an auxiliary class loss to construct a closed-loop self-supervised learning strategy. The resulting pre-trained D²-VideoMAE encoder has the basic function of perceiving the structure and defect morphology of drainage pipes. The specific method is as follows: The overall training loss of D²-VideoMAE consists of three parts: MAE reconstruction loss, contrastive learning loss, and auxiliary classification loss. The definition of MAE reconstruction loss is: (3) in, Yextra Represents the true value, while P = τ · patch · patch Contrastive learning loss is used to measure the similarity between contrastive pairs. Its purpose is to minimize the loss between positive samples (i.e., the loss between samples with and without fog) while maximizing the loss between negative pairs. (4) in, S It is by z 1 and z 2. The constructed similarity matrix, , diag_labels Then corresponding to S Labels of positive samples on the matrix diagonal. An auxiliary classification loss is introduced by using the encoder's CLS token, thus providing weak geometric supervision to mitigate errors caused by CCTV viewpoint drift. (5) (6) Using weighting coefficients λ = 0.2 and μ = 0.1.

[0027] After extensive pre-training on unlabeled CCTV videos, the encoder acquired a powerful defect feature extraction capability, which was subsequently used to train Scan-TM and PP-model.

[0028] The specific steps for training and inference of Scan-TM in step S2 above are as follows: S21: Collect a small amount of CCTV video for soft labeling, only labeling the defect type and the approximate time period of occurrence (i.e., accurate to the second), for fine-tuning training of the model; S22: Construct a Scan-TM model based on a pre-trained D²-VideoMAE encoder. The Scan-TM decoder adopts an attention mechanism structure similar to the encoder. The specific method is as follows: Scan-TM is trained on top of a pre-trained D²-VideoMAE encoder, whose parameters are frozen at this stage. Using limited soft-labeled data, temporal segments of data for each defect occurrence are generated. Scan-TM uses a dedicated decoder to model spatial relationships based on multiple region-labeled features in each frame and to estimate the probability of defect presence. This dedicated decoder consists of a multi-head self-attention mechanism (MHA) and a multilayer perceptron (MLP) module, which are identical to the encoder but have fewer layers.

[0029] S23: During training, the model freezes all parameters of the pre-trained D²-VideoMAE encoder. Based on a small amount of soft-labeled data, it trains only the parameters of the Scan-TM encoder structure, outputting the probability of defect presence and class probability, and performing corresponding loss calculations. The specific method is as follows: Given an input The encoder will generate multiple category token embeddings. These embeddings are stacked into a sequence of categorical features. The sequence is projected and sinusoidal positional encoding is added, then embedded as a query input to the decoder to obtain... A sampling rate of 1 frame per second is used to reduce redundancy and computational overhead. The decoder output can be represented as: (7) The two prediction heads output the probability of the existence of the defect, respectively. and defect category probability The loss function for defect presence and coarse classification is defined as follows: (8) (9) In addition, a background (BG) class is added during training (but not during inference) to ensure clear background classification and suppress fluctuations in predictions. (10) This design can alleviate category imbalance and reduce background interference in long CCTV videos.

[0030] S24: During the inference process, the CCTV video to be predicted is processed by the Scan-TM model to predict the defects in each frame. Based on peak detection and segmentation algorithms, segments in the long CCTV video suspected of having a certain defect are obtained for subsequent inference and identification. The specific method is as follows: For a given sample of raw surveillance video, we first perform a low-rate temporal scan using uniformly sampled frames (typically 1 frame / second). Each sampled frame is encoded using a trained Scan-TM model to generate per-frame probabilities for 16 defect categories.

[0031] To transform noisy probability trajectories into discrete defect events, we apply a category-level peak segmentation strategy to the smoothed probability curves for each category. Specifically, values ​​below a removal threshold are considered background, peaks above a peak threshold are considered event centers, and adjacent peaks are merged or split based on whether their adjacent valley values ​​are below a removal threshold. This strategy effectively separates recurring occurrences of the same defect type while avoiding over-segmentation under slight fluctuations. For each coarse interval, we construct an interval-level prior model using the average of the time posterior values ​​within that interval.

[0032] The specific steps for training and inference of the PP-Model in step S3 above are as follows: S31: PP-Model is built based on the pre-trained D²-VideoMAE encoder. A Prompt parameter layer is overlaid on the encoder for fine-tuning. The detection head adopts a position prior strategy to enhance the performance and interpretability of the model. S32: During training, the model uses the PEFT (Parameter-Efficient Fine-Tuning) strategy to freeze the first few layers of the D²-VideoMAE encoder. Based on a small amount of soft-label data, only the parameters of the last few layers of the encoder, the Prompt parameter, and the detector head are trained. Finally, the type and occurrence time range of the defect are output, and the corresponding loss and location prior loss are calculated. The specific method is as follows: Defect-centric video clips are input into the model, and tokens are generated through the encoder and Prompt module. and categories Linear layers and Softmax layers generate defect classification. Time boundary t s and t e and direction p θ Then, loss calculation is performed. For defect classification, a focus loss combined with label smoothing is used to emphasize difficult samples and minority classes: (11) in, This indicates the label after smoothing, with a smoothing factor of 0.05. γ = 2. Furthermore, entropy regularization loss is introduced to prevent mode collapse in the early stages: (12) The definition of time boundary loss is: (13) For direction prediction, a priori guidance strategy was employed. For each defect category...y The predefined encoding divides the 360° space into eight regions. This prior information from category to geometry helps stabilize defect classification and improves interpretability. The directional prior loss is defined as: (14) S33: During the inference process, the suspected segments output by Scan-TM are input into the PP-Model to ultimately determine the defect status of the segment. A defect report is output based on the Laplacian variance sharpness score, including: defect type, time range of occurrence in the video, and key representative frames. The specific method is as follows: After the original long CCTV video is processed by Scan-TM, a short, high-frame-rate segment around the specified interval is post-processed and refined using PP-Model. This segment model collectively predicts the defect type and orientation and generates start / end boundary distributions. Boundary refinement is only accepted if the boundary distribution is sufficiently reliable; otherwise, coarse boundaries are retained to ensure conservative behavior. Finally, event confidence is calculated by fusing time-level and segment-level posterior values, and a representative keyframe is selected for downstream inspection by maximizing the Laplacian variance sharpness score within the refined interval. The pipeline outputs a structured JSON report containing the defect type, refined timestamp, confidence, prior and posterior values, and the selected keyframe.

[0033] For reference Figures 1-11 As shown, the following describes in detail, through embodiments, a method for detecting defects in urban drainage pipes based on CCTV video provided in this application, including... S1: The D²-VideoMAE encoder, based on a dual-masking strategy and a fog contrast learning strategy, uses a large amount of CCTV video footage of drainage pipes from multiple cities and countries for unsupervised learning to train a pre-trained encoder capable of perceiving pipe structures and defects in CCTV video. S2: Scan-TM (Scan-Aware Temporal Model), based on a pre-trained D²-VideoMAE encoder, is fine-tuned using a small amount of soft-labeled CCTV video data. This model employs a peak segmentation strategy to ultimately output suspected defect time segments from long CCTV videos with a large amount of blank background. S3: PP-Model (Prior-Guided Prompt-Based Model), based on a pre-trained D²-VideoMAE encoder, fine-tuned using a small amount of soft-labeled CCTV video data and a prompt construction. It employs a defect location prior to enhance model performance and interpretability, ultimately outputting the type and time range of defects in CCTV videos. Specifically: The proposed urban drainage pipeline defect detection model framework (PipeMAE-TD) based on CCTV video incorporates both unsupervised learning and soft-label supervised learning during training. The invention process can be found in [link to invention process]. Figure 1 Therefore, in the implementation case, a dataset was constructed containing 960 unlabeled CCTV videos and 128 soft-labeled videos, as shown in Table 1. In actual engineering inspections, each CCTV video typically contains a major and most serious defect, which is then recorded and categorized into the inspection database. To ensure data balance across different defect categories, the same number of videos were selected for each defect type.

[0034] Table 1: Defect Definitions and Video Classifications (Statistics based on the most prevalent defects in each video segment)

[0035] The aim of this study is to identify defect types and their corresponding occurrence intervals from CCTV videos. Therefore, in the soft-label supervised learning process, defect categories and their time ranges are labeled in JSON format. Instead of frame-level labeling, time boundaries are labeled at the second level, thus forming soft time labels. This design significantly reduces labeling costs while still being sufficient to meet the needs of engineering-oriented defect localization tasks.

[0036] See the PipeMAE-TD framework. Figure 2 , Figure 3 , Figure 4 As shown. The training processes for D²-VideoMAE, Scan-TM, and PP-Model are respectively referred to [reference needed]. Figures 5-7 All computing programs run on a Windows system with an AMD EPYC 9534 processor and an NVIDIA Hopper H100 (80GB of memory).

[0037] The training and validation results of the pre-trained model can be found in [link / reference]. Figure 8As shown, the results indicate that during pre-training on 960 unlabeled samples, each training round of D²-VideoMAE takes approximately one hour. A total of 1000 training rounds were conducted, with a total training time of approximately 45 days. All loss terms show a significant decrease in the first 100 rounds, indicating rapid convergence in the early training phase. Between 100 and 350 rounds, the loss changes relatively gradually. After approximately 350 rounds, both the MAE loss and the total loss exhibit a significant accelerated decrease, subsequently stabilizing. The loss finally stabilizes after approximately 800 rounds. This loss trend further reflects the gradual improvement of the model. L_con The continued decline indicates that its robustness to smog and noise has improved. Meanwhile, L_mae The reduction indicates an improved ability to capture and understand the structural features of the pipeline. Furthermore, to better understand the performance of the pre-trained D²-VideoMAE encoder, this implementation example outputs the reconstruction process of the invisible mask during the verification process; see [link to documentation]. Figure 9 As shown. The first row displays the original video frames, the second row displays the visible mask input provided to the model, and the third row displays the reconstructed output. To further visualize the reconstruction error, Figure 9 The third row is presented as a heatmap, with brighter colors indicating greater error. The results show that the model effectively captures the basic structural patterns of the drainage pipes. However, its understanding of sudden or complex pipe defects remains limited. This limitation stems from the unsupervised pre-training model and will be mitigated in subsequent supervised fine-tuning stages.

[0038] Figure 10 The evolution of the training loss of Scan-TM is illustrated. The results show that the training process is generally stable. Within the first 100 epochs, the loss decreases rapidly, indicating effective convergence in the early stages. Between approximately 100 and 600 epochs, the loss decreases more slowly, then reaches a stable value. Thereafter, the loss remains relatively stable. These results demonstrate that the Scan-TM model possesses the ability to perform initial time-segmentation.

[0039] During the training of PP-Model, we employed Sharpness-Aware Minimization (SAM) in the training phase of the temporal decoder to smooth out the minimum and enhance generalization ability. SAM optimizes the worst-case loss within the bounded neighborhood of the parameters, effectively stabilizing temporal predictions, especially under severe background imbalance. Training results show that the model's loss decreases sharply in the first 300 rounds, then exhibits periodic performance fluctuations between 300 and 800 rounds. After 800 rounds, the loss tends to stabilize. Among the various loss components, CLS loss has the largest impact on the total loss, while other losses contribute relatively little. This is because the input of PP-Model consists of pre-cropped segments, making the recognition task easier. To further validate the performance of each module of the model, this case study designed a series of systematic ablation experiments and conducted extensive testing. We removed various components of the model, tested each parameter, and used recall, precision, and F1-score to evaluate performance and computational efficiency, calculated as follows: (16) (17) (18) 1. Two-stage identification strategy To verify the effectiveness of the two-stage identification strategy based on Scan-TM and PP-Model, we directly used the output of PP-Model as the model without the two-stage identification strategy to evaluate basic defect detection performance. The results are shown in Table 2. The results indicate that without this strategy, the model performance becomes extremely unstable, frequently getting trapped in local optima during training. This instability is caused by using a deeper model to identify CCTV videos of drainage pipes, which contain a large amount of irrelevant background scenery, leading to gradient problems. Given limited data and computational resources, these gradient problems cannot be solved at the model level to improve accuracy. Therefore, this strategy is well-suited for drainage pipe defect detection tasks.

[0040] Table 2 Ablation results of the two-stage identification strategy

[0041] 2. Prior Information Strategy for Defect Location: The formation locations of defects in drainage pipes exhibit certain regularities. Therefore, strategies based on prior information can help improve the model's ability to distinguish defect types. We designed a model without a prior guidance strategy as a control. The results are shown in Table 3. The results indicate that this strategy can improve the model's accuracy, although its impact on recall is limited. For defect detection tasks involving fixed occurrence locations, such as drainage pipe defects, adding this strategy can improve model performance and enhance its interpretability.

[0042] Table 3 Ablation results of the location prior strategy

[0043] 3. Different contrastive learning strategies The use of contrastive learning aims to enhance the model's understanding of pipeline structures and scenes. By simulating fog, the model develops a stronger ability to perceive foggy scenes. Therefore, we compared the performance of models using conventional contrastive learning techniques (geometric transformations, such as random cropping and rotation) with those not using contrastive learning. The results are shown in Table 4. The results indicate that contrastive learning itself effectively enhances self-supervised learning. Furthermore, using contrastive samples with fog enhancement can improve the model's performance in specific scenes, resulting in slightly better results.

[0044] Table 4 Ablation results of different contrastive learning strategies

[0045] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. The above preferred embodiments of the present invention are not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the technical solution of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention shall still fall within the scope of the technical solution of the present invention.

Claims

1. A method for detecting defects in urban drainage pipelines based on CCTV video, characterized in that: include S1: A D²-VideoMAE encoder, based on a dual-mask strategy and a fog contrast learning strategy, uses a large amount of CCTV video footage of drainage pipes from multiple cities and countries for unsupervised learning to train a pre-trained encoder capable of perceiving pipe structures and defects in CCTV video footage; S1 includes: S11: Collects a large amount of CCTV video data of drainage pipes from different regions without any processing; S12: Simulate fog and noise to establish original video comparison samples for self-supervised comparison training of the model; S13: Establish encoder and decoder structures, and adopt a dual-mask learning strategy. Set up a large-area mask on the D²-VideoMAE encoder to block video content, while covering a small area on the decoder to restore the original video and the video data of the comparison sample. S14: Compare the real and predicted samples in the masked regions of the original video and the contrast sample. Adjust the model results using MAE loss, contrast loss and an auxiliary class loss to construct a closed-loop self-supervised learning strategy. The obtained pre-trained D²-VideoMAE encoder has the basic function of perceiving the structure and defect morphology of drainage pipes. S2: Scan-TM, based on a pre-trained D²-VideoMAE encoder, is fine-tuned using a small amount of CCTV video data with soft labels. This model employs a peak segmentation strategy and can ultimately output suspected defect time segments from long CCTV videos with a large amount of blank background. S3: PP-Model, based on a pre-trained D²-VideoMAE encoder, uses a small amount of soft-labeled CCTV video data and prompt construction for fine-tuning training. It adopts a defect location prior to enhance the model's performance and interpretability, and can finally output the type of defect in CCTV video and the time range of its occurrence.

2. The method for detecting defects in urban drainage pipelines based on CCTV video according to claim 1, characterized in that: S2 includes: S21: Collect a small amount of CCTV video for soft labeling, only labeling the defect type and the approximate time period of occurrence, for fine-tuning training of the model; S22: The Scan-TM model is built based on the pre-trained D²-VideoMAE encoder. The Scan-TM decoder adopts an attention mechanism structure similar to that of the encoder. S23: During training, the model freezes all parameters of the pre-trained D²-VideoMAE encoder, and only applies Scan-TM based on a small amount of soft-label data. decoder The parameters of the structure are trained to output the probability of the existence and category of defects, and the corresponding loss is calculated. S24: During the inference process, the CCTV video to be predicted will be analyzed by the Scan-TM model to predict the defects in each frame. Based on peak detection and segmentation algorithms, segments in the long CCTV video that are suspected of having a certain defect will be obtained for subsequent inference and identification.

3. The method for detecting defects in urban drainage pipelines based on CCTV video according to claim 1, characterized in that: S3 includes: S31: PP-Model is built based on the pre-trained D²-VideoMAE encoder. A Prompt parameter layer is overlaid on the encoder for fine-tuning. The detection head adopts a position prior strategy to enhance the performance and interpretability of the model. S32: During the training process, the model uses the PEFT strategy to freeze the first few layers of the D²-VideoMAE encoder structure. Based on a small amount of soft-label data, only the parameters of the last few layers of the encoder structure, the Prompt parameter, and the detection head are trained. Finally, the type of defect and the time range of occurrence are output, and the corresponding loss and location prior loss are calculated. S33: During the inference process, the suspected segments output by Scan-TM are input into PP-Model to ultimately determine the defect status of the segment. Based on the Laplacian variance sharpness score, a defect report is output, including: defect type, time range of occurrence in the video, and key representative frames.

4. The method for detecting defects in urban drainage pipelines based on CCTV video according to claim 1, characterized in that: In step S12, fog and noise are artificially introduced into the original video to generate contrast samples. These samples are then used in the contrastive learning process during D²-VideoMAE pre-training. The fog and noise can be described as follows: (1) (2) in, I orig ( x , y () represents the original image. F ( x , y () indicates a fog background. α Controlling the intensity of the fog, N (0, σ 2 () represents Gaussian noise; during data augmentation, the fog intensity and noise variance are randomly selected within a predefined range to enhance robustness and generalization ability; given an input video Then the comparison learning samples can be represented as .

5. The method for detecting defects in urban drainage pipelines based on CCTV video according to claim 1, characterized in that: In step S13, inside the encoder, the encoder mask is confined to a fixed spatial position to capture local appearance features and spatial continuity; the number of visible masks is used... It means that among them The invisible mask ratio is 90%, which forces the model to learn a strong global representation from limited visible information; After 3D block input, input x i =[ x 1, x [2] The vectors are converted into token embeddings, category embeddings, and location embeddings. The invisible mask is processed by a decoder composed of ViViT-based multi-head attention (MHA) and multi-layer perceptron (MLP) to generate intermediate vectors. and categories ; In the decoder, a mask of the invisible region is extracted from a random spatial location through a conventional sampling and rolling offset strategy. The decoder consists of a multi-head self-attention mechanism and a multilayer perceptron module, which are specifically designed to reconstruct the invisible region and thus learn the temporal continuity features. The number of masks in the decoder is The combined TOKEN sequence is Only for the invisible parts Z m Reconstruction will be carried out, and the reconstruction will be carried out as follows: This allows computation to be focused on a small subset of targets, thereby improving training efficiency.

6. The method for detecting defects in urban drainage pipelines based on CCTV video according to claim 1, characterized in that: In S14, the overall training loss of D²-VideoMAE consists of three parts: MAE reconstruction loss, contrastive learning loss, and auxiliary classification loss. The definition of MAE reconstruction loss is: (3) in, Y extra Represents the true value, while P = τ · patch · patch Contrastive learning loss is used to measure the similarity between contrastive pairs. Its purpose is to minimize the loss between positive pairs and maximize the loss between negative pairs. (4) in, S It is by z 1 and z 2. The constructed similarity matrix, , diag_labels Then corresponding to S Labels of positive samples on the matrix diagonal; an auxiliary classification loss is introduced by using the encoder's CLS token, thus providing weak geometric supervision to mitigate errors caused by CCTV viewpoint drift: (5) (6) Using weighting coefficients λ = 0.2 and μ = 0.1; After extensive pre-training on unlabeled CCTV videos, the encoder acquired a powerful defect feature extraction capability, which was subsequently used to train Scan-TM and PP-model.

7. The method for detecting defects in urban drainage pipelines based on CCTV video according to claim 2, characterized in that: In step S22, the Scan-TM is trained on the basis of a pre-trained D²-VideoMAE encoder, during which the encoder parameters are frozen; using limited soft-labeled data, time-segment data of each defect occurrence is generated; Scan-TM uses a dedicated decoder to simulate spatial relationships based on multiple region labeling features of each frame and to estimate the probability of defect existence; the dedicated decoder consists of a multi-head self-attention mechanism (MHA) and a multilayer perceptron (MLP) module with the same encoder but fewer layers.

8. The method for detecting defects in urban drainage pipelines based on CCTV video according to claim 2, characterized in that: In S23, an input is given. The encoder will generate multiple category token embeddings. These embeddings are stacked into a sequence of categorical features. The sequence is projected and sinusoidal positional encoding is added, then embedded as a query input into the decoder to obtain... A sampling rate of 1 frame per second is used to reduce redundancy and computational overhead; the decoder output can be represented as: (7) The two prediction heads output the probability of the presence of the defect, respectively. and defect category probability ; The loss function for defect presence and coarse classification is defined as follows: (8) (9) In addition, a background (BG) class is added during training to ensure clear background classification and suppress prediction fluctuations. (10) This design can alleviate category imbalance and reduce background interference in long CCTV videos.

9. The method for detecting defects in urban drainage pipelines based on CCTV video according to claim 2, characterized in that: In S24, for a given original surveillance video sample, a low-rate temporal scan is first performed by uniformly sampling frames; each sampled frame is encoded by the trained Scan-TM model to generate the probability of each frame for 16 defect categories. To transform noisy probability trajectories into discrete defect events, a category-level peak segmentation strategy is applied to the smooth probability curves of each category. Specifically, values ​​below the removal threshold are considered background, peaks above the peak threshold are considered event centers, and adjacent peaks are merged or split based on whether the adjacent peak valleys are below the removal threshold. For each coarse interval, an interval-level prior model is constructed by averaging the time posterior values ​​within that interval.

Citation Information

Patent Citations

  • Drainage pipeline detection and assessment method using handheld video and closed circuit television

    CN103423596A

  • Pipeline defect detecting and tracking method and device

    CN116703826A

  • Drainage pipeline video defect time point positioning method based on Transform architecture

    CN118447330A