A multi-feature fusion driven high-quality sewer defect sample automatic labeling method, system, device and medium
By employing a multi-feature fusion-driven approach, combined with forward diffusion noise injection and back diffusion iterative denoising techniques, an improved YOLO11-EfficientViT model was constructed. This model addresses the issues of low efficiency and inconsistent quality in manual annotation of drainage pipeline defect sample datasets, achieving efficient and accurate automatic annotation and model optimization.
Patent Information
- Application Number
- CN202510436545.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In existing technologies, manual annotation of drainage pipe defect sample datasets suffers from subjective bias and environmental complexity, resulting in low annotation efficiency and inconsistent quality, which affects model training performance.
A multi-feature fusion-driven approach is adopted. By performing frame extraction and defect visibility screening on the collected drainage pipeline image data, combined with forward diffusion noise injection and back diffusion iterative denoising technology, defect images that meet the annotation requirements are generated. An improved YOLO11-EfficientViT fusion model is constructed for pipe diameter perception training, realizing multi-level quality verification and model iterative optimization of the end-to-end annotation pipeline.
It significantly improves the efficiency and accuracy of defect labeling in drainage pipelines, reduces subjective bias caused by differences in human experience, forms a virtuous cycle of data quality and model performance, and continuously optimizes the defect detection model.
Smart Images

Figure CN120526252B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, system, device, and medium for automatic labeling of high-quality drainage pipeline defect samples driven by multi-feature fusion. Background Technology
[0002] Currently, YOLO-based deep learning models for computer vision are widely used for the rapid localization and identification of pipeline defects in underground drainage pipeline inspection videos. The construction of such models typically relies on two core steps: the creation of a pipeline defect sample dataset and model training. In existing technologies, the creation of the sample dataset requires manual annotation. The specific process involves collecting internal images of drainage pipelines, followed by manual annotation of the defect images by technicians according to industry standards (such as CJJ 181-2012 "Technical Specification for Inspection and Evaluation of Urban Drainage Pipelines"), including the definition of defect types and extents. After verification, the image files and corresponding annotation files are exported.
[0003] However, the creation of sample datasets in existing technologies has significant limitations. First, since the labeling of defect types and extents relies entirely on manual judgment, the differences in experience among different technicians can easily lead to subjective biases in the labeling results, making it difficult to guarantee data consistency and accuracy. Second, the shooting environment inside pipes is complex (e.g., insufficient light, obstruction by dirt, etc.), and the boundaries of defect areas are often blurred, further increasing the difficulty of manual labeling. Therefore, existing labeling methods suffer from low efficiency, long processing time, and high cost, and the labeling quality is greatly affected by human factors, making it difficult to meet the demand for high-quality labeled data for large-scale model training.
[0004] Furthermore, during the model training phase, existing technologies heavily rely on the aforementioned manually labeled sample data. Since subjective errors introduced during labeling directly affect the model's learning performance, the trained model may exhibit potential deficiencies in defect identification accuracy and generalization ability. Therefore, efficiently generating highly consistent labeled data has become a key technical bottleneck in improving the performance of drainage pipeline defect detection models. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a method, system, device and medium for automatic annotation of high-quality drainage pipeline defect samples driven by multi-feature fusion. It solves the technical problems in the prior art where the pipeline defect sample dataset based on manual annotation has low annotation efficiency and inconsistent quality due to subjective judgment bias and environmental complexity, which affects the model training effect.
[0007] (II) Technical Solution
[0008] To achieve the above objectives, the main technical solutions adopted by the present invention include:
[0009] In a first aspect, embodiments of the present invention provide a method for automatically labeling high-quality drainage pipeline defect samples driven by multi-feature fusion, comprising:
[0010] The collected drainage pipe image data is processed by frame extraction.
[0011] Defect visibility screening is performed on the extracted images. For images that do not meet the visibility threshold, forward diffusion noise injection, noise distribution prediction, and back diffusion iterative denoising are performed to obtain defect images that meet the annotation requirements.
[0012] The defect images are sent to the annotation terminal, and the feedback annotation results are used to generate a dataset with contour annotations through multi-level quality verification.
[0013] An improved YOLO11-EfficientViT fusion model was constructed, pre-trained weights were loaded, and the model was trained on the dataset using a pipe diameter-aware training mode to generate a defect detection model with spatial awareness.
[0014] The trained model is loaded into the annotation module, and the model path, defect type and annotation parameters are configured. The end-to-end annotation pipeline is controlled to perform batch defect detection and contour annotation on the new image data. After multi-level quality verification, the results are fed back to the dataset to drive model iterative optimization.
[0015] Optionally, frame extraction processing of the acquired drainage pipe image data includes:
[0016] Pipeline inspection robots, crawling robots, or drones are used to enter the drainage pipeline from the manhole inlet or open port and collect internal images along the longitudinal axis of the pipeline to obtain pipeline inspection video data with continuous spatial coverage.
[0017] Import the pipeline inspection video data into the computer storage medium, and establish a hierarchical storage directory according to the spatiotemporal information of the acquisition to form the original video database;
[0018] Configure the video parsing framework through the command-line interface to determine the mapping relationship between the source video path and the frame-extracted image output path;
[0019] Configure image output format parameters and frame rate parameters; among them, the frame rate parameters are generated by real-time analysis of video data resolution, image quality assessment value and pipe cross-section coverage, and after multi-dimensional weight fusion, they are mapped to a preset frame rate range to generate frame rate parameters that are adaptive to different pipe working conditions.
[0020] The system recursively traverses the storage directory of the original video database, identifies the target video file based on the mapping relationship between the source video path and the output path of the extracted frame image, creates a subdirectory tree with the same name, calls the video decoding program to perform time-series sampling according to the determined frame extraction rate parameter, and then uses serialization naming rules to store the extracted frame images to the corresponding subdirectories.
[0021] Optionally, defect visibility screening is performed on the extracted images. Images that do not meet the visibility threshold are subjected to forward diffusion noise injection, noise distribution prediction, and back diffusion iterative denoising to obtain defect images that meet the annotation requirements, including:
[0022] Based on a preset defect visibility threshold, the extracted images are automatically filtered.
[0023] If the defect visibility threshold is met, the result is directly output to the subsequent annotation process.
[0024] If the defect visibility threshold is not met, progressive noise injection is performed on the corresponding image. The noise intensity is gradually superimposed according to a Gaussian distribution within a preset time step sequence using a preset diffusion model. The noise injection intensity is controlled by the diffusion scheduling coefficient, and the image is gradually degraded into a multi-time step noise feature sequence.
[0025] The noise feature map of each time step in the multi-time step noise feature sequence is input into the pre-trained U-Net noise prediction model to predict the noise component injected at each time step during the forward diffusion process. By comparing the mean square error of the predicted noise component injected at each time step with the actual injected noise, the noise intensity confidence map of each time step is generated and then fused by time-series weighting to obtain the global noise distribution probability map.
[0026] Based on the global noise distribution probability map, the back diffusion process is executed in an iterative manner: in the back time step sequence, the predicted noise component is gradually stripped from the noise feature sequence of multiple time steps, and the noise removal intensity is controlled by the back diffusion scheduling coefficient. The image details are gradually reconstructed through multiple rounds of iteration. When the preset maximum number of iterations is reached or the defect visibility threshold is met, a defect image that meets the annotation requirements is generated.
[0027] Among them, the diffusion scheduling coefficient is a time-dependent parameter used to control the noise injection during the forward process of the diffusion model, and the reverse diffusion scheduling coefficient is a time-dependent parameter used to control the denoising intensity at each step during the reverse diffusion process, and satisfies the reverse correlation with the diffusion scheduling coefficient.
[0028] Optionally, defect images are sent to the annotation terminal, and the feedback annotation results are processed through multi-level quality checks to generate a dataset with contour annotations, including:
[0029] Based on the historical annotation quality and efficiency data of the annotation personnel corresponding to each annotation terminal, the ability levels are divided and associated with differentiated baseline workloads. Among them, the baseline workloads of annotation personnel at the primary, intermediate and advanced ability levels increase progressively according to a preset ratio.
[0030] The system monitors the backlog of tasks to be labeled in real time. When the backlog of tasks drops to the preset threshold ratio of the baseline task volume, the system automatically triggers the task replenishment mechanism and re-allocates the baseline task volume that matches the capability level to the terminal.
[0031] Based on the task backlog, processing rate, and terminal resource utilization of the labeled terminals, a dynamic allocation priority score is generated.
[0032] When a new batch of defective images to be labeled is detected, the new batch is divided into task segments that match the baseline workload of each level according to the baseline workload of the terminal capability level. The available labeling terminals are sorted from high to low according to the priority score, and the task segments are pushed to the corresponding labeling terminals according to the sorting priority.
[0033] Simultaneously push standard guidelines documents containing defect type option sets, contour annotation specifications, and example templates to the corresponding annotation terminals;
[0034] Receive the annotation result file returned by the annotation terminal, parse the defect type label and contour coordinate data in the annotation result file, and perform format checks including file integrity and coordinate range validity;
[0035] Based on the defect type labels in the annotation results file, match them with the category range of the defect type option set to filter out undefined or out-of-bounds types;
[0036] If the labeled result file passes the verification and matching, it is given a qualified label, and a standardized dataset containing image files, label files, and annotation files is generated;
[0037] If the annotation result file fails the inspection or matching, the defect information is returned to the annotation terminal, the corresponding defect image is locked, and the re-annotation process of the same annotation terminal or the reassignment process of the annotation terminal is triggered.
[0038] Optionally, an improved YOLO11-EfficientViT fusion model is constructed, pre-trained weights are loaded, and the model is trained on the dataset using a pipe diameter-aware training mode to generate a defect detection model with spatial awareness capabilities, including:
[0039] Read images and annotation result files from the dataset with contour annotations, automatically divide them into training set, validation set and test set according to preset ratio, and verify the correspondence between images and annotation result files;
[0040] The configuration includes multiple pieces of information such as dataset path, defect type information, input image size, and defect category parameters, and simultaneously sets pipe diameter sensing training parameters including receptive field scaling range and pipe diameter perturbation intensity.
[0041] An improved YOLO11-EfficientViT network structure is constructed, which includes a backbone network with configured EfficientViT modules and a detection head with embedded adaptive receptive field modules. The pre-trained weight file of YOLO11 is loaded from the pre-stored path, the weights of the backbone network are retained, and the detection head modules are randomly initialized.
[0042] Model training is performed based on the training set, and the pipe diameter sensing training mode is enabled during the training process. The size of the detection head receptive field is adjusted according to the scaling range based on the pipe diameter parsed from the annotation result file to match the defect scale of different pipe diameters. At the same time, pipe diameter perturbation is injected to generate multi-scale defect samples. The defect morphology is adjusted through affine transformation to simulate the imaging deformation under different pipe diameters.
[0043] After every N training rounds, the model is evaluated on the validation set to obtain evaluation metrics including segmentation accuracy, precision, and recall, and the learning rate is dynamically adjusted or early stopping is triggered.
[0044] After all training is completed, the performance is verified on the test set, and the qualified models are converted into the specified format.
[0045] Optionally, the trained model is loaded into the annotation module, and the model path, defect type, and annotation parameters are configured. The end-to-end annotation pipeline is controlled to perform batch defect detection and contour annotation on newly added image data. After multi-level quality verification, the results are fed back to the dataset to drive iterative model optimization, including:
[0046] Create a configuration file for the trained defect detection model, and write the model type, model name, model path and the set of pipeline defect types to be detected in the configuration file, and configure the automatic annotation parameters simultaneously.
[0047] Load the trained defect detection model and configuration file into the model management system, register the model name, version number and configuration file path for the annotation module to call;
[0048] Scan the newly added drainage pipeline image data catalog, schedule it to the annotation module in batches, start the end-to-end annotation pipeline, call the model to perform batch defect detection and contour annotation, and perform list matching verification and logic verification;
[0049] If the defect detection and contour annotation results pass the list matching and logical verification, the defect detection and contour annotation results are fed back to the corresponding annotation terminal. Based on the quality results fed back by the standard terminal, the defect detection and contour annotation results are corrected and written into the dataset, the dataset version is updated, and model iteration training is triggered.
[0050] If the defect detection and contour annotation results fail the verification, record the error type and re-inject the data into the annotation queue to await correction.
[0051] Optionally, if the defect detection and contour annotation results pass list matching and logical verification, the defect detection and contour annotation results are fed back to the corresponding annotation terminal. Based on the quality results fed back by the standard terminal, the defect detection and contour annotation results are corrected and written into the dataset. Updating the dataset version and triggering model iterative training includes:
[0052] Based on the defect detection and contour annotation results, the detected defect images are classified according to boundary salience and semantic complexity, into object-type defects with clear boundaries and definition-type defects with vague boundaries, and annotation processing priority labels are generated according to the highest complexity of the defects in the image.
[0053] Establish dynamic matching rules between the ability level of the labeling personnel and the defect type. Among them, the personnel with the basic ability level are assigned to handle definition-type defects and are limited to the baseline task volume; the personnel with the intermediate ability level are assigned to handle definition-type defects and are limited to a preset multiple of the baseline task volume; and the personnel with the advanced ability level are assigned to handle definition-type defects and are limited to the baseline task volume.
[0054] The system monitors the backlog of tasks to be labeled in real time. When the backlog of tasks drops to the preset threshold ratio of the baseline task volume for the corresponding level, an automatic supplementation mechanism is triggered to supplement new tasks that match the corresponding defect type to terminals with each capability level.
[0055] Collect data on the distribution of labeling error types of personnel at each capability level in the defect result quality inspection process. For personnel whose labeling pass rate is greater than the preset pass standard for multiple consecutive batches, upgrade their capability level by at least one level. For personnel whose error rate is greater than the preset error standard for a single batch, reduce their capability level by at least one level and reduce their workload allocation ratio.
[0056] A multidimensional correlation analysis is performed on quality assessment data, dynamic matching rules, and defect type distribution characteristics. Based on the correlation analysis results, the dynamic matching rules and dataset versions are updated to form a co-evolutionary mechanism for improving annotation quality and iterating model parameters.
[0057] Secondly, embodiments of the present invention provide a high-quality automatic annotation system for drainage pipeline image data based on a deep learning model of multi-feature fusion, comprising:
[0058] The frame extraction processing module is used to perform frame extraction processing on the collected drainage pipeline image data.
[0059] The filtering and enhancement module is used to filter the defect visibility of the frame-stripped image. For images that do not meet the visibility threshold, forward diffusion noise injection, noise distribution prediction and back diffusion iterative denoising are performed to obtain defect images that meet the annotation requirements.
[0060] The annotation module is used to send defect images to the annotation terminal and generate a dataset with contour annotations from the feedback annotation results through multi-level quality verification.
[0061] The model building module is used to build an improved YOLO11-EfficientViT fusion model, load pre-trained weights and train it on the dataset using the pipe diameter-aware training mode to generate a defect detection model with spatial awareness capabilities.
[0062] The model deployment module is used to load the trained model into the annotation module, configure the model path, defect type and annotation parameters, control the end-to-end annotation pipeline to perform batch defect detection and contour annotation on new image data, and feed back to the dataset after multi-level quality verification to drive model iterative optimization.
[0063] Thirdly, embodiments of the present invention provide an automatic annotation device for high-quality drainage pipeline image data based on a deep learning model of multi-feature fusion, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the automatic annotation method for high-quality drainage pipeline defect samples driven by multi-feature fusion as described above.
[0064] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the multi-feature fusion-driven automatic annotation method for high-quality drainage pipe defect samples as described above.
[0065] (III) Beneficial Effects
[0066] The beneficial effects of this invention are as follows: This method significantly improves the efficiency, accuracy, and model generalization ability of drainage pipeline defect labeling through systematic technical improvements. Specifically, the beneficial effects are reflected in the following aspects:
[0067] First, to address the issues of image blurring and noise interference caused by complex shooting environments, a combination of forward diffusion noise injection and back diffusion iterative denoising techniques is introduced. After modeling the noise distribution of low-quality images during forward diffusion, back diffusion iterative denoising is used to reconstruct image details, effectively improving the visibility and edge clarity of defective areas. This provides highly reliable input data for subsequent annotation, reducing annotation errors caused by image quality defects from the source.
[0068] Secondly, an end-to-end annotation pipeline is constructed based on the improved YOLO11-EfficientViT fusion model. By combining pre-trained weight inheritance with pipe diameter-aware training mode, the model has a deep understanding of the spatial structural features of the pipeline. In the automatic annotation stage, it can accurately identify the defect morphology in different pipe diameter scenarios, significantly reducing the subjective annotation bias caused by differences in human experience.
[0069] Furthermore, by designing a multi-level quality verification mechanism and a closed loop of manual correction feedback, we can achieve cross-validation and error correction of the annotation results, and dynamically feed the corrected high-quality data back to the training set, forming a virtuous cycle of "model optimization - automatic annotation - manual calibration - data iteration". This breaks through the efficiency bottleneck of traditional manual annotation while continuously improving the consistency of the dataset.
[0070] More importantly, this invention innovatively couples generative and detection models based on a closed-loop iterative optimization architecture: it introduces a combination of forward diffusion noise injection and back diffusion iterative denoising techniques to improve the quality of raw data, thereby reducing annotation difficulty, and improves the detection model to provide high-precision automatic annotation results. Meanwhile, the incremental data generated by the automatic annotation module can feed back into model iteration, forming a positive cycle where data quality and model performance mutually promote each other. This closed-loop technology not only effectively overcomes the inherent drawbacks of manual annotation, such as high subjectivity and cost, but also provides a scalable technical framework for the continuous optimization of drainage pipeline defect detection models. Attached Figure Description
[0071] Figure 1 A flowchart illustrating the method provided in an embodiment of the present invention;
[0072] Figure 2 This is a schematic diagram illustrating the specific process of step S1 of the method provided in this embodiment of the invention;
[0073] Figure 3 This is a detailed flowchart illustrating step S2 of the method provided in this embodiment of the invention;
[0074] Figure 4 Images that do not meet the defect visibility threshold provided in the embodiments of the present invention;
[0075] Figure 5An image that satisfies the defect visibility threshold for the method provided in the embodiments of the present invention;
[0076] Figure 6 A schematic diagram illustrating the forward and reverse diffusion processes of the method provided in this embodiment of the invention;
[0077] Figure 7 This is a detailed flowchart illustrating step S3 of the method provided in this embodiment of the invention;
[0078] Figure 8 A schematic diagram illustrating the defect contour annotation of the method provided in the embodiments of the present invention;
[0079] Figure 9 A schematic diagram illustrating the defect area annotation of the method provided in this embodiment of the invention;
[0080] Figure 10 This is a detailed flowchart illustrating step S4 of the method provided in this embodiment of the invention;
[0081] Figure 11 The model training evaluation metrics provided in the embodiments of the present invention;
[0082] Figure 12 This is a detailed flowchart illustrating step S5 of the method provided in this embodiment of the invention;
[0083] Figure 13 This is a detailed flowchart illustrating step S54 of the method provided in this embodiment of the invention.
[0084] Figure 14 This is a schematic diagram of the overall process of the method provided in the embodiments of the present invention. Detailed Implementation
[0085] To better explain and facilitate understanding of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0086] like Figure 1As shown in the embodiment of the present invention, a high-quality automatic annotation method for drainage pipeline defect samples driven by multi-feature fusion is proposed, comprising: performing frame extraction processing on the collected drainage pipeline image data; performing defect visibility screening on the extracted images, and performing forward diffusion noise injection, noise distribution prediction, and back diffusion iterative denoising on images that do not meet the visibility threshold to obtain defect images that meet the annotation requirements; sending the defect images to the annotation terminal, and generating a dataset with contour annotations by passing the feedback annotation results through multi-level quality verification; constructing an improved YOLO11-EfficientViT fusion model, loading pre-trained weights and training the dataset using a pipe diameter-aware training mode to generate a defect detection model with spatial awareness capabilities; loading the trained model into the annotation module, configuring the model path, defect type, and annotation parameters, controlling the end-to-end annotation pipeline to perform batch defect detection and contour annotation on newly added image data, and feeding back the results to the dataset after multi-level quality verification to drive model iterative optimization.
[0087] This method significantly improves the efficiency, accuracy, and model generalization ability of drainage pipeline defect labeling through systematic technical improvements. The specific beneficial effects are reflected in the following aspects:
[0088] First, to address the issues of image blurring and noise interference caused by complex shooting environments, a combination of forward diffusion noise injection and back diffusion iterative denoising techniques is introduced. After modeling the noise distribution of low-quality images during forward diffusion, back diffusion iterative denoising is used to reconstruct image details, effectively improving the visibility and edge clarity of defective areas. This provides highly reliable input data for subsequent annotation, reducing annotation errors caused by image quality defects from the source.
[0089] Secondly, an end-to-end annotation pipeline is constructed based on the improved YOLO11-EfficientViT fusion model. By combining pre-trained weight inheritance with pipe diameter-aware training mode, the model has a deep understanding of the spatial structural features of the pipeline. In the automatic annotation stage, it can accurately identify the defect morphology in different pipe diameter scenarios, significantly reducing the subjective annotation bias caused by differences in human experience.
[0090] Furthermore, by designing a multi-level quality verification mechanism and a closed loop of manual correction feedback, we can achieve cross-validation and error correction of the annotation results, and dynamically feed the corrected high-quality data back to the training set, forming a virtuous cycle of "model optimization - automatic annotation - manual calibration - data iteration". This breaks through the efficiency bottleneck of traditional manual annotation while continuously improving the consistency of the dataset.
[0091] More importantly, this invention innovatively couples generative and detection models based on a closed-loop iterative optimization architecture: it introduces a combination of forward diffusion noise injection and back diffusion iterative denoising techniques to improve the quality of raw data, thereby reducing annotation difficulty, and improves the detection model to provide high-precision automatic annotation results. Meanwhile, the incremental data generated by the automatic annotation module can feed back into model iteration, forming a positive cycle where data quality and model performance mutually promote each other. This closed-loop technology not only effectively overcomes the inherent drawbacks of manual annotation, such as high subjectivity and cost, but also provides a scalable technical framework for the continuous optimization of drainage pipeline defect detection models.
[0092] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings.
[0093] Specifically, embodiments of the present invention provide a method for automatic annotation of high-quality drainage pipeline defect samples driven by multi-feature fusion, comprising:
[0094] S1. Perform frame extraction processing on the collected drainage pipe image data.
[0095] Furthermore, such as Figure 2 As shown, step S1 includes:
[0096] S11. Using pipeline inspection robots, crawling robots, or drones, enter the drainage pipeline from the manhole inlet or open port, and collect internal images along the pipeline's longitudinal axis to obtain pipeline inspection video data with continuous spatial coverage. Control the pipeline inspection robot, crawling robot, or drone to enter the drainage pipeline from below the rainwater well, sewage well, or the open side of the pipeline, and move at a constant speed along the pipeline's longitudinal axis, continuously capturing a 360° surround video stream inside the pipeline, covering the entire circumference and continuous longitudinal space of the pipeline's inner wall; after reaching the pipeline end, the equipment returns along the same path, outputting the complete pipeline inspection video.
[0097] S12. Import the pipeline inspection video data into the computer storage medium, and establish a hierarchical storage directory according to the spatiotemporal information of the acquisition to form the original video database. Connect the shooting device to the computer port via a data cable, import the drainage pipeline image data (usually in MP4, AVI, or other video data formats), and store it in the computer.
[0098] S13. Configure the video parsing framework via the command-line interface to determine the mapping relationship between the source video path and the extracted frame image output path. The source video path is the root directory of the original video database, and the extracted frame image output path is the root directory of the extracted frame results.
[0099] S14. Configure image output format parameters and frame rate parameters. The frame rate parameter is generated by real-time analysis of video data resolution, image quality assessment value, and pipe cross-section coverage, and then mapping it to a preset frame rate range after multi-dimensional weight fusion, thus generating a frame rate parameter that adapts to different pipe working conditions.
[0100] In a specific embodiment, the adaptive frame rate parameter (adaptiveFPS) for different pipeline operating conditions mainly refers to three factors: image resolution, image quality assessment, and complete pipeline cross-section. It is generated by dynamically fusing multi-dimensional features, and its calculation logic and steps are as follows:
[0101] The formula for adaptive frame rate is as follows:
[0102] adaptiveFps=min(5,max(1,round(2·wres·wquality·wpipe)))
[0103] In the formula, the image resolution weight wres = (1920×1080) / actual image resolution, the image quality assessment weight wquality = average BRISQUE / 50, and the pipe cross-section coverage weight wpipe = 1 / (pipe cross-section coverage + 0.1). The range constraint is: adaptiveFPS = [1, median, 5], which constrains the frame rate range to 1~5 frames / second (fps) to avoid data redundancy due to excessively high frequency or loss of keyframes due to excessively low frequency.
[0104] S15. Recursively traverse the storage directory of the original video database, identify the target video file based on the mapping relationship between the source video path and the output path of the extracted frame image, and create a subdirectory tree with the same name. Call the video decoding program to perform time-series sampling according to the frame extraction rate parameter, and then use the serialization naming rules to store the extracted frame image to the corresponding subdirectory.
[0105] Specifically, a command prompt (PowerShell) and a multimedia processing program (FFmpeg) are prepared to extract frames from all detected videos and export images in JPG format. First, the video import location and image export location are set. Then, the image output format (JPG, PNG, etc.) and frame extraction rate (frames per second) are set. Finally, FFmpeg processes the video files in batches.
[0106] S2. Perform defect visibility screening on the extracted images. For images that do not meet the visibility threshold, perform forward diffusion noise injection, noise distribution prediction, and back diffusion iterative denoising to obtain defect images that meet the annotation requirements.
[0107] In forward diffusion, Gaussian noise is gradually introduced into the training samples, causing them to deviate from the high-dimensional, complex subspace of the original data. Essentially, this process maps the complex, unstructured distribution of the original data to a simpler distribution (such as an isotropic Gaussian distribution) with a defined form that is easy to model and sample. In this way, the model can learn during training how to inversely reconstruct the structural features of the original data distribution from this simpler distribution. At the end of the forward diffusion process, the image data completely loses its original semantic information, and its distribution is effectively transformed into a well-defined, easily processed latent space, thus providing a modeling foundation for the subsequent inverse diffusion process.
[0108] The core idea of "backdiffusion" is to gradually reverse the forward diffusion process. It attempts to recover the interference introduced into the image during the forward diffusion process in a slow and iterative manner. The starting point of backdiffusion is the state at the end of the forward diffusion. The advantage of starting from a simple distribution (such as an isotropic Gaussian distribution) is that it is known how to sample from that distribution, meaning any initial point outside the data subspace can be easily obtained. The goal of this invention is to design a path that returns from this simple space to the complex subspace where the original data resides. However, the problem is that while there are theoretically infinitely many paths starting from a point in a simple distribution, only a very few correctly lead to the data subspace. In diffusion probabilistic models, this process is modeled based on the small perturbations at each step of the forward diffusion process. During forward diffusion, the probability density function (PDF) satisfied by the image changes slightly at each time step. Therefore, in backdiffusion, a deep learning model is used to predict the PDF parameters during the forward diffusion process at each time step, thereby constructing a reasonable reverse transformation path. Once the model has been trained, it can start from any point in a simple distribution and iteratively perform "denoising" operations with the help of the model, eventually returning to the subspace where the original data distribution is located.
[0109] Furthermore, such as Figure 3 As shown, step S2 includes:
[0110] S21. Based on a preset defect visibility threshold, automatically filter the images after frame extraction.
[0111] S22. If the defect visibility threshold is met, the result will be directly output to the subsequent annotation process.
[0112] Based on preset defect visibility thresholds (such as contrast thresholds, defect area ratio thresholds, etc.), the extracted images (which can be images with a single defect or multiple defects) are automatically filtered. If all indicator values are higher than the threshold (e.g., ...), the image will be automatically filtered. Figure 5 (Example), marked as qualified images, are directly output to the annotation process without manual intervention; if any indicator is below the threshold (e.g. Figure 4 Example), marked as an unacceptable image.
[0113] S23. If the defect visibility threshold is not met, progressive noise injection is performed on the corresponding image. The noise intensity is gradually superimposed according to a Gaussian distribution within a preset time step sequence using a preset diffusion model. The noise injection intensity is controlled by the diffusion scheduling coefficient, and the image is gradually degraded into a multi-time step noise feature sequence.
[0114] S24. Input the noise feature map of each time step in the multi-time step noise feature sequence into the pre-trained U-Net noise prediction model to predict the noise component injected at each time step during the forward diffusion process. By comparing the mean square error of the predicted noise component injected at each time step with the actual injected noise, generate the noise intensity confidence map of each time step and perform time-weighted fusion to obtain the global noise distribution probability map.
[0115] S25. Based on the global noise distribution probability map, the back diffusion process is executed in an iterative manner: In the back time step sequence, the predicted noise component is gradually stripped from the multi-time step noise feature sequence, and the noise removal intensity is controlled by the back diffusion scheduling coefficient. The image details are gradually reconstructed through multiple iterations. When the preset maximum number of iterations is reached or the defect visibility threshold is met, a defect image that meets the annotation requirements is generated.
[0116] Among them, the diffusion scheduling coefficient is a time-dependent parameter used to control the noise injection during the forward process of the diffusion model, and the reverse diffusion scheduling coefficient is a time-dependent parameter used to control the denoising intensity at each step during the reverse diffusion process, and satisfies the reverse correlation with the diffusion scheduling coefficient.
[0117] In one specific embodiment, reference is made to Figure 6 For images that do not meet the defect visibility threshold, the following operations are performed:
[0118] First, a forward-diffused p-core is used. θ (X t-1 |X t ), in the preset time step sequence {t1,t2,...,t n In the process, the denoising diffusion probability model (DDPM) is used, based on the diffusion scheduling coefficient β. t Dynamically adjust the noise intensity and gradually inject noise into the image according to a Gaussian distribution to generate a multi-time step noise feature sequence {I t1 ,I t2 ,...,I tn}
[0119] Specifically, the forward diffusion process transforms the original data distribution into an isotropic Gaussian distribution that is easy to model by progressively injecting Gaussian noise into the image. This process is based on a Markov chain, and its core is to progressively perturb the image through a series of conditional probability distributions. The probability density function (PDF) of the forward diffusion process is composed of the product of a series of conditional distributions from time step t = 1 to T, denoted as: In the formula, this forward process is fixed and known, and is usually considered as the unlearnable part of the training process. In this process, all intermediate perturbation images x generated from t=1 to T are... t These are collectively referred to as "latents," and their dimension is the same as the original image x0. The probability distribution used to define this forward diffusion kernel (FDK) is a normal / Gaussian distribution, with the following form: Where, β t It is called the diffusion scheduling coefficient, which is pre-set by the predefined variance scheduler; I is the identity matrix, so the perturbation distribution at each step is an isotropic Gaussian distribution.
[0120] Next, at each time step t, a small Gaussian noise ε is added to the original image, the intensity of which is determined by β in the scheduler. t Controlled by the settings. As the time step T increases, and with a well-defined β... t Under the scheduling strategy, repeated application of the diffusion kernel will gradually transform the original complex data distribution into a form that approximates an isotropic Gaussian distribution.
[0121] Referring to the following formula, the image x can be conveniently extracted from the normal distribution in the following way. t Perform sampling: Where ε represents the noise term, whose value is randomly sampled from the standard Gaussian distribution and is calculated from the previous time-stamped image x before being added. t -1 is used for appropriate scaling.
[0122] This mechanism starts from the original image x0 and gradually adds noise at each time step t=1 to T, thereby gradually perturbing and degrading the image. Ultimately, the image will evolve into a state close to pure noise.
[0123] However, whenever it is necessary to obtain a latent variable sample x at a certain time step t tIn this case, it is usually necessary to execute t sequentially in the Markov chain. -1 The forward sampling step is introduced. To address this computational inefficiency, reparameterized sampling is introduced in DDPM (Denoising Diffusion Probabilistic Models). The diffusion kernel is reparameterized and reconstructed, enabling it to directly jump from the initial image (i.e., time step 0) to any time step t for sampling without iterative steps. To achieve this skip sampling, two additional parameter terms are introduced to redefine the parameters from x0 to x... t Closed expression.
[0124] α t :=1-β t
[0125]
[0126] The above formula represents the cumulative product of α from step 1 to step t. Then, βt is replaced with α. t =1-β t Furthermore, utilizing the additive property of the Gaussian distribution, the forward diffusion process was rewritten so that it can be expressed based on α as follows:
[0127]
[0128] Using the above formula, sampling can be performed directly at any time step t in the Markov chain.
[0129] Next, the noise feature map I at each time step is... ti Input a pre-trained U-Net model to predict the noise component injected at that time step. ti The method employs an automatic hybrid precision (AMP) training approach to improve training efficiency, and the AdamW optimizer iteratively adjusts the neural network weights and biases to minimize the loss function.
[0130] Then, calculate the prediction noise ∈ ti With real injected noise ∈ tj The mean squared error (MSE) is used to optimize the U-Net parameters and reduce the prediction noise. ti Noise intensity confidence plot C converted to the current time step ti The local confidence level of the noise distribution is represented by the diffusion scheduling coefficient βt. Then, a time-weighted average of the confidence map is performed based on the diffusion scheduling coefficient βt to generate the global noise distribution probability map P. g .
[0131] Next, based on the global noise distribution probability map, according to the reverse time step sequence {t′ n ,t′ n-1The iteration is performed on the sequence ,…,t′1}, and at each step t′… i Based on the global noise distribution probability map Pg, the noise region that needs to be removed is located, and the inverse diffusion scheduling coefficient is used as the basis for this determination. Control the noise reduction intensity, gradually from The predicted noise component is stripped away.
[0132] Specifically, after forward diffusion is completed, the backward diffusion process is activated, with the goal of gradually recovering the semantic information of the image from Gaussian noise. This process is based on the Reverse Diffusion Kernel (RDK), and its execution steps and core principles are as follows:
[0133] The reverse diffusion process has the same functional form as the forward diffusion process. Similar to how the forward diffusion kernel (FDK) is defined as a normal distribution, the reverse diffusion kernel can be defined using the same functional form, but with a Gaussian distribution. The reverse diffusion process also constitutes a Markov chain, the difference being that at each time step, the neural network predicts the parameters of the reverse diffusion kernel. During the training phase, the parameter estimates learned by the network should be as close as possible to the true parameters of the posterior distribution of the FDK at each time step.
[0134] The Markov chain of backdiffusion starts from the position where forward diffusion ends, that is, from time step T, at which point the data distribution has been transformed into a (approximately) isotropic Gaussian distribution.
[0135]
[0136] p(x T ):=N(x t ;0,I)
[0137] The probability density function (PDF) of the backdiffusion process is an integral over all possible paths that start from pure noise xT and eventually arrive at a sample that is consistent with the original data distribution.
[0138] Finally, if the number of iterations reaches the preset maximum value or the enhanced image meets the defect visibility threshold, a high-definition defect image is output; otherwise, the next time step is executed.
[0139] S3. Send defect images to the annotation terminal and generate a dataset with contour annotations by performing multi-level quality checks on the feedback annotation results.
[0140] Furthermore, such as Figure 7 As shown, step S3 includes:
[0141] S31. Based on the historical annotation quality and efficiency data of the annotation personnel corresponding to each annotation terminal, classify the ability levels and associate them with differentiated baseline workloads. The baseline workloads of annotation personnel at the primary, intermediate and advanced ability levels increase progressively according to a preset ratio.
[0142] S32. Monitor the backlog of tasks to be labeled in real time. When the backlog of tasks drops to the preset threshold ratio of the baseline task volume, automatically trigger the task replenishment mechanism and re-allocate the baseline task volume to the terminal to match the capability level.
[0143] S33. Generate a dynamic priority allocation score based on the task backlog, processing rate, and terminal resource utilization of the labeled terminals.
[0144] S34. When a new batch of defective images to be labeled is detected, the new batch is divided into task segments that match the baseline task volume of each level according to the baseline task volume of the terminal capability level. The available labeling terminals are sorted from high to low according to the priority score, and the task segments are pushed to the corresponding labeling terminals according to the sorting priority.
[0145] In the image annotation process, a new automated push strategy based on manual annotation has been added. Based on personnel capabilities and task load, defective images are automatically pushed to different annotators' ports to improve annotation efficiency and quality, and to achieve dynamic balance and accurate allocation of tasks.
[0146] The new strategy for pushing manually annotated images is as follows: A capability-matching strategy is implemented based on the annotator's skill level. 1x images are pushed to novice annotators, 2x to intermediate annotators, and 3x to advanced annotators. When the remaining amount is less than 0.2x, images are pushed back to the initial value, and this process is repeated.
[0147] In the case of automatic task assignments, the difficulty of the tasks assigned to annotators is completely random. Therefore, it is necessary to monitor the workload of each annotator in real time and dynamically adjust the assignments, thus incorporating a task load balancing strategy. If an annotator has too many tasks, their priority is automatically reduced; if an annotator is currently idle, they are given priority to be assigned new tasks.
[0148] S35. Simultaneously push standard guidelines documents containing defect type option sets, contour annotation specifications, and example templates to the corresponding annotation terminals.
[0149] Synchronize the standard guidelines document with all annotation terminals. The document includes the following:
[0150] (1) Defect type option set: the defect definition and corresponding number in the "Technical Specification for Inspection and Evaluation of Urban Drainage Pipelines".
[0151] (2) Contour annotation specifications: require polygon closure, vertex coordinate normalization, and complete coverage of defect areas, etc.
[0152] (3) Sample template: Provides correctly labeled sample images and label files as a reference.
[0153] S36. Receive the annotation result file returned by the annotation terminal, parse the defect type label and contour coordinate data in the annotation result file, and perform format verification including file integrity and coordinate range validity.
[0154] S37. Match the defect type labels in the annotation results file with the category range of the defect type option set, and filter out undefined or out-of-bounds types.
[0155] S38. If the annotation result file passes the verification and matching, assign a qualified label and generate a standardized dataset containing image files, label files, and annotation files.
[0156] S39. If the annotation result file fails the inspection or matching, the defect information containing the image ID, defect type number, defect outline and defect description will be returned to the annotation terminal, the corresponding defect image will be locked, and the re-annotation process of the same annotation terminal or the reassignment process of the annotation terminal will be triggered according to the annotation quality determined by the defect information.
[0157] In a specific embodiment, the format of the annotation result file is validated to ensure file integrity (no missing fields or damage) and the validity of the annotation coordinates (contour coordinates are within the valid range of the image). The defect type of the annotation is checked to ensure it is within a preset category option set, filtering out undefined or out-of-bounds type labels. Geometric calculations are used to verify whether the annotation contour completely covers the actual defect area (e.g., ...). Figure 8 , Figure 9 Example), ensure that the defect range does not exceed the “selection” boundary (i.e., the annotation box completely surrounds the defect entity).
[0158] Next, a standardized dataset (containing image files, label files, and TXT format annotation files) is generated from the verified labeled data and written into the pipeline defect sample database, prohibiting subsequent modifications.
[0159] After the annotation work is completed, the annotated dataset requires quality control, so a dynamic adjustment strategy for annotation quality feedback is incorporated. If a defect is found, location information including the image ID and error type code is generated and fed back to the original annotation terminal, automatically locking the defective image and requesting correction. At this point, the severity of the problem is determined based on the error type code. For occasional errors, the terminal is triggered to re-annotate the defective data, while for high-frequency or systemic errors, the annotation terminal's reallocation mechanism is triggered.
[0160] Meanwhile, the historical quality inspection records of each annotation terminal are continuously tracked. By statistically analyzing quality indicators such as error rate and defect type distribution within a window period, a dynamic evaluation model is constructed. For terminals whose annotation quality has significantly declined recently, their task push priority is automatically reduced and the number of assignments is decreased; conversely, for terminals that maintain high-quality annotation and a defect rate below the threshold for multiple consecutive quality inspection cycles, their task priority is increased and the number of assignments is increased. This closed-loop mechanism drives the optimization of task allocation strategies through real-time feedback of quality inspection data, ensuring both the controllability of annotation quality and the optimal allocation of annotation resources.
[0161] S4. Construct an improved YOLO11-EfficientViT fusion model, load pre-trained weights, and train it on the dataset using the pipe diameter sensing training mode to generate a defect detection model with spatial awareness capabilities.
[0162] Furthermore, such as Figure 10 As shown, step S4 includes:
[0163] S41. Read images and annotation result files from the dataset with contour annotations, automatically divide it into training, validation, and test sets according to a preset ratio, and verify the correspondence between images and annotation result files. Read images and corresponding annotation files from the dataset with contour annotations, using 80% as the training set, 10% as the validation set, and the remaining 10% as the test set. The validation module checks whether each image has a matching annotation file, removing isolated files or samples with missing labels.
[0164] S42. Configure multiple information items including dataset path, defect type information, input image size, and defect category parameters, and simultaneously set pipe diameter sensing training parameters including receptive field scaling range and pipe diameter disturbance intensity.
[0165] Next, configure the dataset path (training set / validation set / test set directory) and defect type information (16 categories, including cracks, deformation, corrosion, misalignment, undulation, disjointness, interface material detachment, concealed branch pipe connection, foreign object penetration, leakage, deposition, scaling, obstacles, residual wall, dam root, tree root, and scum) in dataset. Then, set the defect type quantity information and input image size parameters in yolo11-seg.yaml. Simultaneously set the receptive field scaling range (e.g., 0.5×~2.0× baseline size) and pipe diameter disturbance intensity (e.g., ±15% diameter simulation deviation).
[0166] S43. Construct an improved YOLO11-EfficientViT network structure that includes a backbone network with configured EfficientViT modules and an embedded adaptive receptive field module detection head. Load the pre-trained weight file of YOLO11 from the pre-stored path, retain the weights of the backbone network, and randomly initialize the detection head module.
[0167] In this step, a YOLO11-EfficientViT network was constructed. Its backbone integrates the EfficientViT_M5 module, followed by an SPPF layer and a C2PSA attention module. Its detector head embeds an adaptive receptive field module, dynamically adjusting the feature fusion scale through upsampling, feature concatenation, and convolution. Then, pre-trained YOLO11 weights, such as yolo11-seg.pt, are loaded, and training parameters are set, including mixed-precision training, optimizer, and learning rate. Data augmentation algorithms are enabled to improve the model's generalization ability. Training logs and model weights are saved in a designated directory for later review and analysis.
[0168] S44. The model is trained based on the training set, and the pipe diameter sensing training mode is enabled during the training process. The size of the detection head receptive field is adjusted according to the scaling range based on the pipe diameter obtained from the annotation result file to match the defect scale of different pipe diameters. At the same time, pipe diameter perturbation is injected to generate multi-scale defect samples. The defect morphology is adjusted through affine transformation to simulate the imaging deformation under different pipe diameters.
[0169] In a specific embodiment, the pipe diameter sensing dynamic training includes: before training begins, parsing the pipe diameter information in the annotation file and establishing a diameter-receptive field mapping table; during forward propagation, adjusting the dilation rate of the detection head convolution kernel according to the pipe diameter of the input image and the scaling range to match the physical scale of the defect; then, injecting pipe diameter perturbations and randomly applying affine transformations (scaling ±15%, shearing ±10°) to the training images to simulate the morphological distortion of imaging with different pipe diameters, generating multi-scale defect samples. It is important to emphasize that mixed precision training (AMP) and the AdamW optimizer are enabled, with an initial learning rate set to 0.001, and 16 images are input per batch.
[0170] S45. After every N training rounds, the model is evaluated on the validation set to obtain evaluation metrics including segmentation accuracy, precision, and recall. The learning rate is then dynamically adjusted or early stopping is triggered. For example... Figure 11 As shown, the model was evaluated on the validation set, with key metrics including segmentation accuracy (mask mAP50), precision, and recall. Based on the evaluation results, parameter settings were adjusted or early stopping was triggered, and data augmentation strategies were optimized to improve the detection model's performance.
[0171] S46. After all training is complete, perform performance verification on the test set and convert the qualified models to the specified format. After training, automatically detect and export the model in ONNX format, such as yolo11-seg-pipeline defects.onnx.
[0172] S5. Load the trained model into the annotation module, configure the model path, defect type and annotation parameters, control the end-to-end annotation pipeline to perform batch defect detection and contour annotation on the new image data, and feed back to the dataset after multi-level quality verification to drive model iterative optimization.
[0173] Furthermore, such as Figure 12 As shown, step S5 includes:
[0174] S51. Create a configuration file for the trained defect detection model, and write the model type, model name, model path, and the set of pipeline defect types to be detected in the configuration file, and configure the automatic annotation parameters simultaneously. Among them, the automatic annotation parameters are: IOU (Intersection over Union) with a value between 0 and 1, and confidence ratio with a value between 0 and 1.
[0175] S52. Load the trained defect detection model and configuration file into the model management system, register the model name, version number and configuration file path for the annotation module to call.
[0176] S53. Scan the newly added drainage pipeline image data catalog, schedule it to the annotation module in batches, start the end-to-end annotation pipeline, call the model to perform batch defect detection and contour annotation, and perform multi-level quality verification. Specifically, the multi-level quality verification includes: Level 1 automatic verification, filtering detection results with confidence levels below the threshold, and verifying whether the defect type is in the configuration list; Level 2 logical verification, verifying the rationality of defect size based on pipe diameter parameters (e.g., crack length must not exceed 3 times the pipe diameter). In addition, a third level of manual sampling inspection can be set: randomly select 5% of the samples for quality inspectors to review the annotation accuracy.
[0177] S54. If the defect detection and contour annotation results pass the list matching and logical verification, the defect detection and contour annotation results are fed back to the corresponding annotation terminal. Based on the quality results fed back by the standard terminal, the defect detection and contour annotation results are corrected and written into the dataset, the dataset version is updated, and the model iterative training is triggered.
[0178] Furthermore, such as Figure 13 As shown, step S54 includes:
[0179] S541. Based on the defect detection and contour annotation results, classify the detected defect images according to boundary salience and semantic complexity, dividing them into object-type defects with clear boundaries and definition-type defects with ambiguous boundaries, and generate annotation processing priority markers based on the highest complexity of the defects in the image.
[0180] S542. Establish dynamic matching rules between the ability level of the labeling personnel and the defect type. Among them, the personnel with the junior ability level are assigned to handle definition-type defects and the baseline task volume is limited. The personnel with the intermediate ability level are assigned to handle definition-type defects and the baseline task volume is limited to a preset multiple of the baseline task volume. The personnel with the senior ability level are assigned to handle definition-type defects and the baseline task volume is limited.
[0181] S543. Monitor the backlog of tasks to be labeled in real time. When the backlog of tasks drops to the preset threshold ratio of the baseline task volume of the corresponding level, trigger the automatic supplementation mechanism to supplement the terminals with new tasks that match the corresponding defect type.
[0182] S544. Collect data on the distribution of labeling error types for personnel of each competency level in the defect result quality inspection process. For personnel whose labeling pass rate is greater than the preset pass standard for multiple consecutive batches, upgrade their competency level by at least one level. For personnel whose error rate is greater than the preset error standard for a single batch, reduce their competency level by at least one level and decrease their workload allocation ratio.
[0183] S545. Conduct multidimensional correlation analysis on quality assessment data, dynamic matching rules, and defect type distribution characteristics. Update dynamic matching rules and dataset versions based on the correlation analysis results to form a co-evolution mechanism for improving annotation quality and iterating model parameters.
[0184] S546. After the automatic image annotation stage, an automated push strategy based on artificial intelligence pre-classification is added. Based on the defect type, difficulty level, personnel ability profile, and task load, defect images are automatically pushed to different annotation personnel ports in a reasonable manner to improve the quality control efficiency and quality of automatic annotation and achieve dynamic balance and accurate allocation of tasks.
[0185] In another specific embodiment, a push strategy for quality inspection and rework is added after automatic labeling is completed, as follows:
[0186] Based on defect detection and contour annotation results, defect images are classified in a dual-dimensional manner: defect type allocation follows two visual perception angles: "form" and "spirit." "Form" specifically refers to defects with clearly defined boundaries, such as object-type defects like tree roots or obstacles; "spirit" specifically refers to definitional defects without clearly defined boundaries, such as cracks or deformations. "Spirit" is considered more complex than "form." Defect complexity is determined by the most difficult level of annotation within the image; an image may contain single or multiple defects, and the image's defect complexity is defined by the highest level of complexity found within the image.
[0187] A competency profile is established based on the annotators' qualifications, experience, and historical annotation accuracy, incorporating a competency-complexity matching strategy. Basic-level personnel are assigned to handle definitional defects with a limited baseline workload; intermediate-level personnel handle definitional defects with a workload limited to a preset multiple of the baseline workload, with the maximum workload capped at 1.5x times the baseline; advanced-level personnel handle definitional defects with the same baseline workload. Each person is assigned x images. When the remaining workload is less than 0.05x, images are pushed back to the corresponding upper limit for the annotators, and this process is repeated.
[0188] Collect error type distribution data during the quality inspection process: For terminals with a pass rate of >95% for three consecutive batches, automatically upgrade their capability level (e.g., from primary to intermediate) and expand their task permissions.
[0189] When the error rate of a single batch exceeds 10%, the process is immediately downgraded, and the task allocation is reduced by 30%. Simultaneously, a multi-dimensional analysis engine is built to correlate and mine the quality inspection error patterns, defect type distribution, and personnel capability profiles. For example, if a sudden increase in the mislabeling rate of a certain type of "problem" defect is detected on the primary terminal, the task allocation weight is automatically adjusted, prioritizing the routing of this type of defect to the advanced terminal. At the same time, updated annotation specifications for this defect type are generated. Ultimately, through iterative versioning of dynamic matching rules, the co-evolution of dataset quality optimization and model training performance is achieved.
[0190] S55. If the defect detection and contour annotation results fail the verification, record the error type and re-inject them into the annotation queue to wait for correction.
[0191] Additionally, this invention provides a high-quality automatic annotation system for drainage pipeline image data based on a deep learning model using multi-feature fusion. The system includes: a frame extraction module for extracting frames from the collected drainage pipeline image data; a filtering and enhancement module for filtering the extracted images based on defect visibility, performing forward diffusion noise injection, noise distribution prediction, and back diffusion iterative denoising on images that do not meet the visibility threshold to obtain defect images that meet the annotation requirements; an annotation module for sending defect images to an annotation terminal and generating a dataset with contour annotations by passing the feedback annotation results through multi-level quality checks; a model building module for building an improved YOLO11-EfficientViT fusion model, loading pre-trained weights, and training the dataset using a pipe diameter-aware training mode to generate a defect detection model with spatial awareness; and a model deployment module for loading the trained model into the annotation module, configuring the model path, defect type, and annotation parameters, controlling the end-to-end annotation pipeline to perform batch defect detection and contour annotation on newly added image data, and feeding back the results to the dataset after multi-level quality checks to drive iterative model optimization.
[0192] Furthermore, embodiments of the present invention provide an automatic annotation device for high-quality drainage pipeline image data based on a deep learning model of multi-feature fusion, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the automatic annotation method for high-quality drainage pipeline defect samples driven by multi-feature fusion as described above.
[0193] Furthermore, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the multi-feature fusion-driven automatic annotation method for high-quality drainage pipe defect samples as described above.
[0194] In summary, the embodiments of the present invention provide a method, system, device, and medium for automatic annotation of high-quality drainage pipeline defect samples driven by multi-feature fusion, such as... Figure 14 As shown, the process is implemented through the following steps: First, video data of the inside of the drainage pipe is collected by a pipe inspection robot or drone. FFmpeg is used to extract frames at a frequency of 3 frames / second to generate images, and images containing defects are selected. For blurry or low-contrast defect images, a Denoising Diffusion Probabilistic Model (DDPM) is used for enhancement: noise is gradually injected during the forward diffusion stage to generate a multi-time-step noise feature sequence. Then, U-Net is used to predict the noise distribution and perform reverse iterative denoising, finally outputting the enhanced defect image.
[0195] Next, the enhanced image is uploaded to the annotation module, where professionals annotate the defect types and outline ranges according to industry standards. The annotation results undergo multi-level quality checks: automatic checks verify label compliance and outline closure, manual spot checks verify defect coverage, unqualified samples are returned for correction, and qualified samples are stored in the defect sample database to ensure data consistency and tamper-proofness.
[0196] Subsequently, an improved YOLO11 model integrating the EfficientViT module was constructed, pre-trained weights were loaded, and the dataset was partitioned. During training, a pipe diameter sensing mode was enabled, dynamically adjusting the receptive field of the detection head according to the pipe diameter, and random affine transformations were injected to simulate different pipe diameter scenarios. The model was iteratively optimized through mixed-precision training and the AdamW optimizer. Segmentation accuracy, precision, and recall were evaluated on the validation set, and the ONNX format model was exported after achieving the required scores.
[0197] Finally, the model is deployed to the annotation system, automatic annotation parameters are configured, and new image data is processed in batches. After multi-level verification of the detection results, qualified annotated data triggers iterative training of the model, while samples that fail are recorded as error types and re-injected into the correction queue, forming a closed-loop optimization chain of "annotation-training-re-annotation".
[0198] In summary, this invention addresses the industry pain points of low data quality, poor annotation efficiency, and weak model generalization in drainage pipeline defect detection through a collaborative end-to-end approach involving data augmentation, intelligent annotation, and model optimization. Specifically:
[0199] (1) To address the issues of blurred and high-noise images caused by darkness, dampness, and interference from the shooting light source inside the pipe, a denoising diffusion probability model is introduced. Noise is gradually injected through forward diffusion to generate latent space features. Then, the noise distribution is predicted by the U-Net segmentation model, and reverse iterative denoising is performed, which significantly improves the clarity and contrast of low-quality images. This technical feature increases the proportion of usable defective data and greatly improves the utilization rate of image data.
[0200] (2) The enhanced data is input into the improved YOLO11-EfficientViT fusion model. Through pipe diameter sensing training and adaptive receptive field module, the model’s adaptability to complex scenarios is enhanced. After training, the ONNX format detection model is exported and deployed to the automatic annotation module to achieve one-click batch defect detection and contour generation. The annotation efficiency of manual annotation is greatly improved.
[0201] (3) After the automatic annotation results are verified at multiple levels, qualified data triggers iterative training of the model, while unqualified samples are manually corrected and re-injected into the training set. By combining transfer learning and incremental learning strategies, the model performance is continuously optimized. On this basis, automatic annotation and manual collaborative correction of secondary samples are supported, forming a dynamic optimization link of "data augmentation-annotation-training-re-annotation", which drives the continuous improvement of detection accuracy and generalization ability.
[0202] Therefore, this invention constructs a complete technical system for the intelligent operation and maintenance of urban underground pipe networks, from data quality improvement and efficient annotation to model self-evolution. While reducing labor costs, it ensures the accuracy, robustness and efficiency of defect detection and engineering implementation.
[0203] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0204] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the claims should be interpreted to include both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0205] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, then this invention should also include these modifications and variations.
Claims
1. A method for automatically labeling high-quality drainage pipeline defect samples driven by multi-feature fusion, characterized in that, include: The collected drainage pipe image data is processed by frame extraction. Defect visibility screening is performed on the extracted images. For images that do not meet the visibility threshold, forward diffusion noise injection, noise distribution prediction, and back diffusion iterative denoising are performed to obtain defect images that meet the annotation requirements. The defect images are sent to the annotation terminal, and the feedback annotation results are used to generate a dataset with contour annotations through multi-level quality verification. An improved YOLO11-EfficientViT fusion model was constructed, pre-trained weights were loaded, and the model was trained on the dataset using the pipe diameter-aware training mode to generate a defect detection model with spatial awareness capabilities. This included: reading images and annotation result files from the dataset with contour annotations, automatically dividing the training set, validation set, and test set according to a preset ratio, and verifying the correspondence between the images and the annotation result files. The configuration includes multiple parameters such as dataset path, defect type, input image size, and defect category. Pipe-aware training parameters, including receptive field scaling range and pipe diameter perturbation intensity, are set simultaneously. An improved YOLO11-EfficientViT network structure is constructed, comprising a backbone network with configured EfficientViT modules and an embedded adaptive receptive field module detection head. Pre-trained YOLO11 weight files are loaded from a pre-stored path, retaining the backbone network weights, and the detection head module is randomly initialized. The model is trained on the training set, and a pipe-aware training mode is enabled during training to adjust the detection head's receptive field size according to the pipe diameter parsed from the annotation result file, matching the defect scale for different pipe diameters. Pipe diameter perturbations are injected to generate multi-scale defect samples, and affine transformations are used to adjust the defect morphology, simulating imaging deformation under different pipe diameters. After every N training epochs, the model is evaluated on the validation set to obtain evaluation metrics including segmentation accuracy, precision, and recall. The learning rate is dynamically adjusted or early stopping is triggered. After all training is completed, performance is verified on the test set, and successful models are converted to a specified format. The trained model is loaded into the annotation module, and the model path, defect type and annotation parameters are configured. The end-to-end annotation pipeline is controlled to perform batch defect detection and contour annotation on the new image data. After multi-level quality verification, the results are fed back to the dataset to drive model iterative optimization.
2. The method for automatic annotation of high-quality drainage pipeline defect samples driven by multi-feature fusion as described in claim 1, characterized in that, The frame extraction process for the collected drainage pipe image data includes: Pipeline inspection robots, crawling robots, or drones are used to enter the drainage pipeline from the manhole inlet or open port and collect internal images along the longitudinal axis of the pipeline to obtain pipeline inspection video data with continuous spatial coverage. Import the pipeline inspection video data into the computer storage medium, and establish a hierarchical storage directory according to the spatiotemporal information of the acquisition to form the original video database; Configure the video parsing framework through the command-line interface to determine the mapping relationship between the source video path and the frame-extracted image output path; Configure image output format parameters and frame rate parameters; among them, the frame rate parameters are generated by real-time analysis of video data resolution, image quality assessment value and pipe cross-section coverage, and after multi-dimensional weight fusion, they are mapped to a preset frame rate range to generate frame rate parameters that are adaptive to different pipe working conditions. The system recursively traverses the storage directory of the original video database, identifies the target video file based on the mapping relationship between the source video path and the output path of the extracted frame image, creates a subdirectory tree with the same name, calls the video decoding program to perform time-series sampling according to the determined frame extraction rate parameter, and then uses serialization naming rules to store the extracted frame images to the corresponding subdirectories.
3. The method for automatic annotation of high-quality drainage pipeline defect samples driven by multi-feature fusion as described in claim 1, characterized in that, Defect visibility screening is performed on the extracted images. Images that do not meet the visibility threshold are subjected to forward diffusion noise injection, noise distribution prediction, and back diffusion iterative denoising to obtain defect images that meet the annotation requirements, including: Based on a preset defect visibility threshold, the extracted images are automatically filtered. If the defect visibility threshold is met, the result is directly output to the subsequent annotation process. If the defect visibility threshold is not met, progressive noise injection is performed on the corresponding image. The noise intensity is gradually superimposed according to a Gaussian distribution within a preset time step sequence using a preset diffusion model. The noise injection intensity is controlled by the diffusion scheduling coefficient, and the image is gradually degraded into a multi-time step noise feature sequence. The noise feature map of each time step in the multi-time step noise feature sequence is input into the pre-trained U-Net noise prediction model to predict the noise component injected at each time step during the forward diffusion process. By comparing the mean square error of the predicted noise component injected at each time step with the actual injected noise, the noise intensity confidence map of each time step is generated and then fused by time-series weighting to obtain the global noise distribution probability map. Based on the global noise distribution probability map, the back diffusion process is executed in an iterative manner: in the back time step sequence, the predicted noise component is gradually stripped from the noise feature sequence of multiple time steps, and the noise removal intensity is controlled by the back diffusion scheduling coefficient. The image details are gradually reconstructed through multiple rounds of iteration. When the preset maximum number of iterations is reached or the defect visibility threshold is met, a defect image that meets the annotation requirements is generated. Among them, the diffusion scheduling coefficient is a time-dependent parameter used to control the noise injection during the forward process of the diffusion model, and the reverse diffusion scheduling coefficient is a time-dependent parameter used to control the denoising intensity at each step during the reverse diffusion process, and satisfies the reverse correlation with the diffusion scheduling coefficient.
4. The method for automatic annotation of high-quality drainage pipeline defect samples driven by multi-feature fusion as described in claim 1, characterized in that, The defect images are sent to the annotation terminal, and the returned annotation results are processed through multi-level quality checks to generate a dataset with contour annotations, including: Based on the historical annotation quality and efficiency data of the annotation personnel corresponding to each annotation terminal, the ability levels are divided and associated with differentiated baseline workloads. Among them, the baseline workloads of annotation personnel at the primary, intermediate and advanced ability levels increase progressively according to a preset ratio. The system monitors the backlog of tasks to be labeled in real time. When the backlog of tasks drops to the preset threshold ratio of the baseline task volume, the system automatically triggers the task replenishment mechanism and re-allocates the baseline task volume that matches the capability level to the terminal. Based on the task backlog, processing rate, and terminal resource utilization of the labeled terminals, a dynamic allocation priority score is generated. When a new batch of defective images to be labeled is detected, the new batch is divided into task segments that match the baseline workload of each level according to the baseline workload of the terminal capability level. The available labeling terminals are sorted from high to low according to the priority score, and the task segments are pushed to the corresponding labeling terminals according to the sorting priority. Simultaneously push standard guidelines documents containing defect type option sets, contour annotation specifications, and example templates to the corresponding annotation terminals; Receive the annotation result file returned by the annotation terminal, parse the defect type label and contour coordinate data in the annotation result file, and perform format checks including file integrity and coordinate range validity; Based on the defect type labels in the annotation results file, match them with the category range of the defect type option set to filter out undefined or out-of-bounds types; If the labeled result file passes the verification and matching, it is given a qualified label, and a standardized dataset containing image files, label files, and annotation files is generated; If the annotation result file fails the inspection or matching, the defect information is returned to the annotation terminal, the corresponding defect image is locked, and the re-annotation process of the same annotation terminal or the reassignment process of the annotation terminal is triggered.
5. The method for automatic annotation of high-quality drainage pipeline defect samples driven by multi-feature fusion as described in claim 1, characterized in that, The trained model is loaded into the annotation module, and the model path, defect type, and annotation parameters are configured. The end-to-end annotation pipeline is controlled to perform batch defect detection and contour annotation on newly added image data. After multi-level quality verification, the results are fed back to the dataset to drive iterative optimization of the model, including: Create a configuration file for the trained defect detection model, and write the model type, model name, model path and the set of pipeline defect types to be detected in the configuration file, and configure the automatic annotation parameters simultaneously. Load the trained defect detection model and configuration file into the model management system, register the model name, version number and configuration file path for the annotation module to call; Scan the newly added drainage pipeline image data catalog, schedule it to the annotation module in batches, start the end-to-end annotation pipeline, call the model to perform batch defect detection and contour annotation, and perform list matching verification and logic verification; If the defect detection and contour annotation results pass the list matching and logical verification, the defect detection and contour annotation results are fed back to the corresponding annotation terminal. Based on the quality results fed back by the standard terminal, the defect detection and contour annotation results are corrected and written into the dataset, the dataset version is updated, and model iteration training is triggered. If the defect detection and contour annotation results fail the verification, record the error type and re-inject the data into the annotation queue to await correction.
6. The method for automatic annotation of high-quality drainage pipeline defect samples driven by multi-feature fusion as described in claim 4, characterized in that, If the defect detection and contour annotation results pass the list matching and logical verification, the defect detection and contour annotation results are fed back to the corresponding annotation terminal. Based on the quality results fed back by the standard terminal, the defect detection and contour annotation results are corrected and written into the dataset. The dataset version is updated, and model iterative training is triggered, including: Based on the defect detection and contour annotation results, the detected defect images are classified according to boundary salience and semantic complexity, into object-type defects with clear boundaries and definition-type defects with vague boundaries, and annotation processing priority labels are generated according to the highest complexity of the defects in the image. Establish dynamic matching rules between the ability level of the labeling personnel and the defect type. Among them, the personnel with the basic ability level are assigned to handle definition-type defects and are limited to the baseline task volume; the personnel with the intermediate ability level are assigned to handle definition-type defects and are limited to a preset multiple of the baseline task volume; and the personnel with the advanced ability level are assigned to handle definition-type defects and are limited to the baseline task volume. The system monitors the backlog of tasks to be labeled in real time. When the backlog of tasks drops to the preset threshold ratio of the baseline task volume for the corresponding level, an automatic supplementation mechanism is triggered to supplement new tasks that match the corresponding defect type to terminals with each capability level. Collect data on the distribution of labeling error types of personnel at each capability level in the defect result quality inspection process. For personnel whose labeling pass rate is greater than the preset pass standard for multiple consecutive batches, upgrade their capability level by at least one level. For personnel whose error rate is greater than the preset error standard for a single batch, reduce their capability level by at least one level and reduce their workload allocation ratio. A multidimensional correlation analysis is performed on quality assessment data, dynamic matching rules, and defect type distribution characteristics. Based on the correlation analysis results, the dynamic matching rules and dataset versions are updated to form a co-evolutionary mechanism for improving annotation quality and iterating model parameters.
7. A high-quality automatic annotation system for drainage pipeline image data based on a deep learning model of multi-feature fusion, characterized in that, include: The frame extraction processing module is used to perform frame extraction processing on the collected drainage pipeline image data. The filtering and enhancement module is used to filter the defect visibility of the frame-stripped image. For images that do not meet the visibility threshold, forward diffusion noise injection, noise distribution prediction and back diffusion iterative denoising are performed to obtain defect images that meet the annotation requirements. The annotation module is used to send defect images to the annotation terminal and generate a dataset with contour annotations from the feedback annotation results through multi-level quality verification. The model building module is used to build an improved YOLO11-EfficientViT fusion model, load pre-trained weights and train the model on the dataset using the pipe diameter-aware training mode to generate a defect detection model with spatial awareness capabilities. This includes: reading images and annotation result files from the dataset with contour annotations, automatically dividing the training set, validation set and test set according to a preset ratio, and verifying the correspondence between the images and the annotation result files. The configuration includes multiple parameters such as dataset path, defect type, input image size, and defect category. Pipe-aware training parameters, including receptive field scaling range and pipe diameter perturbation intensity, are set simultaneously. An improved YOLO11-EfficientViT network structure is constructed, comprising a backbone network with configured EfficientViT modules and an embedded adaptive receptive field module detection head. Pre-trained YOLO11 weight files are loaded from a pre-stored path, retaining the backbone network weights, and the detection head module is randomly initialized. The model is trained on the training set, and a pipe-aware training mode is enabled during training to adjust the detection head's receptive field size according to the pipe diameter parsed from the annotation result file, matching the defect scale for different pipe diameters. Pipe diameter perturbations are injected to generate multi-scale defect samples, and affine transformations are used to adjust the defect morphology, simulating imaging deformation under different pipe diameters. After every N training epochs, the model is evaluated on the validation set to obtain evaluation metrics including segmentation accuracy, precision, and recall. The learning rate is dynamically adjusted or early stopping is triggered. After all training is completed, performance is verified on the test set, and successful models are converted to a specified format. The model deployment module is used to load the trained model into the annotation module, configure the model path, defect type and annotation parameters, control the end-to-end annotation pipeline to perform batch defect detection and contour annotation on new image data, and feed back to the dataset after multi-level quality verification to drive model iterative optimization.
8. A high-quality automatic annotation device for drainage pipeline image data based on a deep learning model of multi-feature fusion, characterized in that, include: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores instructions that can be executed by at least one processor, which are executed by at least one processor to enable the at least one processor to perform the multi-feature fusion-driven automatic annotation method for high-quality drainage pipe defect samples as described in any one of claims 1-6.
9. A computer-readable storage medium storing computer-executable instructions thereon, characterized in that, When the executable instructions are executed by the processor, they implement the high-quality drainage pipeline defect sample automatic annotation method driven by multi-feature fusion as described in any one of claims 1-6.
Citation Information
Patent Citations
Tunnel water seepage data set generation method based on de-noising diffusion probability model
CN119169410A
Training data determination method and apparatus, target detection method and apparatus, device, and medium
WO2024193681A1