Intelligent defect identification and duplicate removal method for drainage pipeline detection video
By combining target tracking technology and model confidence analysis, the sliding window mechanism is used to automatically screen the defect keyframes of the drainage pipeline to detect video, solving the problems of large errors and low efficiency in the existing technology, and achieving efficient and accurate defect identification and evaluation.
Patent Information
- Application Number
- CN202510691941.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The reliance on manual interpretation in existing drainage pipeline inspection results in large errors, strong subjectivity, time-consuming and low efficiency. The existing AI detection still requires manual recording of defect types and locations, and there is a lack of a method of automatically screening keyframes.
The pre-trained pipeline defect recognition model is used to combine target tracking technology and model confidence analysis, and set the effective distance interval, use the sliding window mechanism to select the most representative and accurate recognition video frames, and automatically filter the defect keyframes.
It realizes efficient and accurate extraction of representative defect keyframes from a large number of detected video frames, reduces manual intervention, and improves the efficiency and accuracy of defect identification and evaluation.
Smart Images

Figure CN120598889A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pipeline detection, and in particular to a method for intelligent defect recognition and deduplication in drainage pipeline detection videos. Background Art
[0002] The existing inspection and evaluation of drainage pipes are mainly based on manual interpretation, which has the main disadvantages of high interpretation error, strong subjectivity, time-consuming and labor-intensive, high cost and low efficiency.
[0003] The current YOLO-based computer vision deep learning artificial intelligence model can quickly and effectively identify pipeline defects in underground drainage pipeline inspection videos, effectively reducing the manual review of the types and number of defects in drainage pipeline image data. However, it is still necessary to manually record the type, number, and distance position of the defects in the pipeline, including the following steps: 1) Collect image data of all pipeline defect types; 2) Screen pipeline defect images; 3) Annotate images of all pipeline types; 4) Form a pipeline defect sample database; 5) Build and train an image segmentation AI intelligent detection model framework; 6) Construct an automatic detection, identification, tracking, and statistical model for drainage pipeline defects; 7) Output the interpretation video and statistical results. Even with the current use of AI automatic interpretation, it is still necessary to manually continue to count the types and numbers of defects and the longitudinal position of the drainage pipeline, from which to select the key frame images with the best recognition effect for the defects.
[0004] Therefore, how to further reduce manual work and realize automatic screening of defect key frame images has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] (1) Technical issues to be resolved
[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a method for intelligent defect recognition and deduplication in drainage pipe inspection videos, which accurately selects the most representative defect key frames and improves the efficiency and accuracy of defect target recognition and evaluation.
[0007] (2) Technical solution
[0008] In order to achieve the above objectives, the main technical solutions adopted by the present invention include:
[0009] In a first aspect, an embodiment of the present invention provides a method for intelligently identifying and deduplicating defects in drainage pipe inspection videos, comprising:
[0010] S100: Obtain each frame of an image and a frame number in a video of a town drainage pipeline captured by a pipeline robot, wherein each frame of the image is annotated with attribute information collected by the pipeline robot;
[0011] S200, analyzing each frame of the image using a pre-trained pipeline defect recognition model to obtain a recognition result for each frame of the image, the recognition result including: a category of the pipeline defect in the current frame, a confidence level of each pipeline defect, a pipeline defect mask area, and tracking number information for tracking the pipeline defect in the current frame;
[0012] S300, extracting a valid frame image to which each pipeline defect belongs and a motion trajectory of the pipeline defect based on distance information in attribute information of all frame images and the recognition result of the pipeline defect;
[0013] S400: Comprehensively analyze the motion trajectory of the pipeline defect corresponding to each valid frame image of the pipeline defect, and the category, confidence level, and mask area of the pipeline defect in the recognition result corresponding to each valid frame image, to select a defect key frame image set for pipeline defect recognition and tracking analysis.
[0014] The method for intelligent defect identification and deduplication in drainage pipe inspection videos according to an embodiment of the present invention combines target tracking technology, identification distance information and model confidence analysis, sets an effective distance interval and uses a sliding window mechanism to select the most representative and most accurate video frame of each defect target in the motion trajectory. Compared with the existing technology, it can automatically, accurately and efficiently extract key frame data (the best representative frame number of each defect and the distance value corresponding to the frame) that can represent the characteristics of each defect in automatic detection and the best recognition effect from a large number of inspection video frames, thereby facilitating subsequent analysis and evaluation of the defects.
[0015] Optionally, before S200, the method further includes:
[0016] The pipeline defect training data set is used to train and verify the pipeline defect recognition model to obtain the trained pipeline defect recognition model;
[0017] Each image in the pipeline defect training dataset is pre-labeled with a defect category and a defect mask area; and for images of a single defect or multiple defect categories, an image enhancement algorithm is used to process the images to obtain training data that meets the preset image quality.
[0018] Optionally, an image enhancement algorithm is used to obtain training data that meets a preset image quality, including:
[0019] Use DnCNN and SRCNN to process images with single defect or multiple defect categories;
[0020] DnCNN includes the first input layer, the middle layer and the first output layer.
[0021] The first input layer extracts features from the input image using the following formula:
[0022]
[0023] Among them, X is the input image, and its dimension is
[0024] Conv2d is a two-dimensional convolution operation, k1=5 means its convolution kernel is 5×5, d=2 means its stride is 2, g=3 means the number of groups in its grouped convolution is 3, padding=same means its padding method is the same, (3→32) means its convolution operation converts 3 channels into 32 channels,
[0025] tanh is the hyperbolic tangent function, F1(X) is the output of the first input layer,
[0026] The middle layer consists of 10 alternating convolutional layers.
[0027]
[0028] in, is the output of each layer in the middle layer, X m is the input of each layer in the middle layer,
[0029] Represents a two-dimensional convolution operation, the convolution kernel k2 is 3×3, the padding value p1 is 1, and the 32 channels are converted to 64 channels.
[0030] Represents a two-dimensional convolution operation, the convolution kernel k3 is 3×3, the padding value p2 is 1, and the 64 channels are converted to 32 channels.
[0031] BatchNorm means batch normalization, LeakyReLU is the activation function,
[0032] F1(X) is the output of the first input layer, F m (X) is the output of the middle layer;
[0033] The first output layer is,
[0034]
[0035] Among them, F m (X) is the output of the middle layer, f DnCNN (X) is the output of the first output layer,
[0036] Represents a two-dimensional convolution operation, with the convolution kernel k4 being 5×5 and the padding value p3 being 2, converting 32 channels into 3 channels;
[0037] SRCNN takes the output of DnCNN as input, and the formula is as follows:
[0038]
[0039] Among them, f SRCNN (Y) is the output of SRCNN, Y is the input of SRCNN,
[0040] Represents a two-dimensional convolution operation, the convolution kernel k5 is 9×9, the padding value p4 is 4, and the 3 channels are converted to 64 channels.
[0041] Represents a two-dimensional convolution operation, the convolution kernel k6 is 1×1, the padding value p5 is 0, and the 64 channels are converted to 64 channels.
[0042] Represents a two-dimensional convolution operation, the convolution kernel k7 is 5×5, the padding value p6 is 2, and the 64 channels are converted to 3 channels.
[0043] ReLU stands for Rectified Linear Function.
[0044] Optionally, the attribute information of each frame of image includes: time, robot identification, distance information, and pipeline identification annotated on each frame of image;
[0045] The S300 includes:
[0046] S301, automatically identifying attribute information in each frame of image;
[0047] S302 : Extracting the valid frame image to which each pipeline defect belongs and the motion trajectory of the pipeline defect based on the recognition result of each frame image and the distance information in the attribute information.
[0048] Optionally, the step 301 includes:
[0049] The attribute information of each frame of the pipeline video is obtained by using OCR recognition technology, and the attribute information is parsed based on regular expressions to obtain distance information.
[0050] Optionally, the step 302 includes:
[0051] S302-1. For each defect, obtain the image set T where the defect appears based on the category and tracking number information of the pipeline defect in the recognition result. k ,
[0052] T k ={(F i , ID i , D i , C i ,Pi )|i=1,2,...,N}
[0053] Among them, F i is the frame number of the i-th frame image, ID i is the defect category and tracking number information corresponding to the i-th frame image, D i is the recognition distance in the distance information of the i-th frame image and satisfies the ascending order of distance, that is, D1≤D2≤…≤D N , C i is the confidence level of the model recognition corresponding to the i-th frame image, 0≤C i ≤1, P i is the i-th frame image,
[0054] S302-2. Image set T based on defect occurrence k The distance information in the image is used to obtain the motion trajectory of the defect.
[0055] D total =D N -D1
[0056] Among them, D total is the total movement distance of the defect, D N is the identification distance in the distance information of the image where the defect finally appears, and D1 is the identification distance in the distance information of the image where the defect initially appears;
[0057] S302-3. Extract valid frame images from the defective image set according to the following formula:
[0058]
[0059] in, is a valid frame image set,
[0060] D min is the effective minimum distance, D min =D1+0.7×D total ,
[0061] D max is the effective maximum distance, D max ×D1+0.9×D total .
[0062] Optionally, the S400 includes:
[0063] S401. In the valid frame image set, a fixed window length W is used to form (M-W+1) smoothing windows, where M is the number of valid frame images in the valid frame image set. The average confidence of each sliding window is calculated using the following formula:
[0064]
[0065] Among them, MC j is the average confidence of the j-th sliding window, C i is the confidence of the model recognition corresponding to the i-th frame image;
[0066] S402: From all smooth windows, select the window with the largest average confidence as the optimal window.
[0067]
[0068]
[0069] Among them, j * is the j value of the optimal window, Indicates the parameter value that maximizes the following expression within the range of 1≤j≤M-W+1. represents the image set within the optimal window;
[0070] S403: Select the image with the highest confidence level from the image set of the optimal window as the key frame image of the defect.
[0071]
[0072] Among them, F best Indicates the frame number of the key frame image, Indicates The parameter value that makes the following expression obtain the maximum value within the range, C i is the confidence of the model recognition corresponding to the i-th frame image;
[0073] S404: Construct a key frame image set based on the key frame image of each defect.
[0074] Optionally, the pipeline defect recognition model is a pipeline defect recognition model based on the YOLO12 algorithm, including a second input layer, a feature extraction layer, a feature fusion layer, an upsampling layer, and a second output layer;
[0075] The second input layer is used to receive the input drainage pipe image.
[0076] The feature layer includes convolution layer, C3k2 layer, A2C2f layer, which are used for preliminary feature extraction and attention processing.
[0077] The fusion layer is used to splice the feature maps output by different levels of the feature layer.
[0078] The upsampling layer is used to increase the resolution of the feature map output by the fusion layer.
[0079] The second output layer is used to output segmentation and detection results.
[0080] In a second aspect, an embodiment of the present invention provides a device for intelligently identifying and deduplicating defects in drainage pipe inspection videos, comprising:
[0081] An image acquisition module is used to obtain each frame of the video of the urban drainage pipeline collected by the pipeline robot and its frame number. Each frame of the video is marked with the attribute information of the pipeline robot at the time of collection;
[0082] An analysis and recognition module is used to analyze each frame of the image using a pre-trained pipeline defect recognition model to obtain a recognition result for each frame of the image. The recognition result includes: the category of the pipeline defect in the current frame, the confidence level of each pipeline defect, the pipeline defect mask area, and tracking number information for tracking the pipeline defect in the current frame;
[0083] An effective frame extraction module is used to extract the effective frame image of each pipeline defect and the motion trajectory of the pipeline defect based on the distance information in the attribute information of all frame images and the recognition result of the pipeline defect;
[0084] The screening module is used to comprehensively analyze the motion trajectory of the pipeline defect corresponding to each valid frame image of the pipeline defect, as well as the category, confidence level, and pipeline defect mask area of the recognition result corresponding to each valid frame image, and screen out the defect key frame image set for pipeline defect identification and tracking analysis.
[0085] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program in the memory, specifically performing the steps of the method for intelligent defect identification and deduplication in drainage pipe inspection videos as described in any one of the first aspects above.
[0086] (3) Beneficial effects
[0087] The beneficial effects of the present invention are as follows: a method for intelligent defect recognition and deduplication for drainage pipe inspection videos of the present invention, due to the combination of target tracking technology, recognition distance information and model confidence analysis, sets an effective distance interval and uses a sliding window mechanism to select the most representative and most accurate video frame of each defect target in the motion trajectory. Compared with the existing technology, it can automatically, accurately and efficiently extract key frame data (the best representative frame number of each defect and the distance value corresponding to the frame) that can represent the characteristics of each defect in automatic detection and the best recognition effect from a large number of inspection video frames, which is convenient for subsequent analysis and evaluation of the defects. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1A flowchart of a method for intelligent defect identification and deduplication in drainage pipe inspection videos provided by one embodiment of the present invention;
[0089] Figure 2 This is a schematic diagram of the structure of the pipeline defect recognition model;
[0090] Figure 3 A flowchart of a method for intelligent defect identification and deduplication in drainage pipe inspection videos provided by another embodiment of the present invention;
[0091] Figure 4 The image with defects used for training the model before being processed by the image enhancement algorithm;
[0092] Figure 5 The image with defects processed by the image enhancement algorithm is used to train the model;
[0093] Figure 6 is a schematic diagram of an image frame carrying a defect stored in the first stage according to another embodiment of the present invention;
[0094] Figure 7 A schematic diagram of one of the image frames with defect annotations stored in the first stage according to another embodiment of the present invention;
[0095] Figure 8 FIG. 1 is a schematic diagram of a video interpretation stored in the first stage according to another embodiment of the present invention. DETAILED DESCRIPTION
[0096] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0097] This embodiment relates to a video target detection and tracking method based on computer vision and artificial intelligence technology, and in particular to a technical method for instance segmentation of defect targets, continuous multi-target tracking, and automatic recognition of numerical information in images during the detection of urban drainage pipelines.
[0098] This example uses a deep learning model based on the YOLO12 algorithm, suitable for segmentation, detection, identification, tracking, and statistical modeling of drainage pipe defect types. Based on image data of pipe defects, the automatic detection and identification model is trained on the YOLO12 algorithm to improve detection and identification accuracy and stability.
[0099] This embodiment uses a model to automatically interpret the pipeline video and automatically records the number and longitudinal positions of defects in the pipeline.
[0100] The pipeline videos used for screening and the pipeline defect training data sets used for training and verifying the pipeline defect recognition model in this embodiment are all derived from pipeline robots. For example, they can be drainage pipe image data collected by pipeline inspection robots, crawling robots equipped with cameras, or drones. These devices enter from rainwater wells, sewage wells, or open pipe ports, advance longitudinally along the inside of the pipe, and simultaneously capture video or image data from the inside of the pipe. After reaching the other end of the pipe, the device returns to the ground with the recorded image data. After the acquisition is completed, the shooting device is connected to the port of the host device (such as a computer) via a data cable, and the drainage pipe image data (generally in video formats such as mp4 and avi) is imported and stored in the host device.
[0101] Example 1
[0102] This embodiment provides a method for intelligent defect recognition and deduplication for drainage pipe inspection videos, which is used to identify pipe defects from drainage pipe videos and select key frame images with the best defect recognition effect for each defect. The method of this embodiment can be implemented on any host computer device. Figure 1 , the method of this embodiment includes:
[0103] S100: For a pipeline video of a town drainage pipeline collected by a pipeline robot, obtain each frame image and frame sequence number in the pipeline video, wherein each frame image is marked with attribute information collected by the pipeline robot.
[0104] Specifically, the attribute information of each frame of image includes: time, robot identification, distance information, pipeline identification, etc. marked on each frame of image. These attribute information are automatically added during the pipeline robot acquisition process.
[0105] In step S100, the host computer automatically reads the video frame, obtains its width, height and frame rate information, and creates a corresponding video storage path and image output directory based on this to store the video and image data generated during the detection and recognition process, as well as the corresponding text recognition and tracking information.
[0106] S200: Analyze each frame of image using a pre-trained pipeline defect recognition model to obtain a recognition result for each frame of image. The recognition result includes: the category of the pipeline defect in the current frame, the confidence level of each pipeline defect, the pipeline defect mask area, and tracking number information for tracking the pipeline defect in the current frame. The tracking number information is unique for each defect.
[0107] Specifically, see Figure 2 , the pipeline defect recognition model may be a pipeline defect recognition model obtained based on the YOLO12 algorithm, including a second input layer, a feature extraction layer, a feature fusion layer, an upsampling layer, and a second output layer;
[0108] The second input layer is used to receive the input drainage pipe image.
[0109] The feature extraction layer includes a convolutional layer, a C3k2 layer, and an A2C2f layer, which are used to perform preliminary feature extraction and attention processing.
[0110] The feature fusion layer is used to splice the feature maps output by different levels of the feature layer.
[0111] The upsampling layer is used to increase the resolution of the feature map output by the fusion layer.
[0112] The second output layer is used to output segmentation and detection results.
[0113] In subsequent steps, this embodiment extracts the best valid frame and confidence analysis of the detected pipeline defect based on the above detection and recognition results.
[0114] In step S200, the pre-trained pipeline defect recognition model is loaded and the input video data is analyzed frame by frame to achieve automatic instance segmentation and multi-target tracking of drainage pipeline defects. More specifically, after the host computer reads the video data frame by frame, it performs target instance segmentation detection and continuous tracking through the YOLO12 track interface.
[0115] In practical applications, the results after model processing are stored separately in the form of files. For example, the storage name of the file can be pre-set, such as storage Figure 6 The image shown is in the folder of the original defect image. Figure 7 The images shown are stored in a folder with labeled defect images. Of course, the complete video also needs to be stored, such as Figure 8 As shown, the video after adding defect annotation information can be stored separately. At this time, a defect tracking information record table is also set separately for the separately stored video, which includes: frame number, tracking number (corresponding to the tracking number in the recognition result), defect type (corresponding to the category of the defect in the recognition result), confidence, and distance (corresponding to the recognition distance in the image frame information).
[0116] The information record table may be obtained by summarizing the recognition results of each frame mentioned above.
[0117] S300 , extracting a valid frame image to which each pipeline defect belongs and a motion trajectory of the pipeline defect based on distance information in all frame image attribute information and the pipeline defect recognition result.
[0118] Specifically, the S300 includes:
[0119] S301 , automatically identifying attribute information in each frame of the image. More specifically, using OCR recognition technology to obtain attribute information in each frame of the pipeline video, parsing the attribute information based on regular expressions to obtain distance information.
[0120] In order to achieve automatic recognition of specific values (such as distance information) during pipeline inspection, this embodiment integrates PaddleOCR text recognition technology to perform optical character recognition (OCR) on a specified area of the current video frame image (such as the area approximately 10% of the bottom of the image) to extract the distance value information presented by the image attribute information of each frame in the pipeline video in real time.
[0121] The text data after OCR recognition is parsed twice using regular expressions to extract numerical information containing clear distance unit identification (such as meters, centimeters, kilometers, m, cm, km, etc.). After successful extraction, the host computer judges the validity and stability of the numerical value: that is, by comparing the distance value recognized by the current frame with the distance value recorded in the previous frame, the corresponding change threshold is set (such as a change within 50%) to ensure data continuity and accuracy. Only when this threshold condition is met is the current distance value updated and stored, and the storage operation of the frame image and corresponding tracking data is executed. Frame data that does not pass the threshold judgment or fails to be recognized is automatically ignored to avoid the accumulation and interference of erroneous information.
[0122] S302 : Extracting the valid frame image to which each pipeline defect belongs and the motion trajectory of the pipeline defect based on the recognition result of each frame image and the distance information in the attribute information.
[0123] More specifically, the step 302 includes:
[0124] S302-1. For each defect, obtain the image set T where the defect appears based on the category and tracking number information of the pipeline defect in the recognition result. k ,
[0125] T k ={(F i , ID i , D i , C i , P i )|i=1,2,...,N}
[0126] Among them, F i is the frame number of the i-th frame image, ID i is the defect category and tracking number information corresponding to the i-th frame image, d i is the recognition distance in the distance information of the i-th frame image and satisfies the ascending order of distance, that is, D1≤D2≤…≤D N , C iis the confidence level of the model recognition corresponding to the i-th frame image, 0≤C i ≤1, P i is the u-th frame image. The image set T where the defect appears k Contains N frames of images.
[0127] S302-2. Image set T based on defect occurrence k The distance information in the image is used to obtain the motion trajectory of the defect.
[0128] D total =D N -D1
[0129] Among them, D total is the total movement distance of the defect, D N is the identification distance in the distance information of the image where the defect finally appears, and D1 is the identification distance in the distance information of the image where the defect initially appears.
[0130] S302-3. Extract valid frame images from the defective image set according to the following formula:
[0131]
[0132] in, is a valid frame image set,
[0133] D min is the effective minimum distance, D min =D1+0.7×D total ,
[0134] D max is the effective maximum distance, D max =D1+0.9×D total .
[0135] S400: Comprehensively analyze the motion trajectory of the pipeline defect corresponding to each valid frame image of the pipeline defect, and the category, confidence level, and mask area of the pipeline defect in the recognition result corresponding to each valid frame image, and select a key frame image set of the defect for pipeline defect recognition and post-processing.
[0136] Specifically, the S400 includes:
[0137] S401. In the valid frame image set, a fixed window length W is used to form (M-W+1) smoothing windows, where M is the number of valid frame images in the valid frame image set. The average confidence of each sliding window is calculated using the following formula:
[0138]
[0139] Among them, MCj is the average confidence of the j-th sliding window, C i is the confidence level of the model recognition corresponding to the i-th frame image.
[0140] S402: From all smooth windows, select the window with the largest average confidence as the optimal window.
[0141]
[0142] Among them, j * is the j value of the optimal window, Indicates the parameter value that maximizes the following expression within the range of 1≤j≤M-W+1. Represents the set of images within the optimal window.
[0143] S403: Select the image with the highest confidence level from the image set of the optimal window as the key frame image of the defect.
[0144]
[0145] Among them, F best Indicates the frame number of the key frame image, Indicates The parameter value that makes the following expression obtain the maximum value within the range, C i is the confidence level of the model recognition corresponding to the i-th frame image.
[0146] The final selected key frame image can be fully expressed as:
[0147] F best , ID best , D best , C best , P best
[0148] Among them, F best Indicates the frame number of the key frame image, ID best Indicates the defect category and tracking number (there are one or more defects) corresponding to the key frame image, D best Represents the recognition distance in the key frame image distance information, C best Indicates the confidence of the key frame image (this value is the highest in the window), P best Represents a keyframe image.
[0149] S404: Construct a key frame image set based on the key frame image of each defect.
[0150] Steps S300 to S400 are used to automatically select key video frames based on the target tracking results. Specifically, they involve a comprehensive analysis of the motion trajectory, confidence level, and distance information of target defects in urban drainage pipe inspection videos, thereby accurately and automatically selecting the most representative image frames to improve the efficiency and accuracy of defect target identification and assessment.
[0151] It should be noted that you can first define a series of file paths, including the original tracking information text file (track_info.txt), the folder where the detected frame images are located (detected_frames), and the output folder after the keyframes are selected (selected_frames). Make sure the corresponding folder structure has been established and has the necessary access permissions.
[0152] In the specific implementation process, the previously saved target tracking log file is first read. Each line of log information is parsed using regular expressions to extract the frame number (framenumber), tracking number (track_id), target category (class_name), model recognition confidence (confidence), and real-time recognition distance information (distance) corresponding to each frame image. The parsed data is categorized and stored according to the unique identifier of the tracking target, namely the tracking number and category, so that all relevant frames of information for each tracking target can be processed uniformly.
[0153] Each tracked defect object (identified by its tracking number and category as a unique combination) is then sorted by frame number to determine the initial and final identification distances of the target within the video clip, thereby calculating the target's total motion distance. Based on this motion distance, a valid range for keyframe selection is further defined, specifically set between 70% and 90% of the target's total motion distance. Video frames within this range typically represent stable states during the target's motion, making them suitable for subsequent evaluation and analysis tasks.
[0154] Within the selected effective distance range, a sliding window mechanism (e.g., a window length of 5 frames) is used to traverse all valid frames. The average confidence of the model prediction within each window is calculated, and the window position with the highest average confidence during the sliding process is recorded. Within this optimal window, the image frame with the highest single-frame confidence is further selected as the representative frame to ensure that the selected image is in the optimal state for target recognition accuracy.
[0155] After completing the keyframe selection process, the finalized keyframes are copied and saved to a predefined output folder. A text log file (selected_frames.txt) containing complete tracking and identification information is also generated. Before copying the frame files, an additional file existence check step can be added to avoid anomalies caused by missing files. Detailed debug output information is also included for each operation step to monitor the status and results of the entire process in real time.
[0156] This embodiment combines target tracking technology, identification distance information and model confidence analysis, sets an effective distance interval and uses a sliding window mechanism to select the most representative and most accurate video frames in the motion trajectory of each defect target, thereby significantly improving the efficiency, accuracy and automation level of defect identification and subsequent analysis.
[0157] The pipeline defect recognition and post-processing of this embodiment can be the recognition and tracking of all key frames, for example, through post-processing through automatic recognition and tracking models. The method of this embodiment greatly improves the intelligent detection effect of drainage pipeline defects and can track each defect. Even if N defects of a certain type are detected in a piece of image data, the type of defect, tracking number, frame number of the key frame in which it is located, distance information corresponding to the key frame, and confidence level can be directly given (generally one second contains 24-30 frames, and detection and recognition will be repeated. This problem can be avoided through tracking).
[0158] Example 2
[0159] Based on the first embodiment, this embodiment adds a step of training and verifying the pipeline defect recognition model before S200.
[0160] See also Figure 3 , the method further comprises:
[0161] The pipeline defect training dataset is used to train and verify the pipeline defect recognition model to obtain the trained pipeline defect recognition model.
[0162] That is, before performing the video frame analysis in the pipeline video, the pipeline defect recognition model is first trained and constructed, that is, the pipeline defect recognition model is trained and verified using a pre-given pipeline defect training dataset to obtain a trained pipeline defect recognition model.
[0163] Next, the trained pipeline defect training model is used to analyze each video frame in the pipeline video to remove duplicates, as described in the steps of Example 1 above. That is, the recognition results after using the model to analyze each frame in S200 can include all image frames containing defects, all image frames with defect annotations, and a record table of all defect tracking information (the record table records the frame number, tracking sequence, defect type, confidence level, distance, etc.). After processing in S300, key frame information for each defect is finally obtained (this key frame information may include: the original defect image to which the key frame belongs, the defect image with annotations, the defect information record table, etc.) for use in other inspection and evaluation reports or other tracking analysis processes.
[0164] Each image in the above pipeline defect training dataset is pre-labeled with the defect category and defect mask area; and for images of single defect or multiple defect categories, an image enhancement algorithm is used to process them to obtain training data that meets the preset image quality.
[0165] The pipeline defect training dataset is obtained by extracting, filtering, and labeling pipeline videos collected by the pipeline robot for training. To better illustrate the process of obtaining the pipeline defect training dataset, steps 1 to 3 are added to the following description.
[0166] Step 1: Extract frames from the pipeline videos collected by the pipeline robot for training. Specifically, you can use a command prompt (such as PowerShell) to call a multimedia processing program (such as FFmpeg) to batch extract frames from the pipeline videos used for training. Specify the source folder for storing video files and the output folder for extracted images. Set the output image format (such as png) and the frame extraction frequency (fps). Recursively search all video files (.mp4 and .avi formats) in the source folder. For each video, create a subfolder named after the video file name to store the corresponding extracted images. Use the FFmpeg tool to batch extract video frames at the specified frequency (fps) and save them to the corresponding subfolder.
[0167] When extracting frames, you can also use the adaptive frame extraction method. The formula for the adaptive frame extraction rate is as follows:
[0168] adaptiveFps=min(5,max(1,round(2·wres·wquality·wpipe)))
[0169] When wres = 1920 × 1080 / actual image resolution
[0170] wquality = average brisque / 50
[0171] wpipe=1 / (pipe section coverage + 0.1)
[0172] Range constraint: adaptiveFPS = [1, middle, 5]
[0173] For the above-mentioned adaptive frame extraction, an example is as follows:
[0174] If wres=0.8, wquality=1.2, wpipe=1.5: adaptiveFps=min(5,max(1,round(2×0.8×1.2×1.5)))=min(5,max(1,round(2.88)))=min(5,3)=3.
[0175] Step 2: From the extracted pipeline image data, select images with defects. These images can contain a single defect or multiple defect types. First, check whether the type and scope of the defect in the image are clearly visible. Images that meet the requirements will directly enter the next stage of data annotation and model training.
[0176] For images with unclear defects and substandard quality, image preprocessing techniques are used to improve image quality, ensuring that the defect type and boundary range are more clearly defined. More specifically, image enhancement algorithms (such as computer vision-based image enhancement models) are used to improve image clarity and detail, ensuring accurate identification and analysis of defect information. For example, Deep Neural Network for Image Denoising (DnCNN) and Super Resolution Convolutional Neural Network (SRCNN) can be used to improve image quality.
[0177] Image enhancement algorithms are used to obtain training data that meets the preset image quality, including: using DnCNN and SRCNN to process images of single defects or multiple defect categories.
[0178] DnCNN includes the first input layer, the middle layer and the first output layer. The first input layer extracts features from the input image using the following formula:
[0179]
[0180] Among them, X is the input image, and its dimension is
[0181] Conv2d is a two-dimensional convolution operation, k1=5 means its convolution kernel is 5×5, d=2 means its stride is 2, g=3 means the number of groups in its grouped convolution is 3, padding=same means its padding method is the same, (3→32) means its convolution operation converts 3 channels into 32 channels,
[0182] tanh is the hyperbolic tangent function, F1(X) is the output of the first input layer,
[0183] The middle layer consists of 10 alternating convolutional layers.
[0184]
[0185] in, is the output of each layer in the middle layer, X m is the input of each layer in the middle layer,
[0186] Represents a two-dimensional convolution operation, the convolution kernel k2 is 3×3, the padding value p1 is 1, and the 32 channels are converted to 64 channels.
[0187] Represents a two-dimensional convolution operation, the convolution kernel k3 is 3×3, the padding value p2 is 1, and the 64 channels are converted to 32 channels.
[0188] BatchNorm means batch normalization, LeakyReLU is the activation function,
[0189] F1(X) is the output of the first input layer, F m (X) is the output of the middle layer;
[0190] The first output layer is,
[0191]
[0192] Among them, F m (X) is the output of the middle layer, F DnCNN (X) is the output of the first output layer,
[0193] Represents a two-dimensional convolution operation, with the convolution kernel k4 being 5×5 and the padding value p3 being 2, converting 32 channels into 3 channels.
[0194] The overall definition of DnCNN is: Y = Xf DnCNN (X).
[0195] SRCNN takes the output of DnCNN as input, and the formula is as follows:
[0196]
[0197] Among them, f SRCNN (Y) is the output of SRCNN, Y is the input of SRCNN,
[0198] Represents a two-dimensional convolution operation, the convolution kernel k5 is 9×9, the padding value p4 is 4, and the 3 channels are converted to 64 channels.
[0199] Represents a two-dimensional convolution operation, the convolution kernel k6 is 1×1, the padding value p5 is 0, and the 64 channels are converted to 64 channels.
[0200] Represents a two-dimensional convolution operation, the convolution kernel k7 is 5×5, the padding value p6 is 2, and the 64 channels are converted to 3 channels.
[0201] ReLU stands for Rectified Linear Function.
[0202] The overall definition of the image enhancement algorithm is:
[0203] Z=F SRCNN (XF DnCNN (X)
[0204] Where Z is the output image after the image enhancement algorithm. This formula expresses the end-to-end structure of the image enhancement algorithm, which first uses the DnCNN module to reduce noise (or remove redundant information), and then uses the SRCNN module to enhance details, thereby outputting a high-quality image Z.
[0205] See also Figures 4-5 After adopting the image enhancement algorithm, the image clarity is improved and the details are more obvious, which facilitates the accurate identification and analysis of subsequent defect information.
[0206] It should be noted that other AI models can also be used to improve image quality, such as the Denoising Diffusion Probabilistic Model (DDPM).
[0207] Step 3: Label images with defects.
[0208] Specifically, the selected pipeline defect images can be uploaded in batches to the image data annotation module, and the outlines of each defect type can be marked according to the defect definition in the "Technical Regulations for Inspection and Assessment of Urban Drainage Pipelines", and the scope of the defects can be clearly marked. This step can be completed by professionally trained annotation personnel.
[0209] This embodiment can automatically identify all defect types. For example, the defect types may include the following:
[0210] 0: rupture; 1: deformation; 2: corrosion; 3: misalignment; 4: undulation; 5: disconnection;
[0211] 6: Interface material falling off; 7: Concealed connection of branch pipes; 8: Penetration of foreign objects; 9: Leakage; 10: Deposition; 11: Scaling; 12: Obstacles; 13: Remaining walls and dam roots; 14: Tree roots;
[0212] 15: scum.
[0213] Check whether the defects marked by outlines in the marked images are selected according to the correct defect type and whether the defect range is completely within the selected range. Unqualified samples are returned to the defect marking program for re-editing.
[0214] Export qualified annotated data as images, labels, and annotation files (txt format) and store them in the pipeline defect sample database. The database content will not be modified.
[0215] For example, the sample database can be stored as follows:
[0216] #Sample database path
[0217] Sample library path: path / to / dataset
[0218] Training set:path / to / dataset / train
[0219] Validation set: path / to / dataset / validation
[0220] Test set: path / to / dataset / test.
[0221] Model training is based on the multi-feature fusion deep learning model YOLO12, which supports a variety of computer vision tasks, including image segmentation. It achieves fast real-time inference while maintaining high accuracy, and is compatible with various hardware environments and cross-platform deployment.
[0222] During training, import the training file (the pipeline defect training dataset) from the sample library into the model training module. Divide the dataset into 80% for training, 10% for validation, and the remaining 10% for testing, ensuring that the images and label files correspond. Install relevant dependencies, such as CUDA, cuDNN, PyTorch, and the Ultralytics toolkit. Next, configure the dataset path and defect type information in the dataset.yaml file, and set parameters such as the number of defect types and input image size in the yolo11-seg.yaml file.
[0223] Use the YOLO12 algorithm model to load pre-trained YOLO12 weights, such as yolo12-seg.pt, and set the training parameters, including hyperparameters such as mixed-precision training, optimizer, and learning rate. To improve the model's generalization capabilities, enable data augmentation. Training logs and model weights are saved in a designated directory for easy review and analysis.
[0224] Evaluate the model on the validation set, using key metrics including segmentation accuracy (mask mAP50), precision, and recall. Adjust parameter settings and optimize data augmentation strategies based on the evaluation results to improve the performance of the detection model.
[0225] After training is complete, the detection model will be automatically exported to ONNX format, such as yolo12-seg-pipeline-defects.onnx.
[0226] This embodiment combines the convolutional neural network combination model of DnCNN+SRCNN to improve the quality of pipeline image data, reduce negative images such as image noise, blur, and high contrast caused by the darkness and humidity inside the pipeline and the shooting light source, and greatly increase the utilization rate of pipeline image data.
[0227] Example 3
[0228] This embodiment provides a device for intelligent defect recognition and deduplication in drainage pipe inspection videos, comprising:
[0229] An image acquisition module is used to obtain each frame of the video of the urban drainage pipeline collected by the pipeline robot and its frame number. Each frame of the video is marked with the attribute information of the pipeline robot at the time of collection;
[0230] An analysis and recognition module is used to analyze each frame of the image using a pre-trained pipeline defect recognition model to obtain a recognition result for each frame of the image. The recognition result includes: the category of the pipeline defect in the current frame, the confidence level of each pipeline defect, the pipeline defect mask area, and tracking number information for tracking the pipeline defect in the current frame;
[0231] An effective frame extraction module is used to extract the effective frame image of each pipeline defect and the motion trajectory of the pipeline defect based on the distance information in the attribute information of all frame images and the recognition result of the pipeline defect;
[0232] The screening module is used to comprehensively analyze the motion trajectory of the pipeline defect corresponding to each valid frame image of the pipeline defect, as well as the category, confidence level, and pipeline defect mask area of the recognition result corresponding to each valid frame image, and screen out the key frame image set of the defect for pipeline defect identification and tracking analysis.
[0233] The device of this embodiment can be used to screen and track key frames in pipeline video frames to achieve subsequent further analysis based on pipeline defect video frames.
[0234] This embodiment further provides an electronic device, including: a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program in the memory, specifically performing the steps of the method of the above-mentioned embodiment 1 or embodiment 2.
[0235] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0236] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.
[0237] It should be noted that, in the claims, any reference signs placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims enumerating several means, several of these means may be embodied by one and the same hardware. The use of the words first, second, third etc. is for convenience only and does not indicate any order. These words may be understood as part of the component name.
[0238] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0239] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments after learning the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0240] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the claims and their equivalents, the present invention shall also include such modifications and variations.
Claims
1. A method for intelligent defect identification and deduplication in drainage pipe inspection videos, characterized in that: include: S100: Obtain each frame of an image and a frame number in a video of a town drainage pipeline captured by a pipeline robot, wherein each frame of the image is annotated with attribute information collected by the pipeline robot; S200, analyzing each frame of the image using a pre-trained pipeline defect recognition model to obtain a recognition result for each frame of the image, the recognition result including: a category of the pipeline defect in the current frame, a confidence level of each pipeline defect, a pipeline defect mask area, and tracking number information for tracking the pipeline defect in the current frame; S300, extracting a valid frame image to which each pipeline defect belongs and a motion trajectory of the pipeline defect based on distance information in attribute information of all frame images and the recognition result of the pipeline defect; S400: Comprehensively analyze the motion trajectory of the pipeline defect corresponding to each valid frame image of the pipeline defect, and the category, confidence level, and mask area of the pipeline defect in the recognition result corresponding to each valid frame image, to select a defect key frame image set for pipeline defect recognition and tracking analysis.
2. The method according to claim 1, characterized in that Before S200, the method further includes: The pipeline defect recognition model is trained and verified using the pipeline defect training dataset to obtain a trained pipeline defect recognition model; Each image in the pipeline defect training dataset is pre-labeled with a defect category and a defect mask area; and for images of a single defect or multiple defect categories, an image enhancement algorithm is used to process the images to obtain training data that meets the preset image quality.
3. The method according to claim 2, characterized in that Image enhancement algorithms are used to obtain training data that meets the preset image quality, including: Use DnCNN and SRCNN to process images with single defect or multiple defect categories; DnCNN includes the first input layer, the middle layer and the first output layer. The first input layer extracts features from the input image using the following formula: Among them, X is the input image, and its dimension is Conv2d is a two-dimensional convolution operation, k1=5 means its convolution kernel is 5×5, d=2 means its stride is 2, g=3 means the number of groups in its grouped convolution is 3, padding=same means its padding method is the same, (3→32) means its convolution operation converts 3 channels into 32 channels, tanh is the hyperbolic tangent function, F1(X) is the output of the first input layer, The middle layer consists of 10 alternating convolutional layers. in, is the output of each layer in the middle layer, X m is the input of each layer in the middle layer, Represents a two-dimensional convolution operation, the convolution kernel k2 is 3×3, the padding value p1 is 1, and the 32 channels are converted to 64 channels. Represents a two-dimensional convolution operation, the convolution kernel k3 is 3×3, the padding value p2 is 1, and the 64 channels are converted to 32 channels. BatchNorm means batch normalization, LeakyReLU is the activation function, F1(X) is the output of the first input layer, F m (X) is the output of the middle layer; The first output layer is, Among them, F m (X) is the output of the middle layer, f DnCNN (X) is the output of the first output layer, Represents a two-dimensional convolution operation, with the convolution kernel k4 being 5×5 and the padding value p3 being 2, converting 32 channels into 3 channels; SRCNN takes the output of DnCNN as input, and the formula is as follows: Among them, f SRCNN (Y) is the output of SRCNN, Y is the input of SRCNN, Represents a two-dimensional convolution operation, the convolution kernel k5 is 9×9, the padding value p4 is 4, and the 3 channels are converted to 64 channels. Represents a two-dimensional convolution operation, the convolution kernel k6 is 1×1, the padding value p5 is 0, and the 64 channels are converted to 64 channels. Represents a two-dimensional convolution operation, the convolution kernel k7 is 5×5, the padding value p6 is 2, and the 64 channels are converted to 3 channels. ReLU stands for Rectified Linear Function.
4. The method according to claim 1, wherein The attribute information of each frame of image includes: time, robot identification, distance information, and pipeline identification marked on each frame of image; The S300 includes: S301, automatically identifying attribute information in each frame of image; S302 : Extracting the valid frame image to which each pipeline defect belongs and the motion trajectory of the pipeline defect based on the recognition result of each frame image and the distance information in the attribute information.
5. The method according to claim 4, characterized in that The 301 includes: The attribute information of each frame of the pipeline video is obtained by using OCR recognition technology, and the attribute information is parsed based on regular expressions to obtain distance information.
6. The method according to claim 4, characterized in that The 302 includes: S302-1. For each defect, obtain the image set T where the defect appears based on the category and tracking number information of the pipeline defect in the recognition result. k , T k ={(F i ,ID i ,D i ,C i ,P i )|i=1,2,...,N} Among them, F i is the frame number of the i-th frame image, ID i is the defect category and tracking number information corresponding to the i-th frame image, D i is the recognition distance in the distance information of the i-th frame image and satisfies the ascending order of distance, that is, D1≤D2≤…≤D N , C i is the confidence level of the model recognition corresponding to the i-th frame image, 0≤C i ≤1, P i is the i-th frame image, S302-2. Image set T based on defect occurrence k The distance information in the image is used to obtain the motion trajectory of the defect. D total =D N -D1 Among them, D total is the total movement distance of the defect, D N is the identification distance in the distance information of the image where the defect finally appears, and D1 is the identification distance in the distance information of the image where the defect initially appears; S302-3. Extract valid frame images from the defective image set according to the following formula: in, is a valid frame image set, D min is the effective minimum distance, D min =D1+0.7×D total , D max is the effective maximum distance, D max =D1+0.9×D total .
7. The method according to claim 4, characterized in that The S400 includes: S401. In the valid frame image set, a fixed window length W is used to form (M-W+1) smoothing windows, where M is the number of valid frame images in the valid frame image set. The average confidence of each sliding window is calculated using the following formula: Among them, MC j is the average confidence of the j-th sliding window, C i is the confidence of the model recognition corresponding to the i-th frame image; S402: From all smooth windows, select the window with the largest average confidence as the optimal window. Among them, j * is the j value of the optimal window, Indicates the parameter value that maximizes the following expression within the range of 1≤j≤M-W+1. represents the image set within the optimal window; S403: Select the image with the highest confidence level from the image set of the optimal window as the key frame image of the defect. Among them, F best Indicates the frame number of the key frame image, Indicates The parameter value that makes the following expression obtain the maximum value within the range, C i is the confidence of the model recognition corresponding to the i-th frame image; S404: Construct a key frame image set based on the key frame image of each defect.
8. The method according to claim 1, characterized in that The pipeline defect recognition model is a pipeline defect recognition model based on the YOLO12 algorithm, including a second input layer, a feature extraction layer, a feature fusion layer, an upsampling layer, and a second output layer; The second input layer is used to receive the input drainage pipe image. The feature layer includes convolution layer, C3k2 layer, A2C2f layer, which are used for preliminary feature extraction and attention processing. The fusion layer is used to splice the feature maps output by different levels of the feature layer. The upsampling layer is used to increase the resolution of the feature map output by the fusion layer. The second output layer is used to output segmentation and detection results.
9. A device for intelligent defect recognition and deduplication in drainage pipe inspection videos, characterized in that: include: An image acquisition module is used to obtain each frame of the video of the urban drainage pipeline collected by the pipeline robot and its frame number. Each frame of the video is marked with the attribute information of the pipeline robot at the time of collection; An analysis and recognition module is used to analyze each frame of the image using a pre-trained pipeline defect recognition model to obtain a recognition result for each frame of the image. The recognition result includes: the category of the pipeline defect in the current frame, the confidence level of each pipeline defect, the pipeline defect mask area, and tracking number information for tracking the pipeline defect in the current frame; An effective frame extraction module is used to extract the effective frame image of each pipeline defect and the motion trajectory of the pipeline defect based on the distance information in the attribute information of all frame images and the recognition result of the pipeline defect; The screening module is used to comprehensively analyze the motion trajectory of the pipeline defect corresponding to each valid frame image of the pipeline defect, as well as the category, confidence level, and pipeline defect mask area of the recognition result corresponding to each valid frame image, and screen out the defect key frame image set for pipeline defect identification and tracking analysis.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program in the memory, specifically performing the steps of the method for intelligent defect identification and deduplication in drainage pipe inspection videos as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Drainage pipeline defect detection method and system based on deep learning
CN113469177A
Pipeline defect detection method and device, electronic equipment and storage medium
CN116416208A
Drainage pipeline defect detection method based on convolutional neural network
CN117952957A
Drainage pipeline defect detection processing method based on YOLO model
CN118230063A
Pipeline defect detection method and system and electronic equipment
CN119600018A
Cited By
Drainage pipeline video defect processing method based on deep learning and sequence interpretation
CN122156836A