A method for training a pipeline defect detection model, a detection method, a device, and a storage medium
By introducing a relationship-aware self-attention mechanism and an improved RSwinTransformer backbone network, the loss function is optimized, and the existing model lacks perception ability in pipeline defect detection is solved, and more efficient detection of complex structure defects is achieved, which improves detection accuracy and robustness.
Patent Information
- Application Number
- CN202510630226.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing image object detection model lacks the ability to perceive complex structure defects in pipeline defect detection, especially in irregular shapes, blurred edges and background noise, resulting in poor detection effect.
A pipeline defect detection model training method is adopted, and the relationship-aware self-attention mechanism (RASA) is introduced to build a backbone network based on the improved RSwinTransformer. Through multi-scale adjustment and collaborative label allocation mechanism of auxiliary heads, the modeling ability of spatial geometric relationships between targets is enhanced, and the loss function design is optimized to improve the large-object detection performance.
It improves the perception and recognition capabilities of the pipeline defect detection model in complex structural scenarios, improves the accuracy, robustness and real-timeness of detection, and is suitable for detection scenarios such as underground pipeline networks with complex target scale distribution and changes.
Smart Images

Figure CN120147623B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image target detection, and in particular to a method for training a pipeline defect detection model, a detection method, a device, and a storage medium. Background Art
[0002] Pipelines such as water supply and drainage pipelines may have defects due to production and assembly errors or aging during use. Typical defects include deformation, offset, and rupture. Traditional pipeline defect detection relies on manual observation. However, for pipelines located in inaccessible or dangerous locations such as underground pipe networks, it is difficult to detect pipeline defects. Even after using a robot to take videos of the pipelines and then having a human view the videos for analysis, there are still the disadvantages of low efficiency and being prone to errors.
[0003] Some related technologies use an image target detection model, with the positions where there are defects on the pipeline as the detection targets, and perform target detection on the image data captured by the robot, thereby achieving automatic detection of pipeline defects. However, the current image target detection models have problems such as insufficient ability to perceive complex structure defects and low detection efficiency in the case of images with irregular shapes, blurred edges, and background noise, which are exactly the characteristics of the pipeline images where the defects to be detected are located. Therefore, the current image target detection technology has poor application effects in pipeline defect detection. Summary of the Invention
[0004] Aiming at the technical problems that the current image target detection technology has insufficient ability to perceive complex structure defects, low detection efficiency in the case of images with irregular shapes, blurred edges, and background noise, and poor application effects in pipeline defect detection, the purpose of the present invention is to provide a method for training a pipeline defect detection model, a detection method, a device, and a storage medium.
[0005] On the one hand, an embodiment of the present invention includes a method for training a pipeline defect detection model, and the method for training the pipeline defect detection model includes the following steps:
[0006] Obtain a first detection model; the first detection model includes a backbone network, a Transformer encoding block, a Transformer decoding block, a multi-scale adjuster, and at least one auxiliary head, and the backbone network includes multiple stages;
[0007] Obtain training images and ground truth boxes; the ground truth boxes represent the true positions of the defect parts included in the training images;
[0008] Input the training images into the backbone network, and each of the stages in the backbone network processes the training images in sequence through a relationship-aware self-attention mechanism to obtain image feature data;
[0009] Input the image feature data into the Transformer encoding block for encoding to obtain encoded feature data;
[0010] Input the encoded feature data into the multi-scale adjuster for multi-scale adjustment to obtain a feature pyramid;
[0011] Traverse all the auxiliary heads;
[0012] For any traversed auxiliary head, the auxiliary head processes the feature pyramid and obtains multiple positive sample positions according to the ground truth box, performs position encoding on each positive sample position to obtain customized query information, inputs the customized query information, learnable query information, and the encoded feature data into the Transformer decoding block for decoding, obtains the multi-scale prediction results of the Transformer decoding block, determines the multi-scale aggregation loss function according to the multi-scale prediction results, and trains the first detection model according to the multi-scale aggregation loss function.
[0013] Furthermore, the backbone network further includes an image partitioning module;
[0014] The image partitioning module is used to partition the training image into multiple image patches;
[0015] The first stage includes a linear embedding module and an RSwinTransformer module, and the other stages except the first stage include an image patch fusion embedding module and an RSwinTransformer module;
[0016] The input data of the first stage is each image patch; the input data of any other stage except the first stage is the output data of the previous stage;
[0017] The linear embedding module is used to perform dimensionality conversion on the input data;
[0018] The image patch fusion embedding module is used to perform spatial dimensionality reduction on the input data.
[0019] Furthermore, the RSwinTransformer module includes a relation-aware self-attention unit and a multi-layer perceptron unit;
[0020] The relation-aware self-attention unit is used to process the input data of the RSwinTransformer module through the relation-aware self-attention mechanism, and then fuse it with the input data of the RSwinTransformer module to obtain the processing result of the relation-aware self-attention unit;
[0021] The multi-layer perceptron unit is used to process the processing result of the relation-aware self-attention unit, and then fuse it with the processing result of the relation-aware self-attention unit to obtain the output data of the RSwinTransformer module.
[0022] Further, the relation-aware self-attention unit includes a receiving subunit, a first matrix multiplication subunit, a scaling subunit, a geometric weight subunit, a weighted Softmax subunit, a second matrix multiplication subunit, and a multi-layer perceptron subunit;
[0023] The receiving subunit is used to receive and obtain multiple data units of the input data of the RSwinTransformer module, and disassemble the data units into content features and position features;
[0024] The first matrix multiplication subunit is used to perform arithmetic processing according to the formula where, is the content feature corresponding to the th data unit, is the content feature corresponding to the th data unit, and are the weight matrices of the first matrix multiplication subunit, represents the matrix transpose operation;
[0025] The scaling subunit is used to perform arithmetic processing according to the formula to obtain the content weight ; is the scaling coefficient;
[0026] The geometric weight subunit is used to perform arithmetic processing according to the formula
[0027]
[0028]
[0029] to obtain the geometric weight ; where, is the position feature corresponding to the th data unit, is the position feature corresponding to the th data unit, and are the spatial coordinates corresponding to the th data unit, is the width corresponding to the th data unit, is the height corresponding to the th data unit, and is the spatial coordinate corresponding to the th data unit, is the width corresponding to the th data unit, is the height corresponding to the th data unit, represents the GIOU operation, is the weight matrix of the geometric weight sub-unit, represents encoding;
[0030] The weighted Softmax sub-unit is used to perform operation processing according to the formula
[0031]
[0032] to obtain the total attention weight ;
[0033] The second matrix multiplication sub-unit is used to perform operation processing according to the formula
[0034]
[0035] ; where is the number of data units corresponding to the input data of the RSwinTransformer module, is the weight matrix of the second matrix multiplication sub-unit;
[0036] The multi-layer perceptron sub-unit is used to perform operation processing according to the formula
[0037]
[0038] to obtain the processing result of the relation-aware self-attention unit ; where represents concatenation processing, represents multi-layer perceptron processing.
[0039] Furthermore, determining the multi-scale aggregation loss function according to the multi-scale prediction result includes:
[0040] Calculating the multi-scale aggregation loss function according to the formula
[0041]
[0042] ; where is the th layer prediction result in the multi-scale prediction result, is the scale weighting coefficient corresponding to the th layer, is the is an adaptive threshold focal loss function, is the total number of layers of the prediction results included in the multi-scale prediction results.
[0043] Furthermore,
[0044] is the average target size corresponding to the th layer in the feature pyramid.
[0045] On the other hand, an embodiment of the present invention further includes a pipeline defect detection method. The pipeline defect detection model training method includes the following steps:
[0046] Obtain a second detection model; the second detection model is obtained based on the first detection model in the pipeline defect detection model training method;
[0047] Obtain an image to be detected; the image to be detected includes pipeline image content;
[0048] Input the image to be detected into the second detection model for processing;
[0049] Determine the pipeline defect position in the image to be detected according to the processing result of the second detection model.
[0050] Furthermore, the obtaining of the second detection model includes:
[0051] Obtain the first detection model;
[0052] Discard all the auxiliary heads in the first detection model to obtain the second detection model.
[0053] On the other hand, an embodiment of the present invention further includes a computer device, including a memory and a processor. The memory is used to store at least one program, and the processor is used to load at least one program to execute the pipeline defect detection model training method in the embodiment.
[0054] On the other hand, an embodiment of the present invention further includes a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to execute the pipeline defect detection model training method in the embodiment when executed by the processor.
[0055] The beneficial effects of the present invention are as follows: The pipeline defect detection model training method in the embodiment introduces a relationship-aware self-attention mechanism (RASA) to construct a backbone network based on the improved RSwinTransformer, enhancing the first detection model's ability to model the spatial geometric relationships between targets, improving the first detection model's perception and recognition capabilities in complex structure scenarios such as underground pipe networks, having advantages in terms of accuracy, robustness, and real-time performance, and being applicable to actual detection scenarios with complex target scale distributions and changes, especially for pipelines such as underground pipe networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a schematic diagram of the steps of the pipeline defect detection model training method in the embodiment;
[0057] Figure 2 It is a schematic diagram of the structure of the first detection model in the embodiment;
[0058] Figure 3 It is a schematic diagram of the structure of the backbone network in the embodiment;
[0059] Figure 4 It is a schematic diagram of the structure of RswinTransformer in the embodiment;
[0060] Figure 5 It is a schematic diagram of the structure of the relationship-aware self-attention unit in the embodiment;
[0061] Figure 6 It is a schematic diagram of the structure of the second detection model in the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0062] The current object detection models have the following disadvantages:
[0063] Insufficient spatial perception ability: Since traditional ViT or CNN is still used as the backbone network, its attention mechanism has limited ability to model the complex spatial relationships between targets. Especially in the underground pipe network scenario, in the face of problems such as non-regular deformation, local corrosion, complex lighting, and background interference, the current object detection models are difficult to effectively capture the geometric and semantic dependencies between targets;
[0064] Weak scale adaptability and difficulty in optimizing large object detection: Using loss functions in the form of Focal Loss or its variants can enhance the attention to difficult samples, but the supervision signal for large-size targets is not strong enough. Especially in the multi-scale feature fusion framework, the influence of different scales on the detection performance is not considered, resulting in the risk of missed detection and inaccurate positioning in the underground pipeline defect detection task dominated by large objects;
[0065] Insufficient utilization of auxiliary supervision: Even if some object detection models introduce the auxiliary supervision mechanism of One-to-Many auxiliary heads, the matching degree between their label assignment methods and feature extraction capabilities is insufficient, and the potential of intermediate feature maps is not fully explored, restricting the learning efficiency and generalization ability of the decoder.
[0066] In view of the technical problems existing in the current object detection models, a pipeline defect detection model training method is provided in this embodiment. As Figure 1 shown, the pipeline defect detection model training method includes the following steps:
[0067] P1. Obtain the first detection model;
[0068] P2. Obtain training images and ground truth boxes;
[0069] P3. Input the training images into the backbone network, and each stage in the backbone network processes the training images in turn through the relation-aware self-attention mechanism to obtain image feature data;
[0070] P4. Input the image feature data into the Transformer encoding block for encoding to obtain encoded feature data;
[0071] P5. Input the encoded feature data into the multi-scale adjuster for multi-scale adjustment to obtain a feature pyramid;
[0072] P6. Traverse all auxiliary heads:
[0073] For any traversed auxiliary head, the auxiliary head processes the feature pyramid and obtains multiple positive sample positions according to the ground truth boxes, performs position encoding on each positive sample position to obtain customized query information, inputs the customized query information, learnable query information, and encoded feature data into the Transformer decoding block for decoding, obtains the multi-scale prediction results of the Transformer decoding block, determines the multi-scale aggregation loss function according to the multi-scale prediction results, and trains the first detection model according to the multi-scale aggregation loss function.
[0074] In step P1, the established first detection model is as Figure 2 shown. Referring to Figure 2 , the first detection model is a Co-DETR (DETRs with Collaborative Hybrid Assignments Training) model, including components such as a backbone network, a Transformer encoding block, a Transformer decoding block, a multi-scale adjuster, and multiple auxiliary heads.
[0075] In this embodiment, the structure of the backbone network in the first detection model is as Figure 3As shown in the reference Figure 3 , the backbone network includes an image partitioning module and multiple stages, which are sorted and connected in sequence according to the processing order of the data. Among them, the first stage includes a linear embedding module and an RSwin Transformer module, and the other stages except the first stage include an image patch fusion embedding module and an RSwin Transformer module. That is, each stage includes an RSwin Transformer module, so that each stage can process data through a relation-aware self-attention mechanism.
[0076] In this embodiment, the structures of the RSwin Transformer modules included in different stages are the same. Taking one of the RSwin Transformer modules as an example, its structure is as Figure 4 shown in the reference Figure 4 . The RSwin Transformer module includes a relation-aware self-attention unit, a multi-layer perceptron unit, and processing parts such as layer normalization. Among them, the relation-aware self-attention unit is used to process the input data of the RSwin Transformer module through the relation-aware self-attention mechanism, and then fuse it with the input data of the RSwin Transformer module to obtain the processing result of the relation-aware self-attention unit; the multi-layer perceptron unit is used to process the processing result of the relation-aware self-attention unit, and then fuse it with the processing result of the relation-aware self-attention unit to obtain the output data of the RSwin Transformer module.
[0077] When performing step P2, high-definition video acquisition can be carried out on the inside of the underground pipe network through a camera device mounted on the pipeline robot, so as to obtain an image of a pipeline with typical defects such as deformation, misalignment, and rupture as the training image. Then, according to relevant defect detection technical standards (such as the Technical Specification for Inspection and Assessment of Urban Drainage Pipelines CJJ181-2012), a labeling tool (such as Labelimg) is used to manually select the true positions of pipeline defects in the training image, so as to obtain the ground truth box corresponding to the training image.
[0078] In step P2, multiple training images and their respective ground truth boxes can be obtained. Among them, each training image can have one or more ground truth boxes. In this embodiment, one training image and its corresponding one ground truth box are taken as an example for illustration.
[0079] In step P3, the training image is input into the backbone network, so that each stage in the backbone network processes the training image in sequence through the relation-aware self-attention mechanism to obtain image feature data.
[0080] Specifically, referring to Figure 3 , after the training image is input into the backbone network, the image partitioning module in the backbone network receives the training image and divides the two-dimensional training image into multiple image patches of a fixed size. For example, the size of the training image can be H×W×3 (specifically 224×224×3). After being divided by the image partitioning module, if the size of each image patch is set to 4×4, then (H / 4)×(W / 4) image patches can be obtained. The image partitioning module can also flatten each image patch into a vector with a length of 4×4×3 = 48, so that the image patches are output in the form of vectors, that is, the image partitioning module outputs a vector of N×48 dimensions, where N=(H / 4)×(W / 4) is the total number of image patches.
[0081] Referring to Figure 3 , each image patch output by the image partitioning module serves as the input data for the first stage (i.e., Figure 3 stage 1 in ), and is processed by stage 1.
[0082] The linear embedding module in stage 1 can perform dimensionality conversion on the input data. Specifically, the linear embedding module can convert the vector corresponding to each image patch into a specified feature dimension C (C can specifically be 96, 128, 256, etc.) through a fully connected layer, so that the shape of the tensor output by the linear embedding module becomes N×C. The linear embedding module can also perform a reshape operation so that the output data size is (H / 4)×(W / 4)×C, which is convenient for the RSwinTransformer module in stage 1 to process.
[0083] Referring to Figure 4 , after the RSwinTransformer module in stage 1 receives the data output by the linear embedding module, it enters the relation-aware self-attention unit after layer normalization for reception. Each sub-unit in the relation-aware self-attention unit performs corresponding data processing respectively.
[0084] Referring to Figure 5 , in the relation-aware self-attention unit of stage 1, the receiving sub-unit receives the input data of the RSwinTransformer module. Since the image partitioning module before stage 1 has divided the training image into multiple image patches, Figure 5 the input data of the RSwinTransformer module received by the receiving sub-unit in has multiple data units, each data unit corresponds to an image patch respectively, and each data unit can be used as a token and processed by each sub-unit in the relation-aware self-attention unit.
[0085] In this embodiment, for the N received data units (tokens), the receiving subunit disassembles each data unit (token) into corresponding content features and position features respectively. Among them, the content feature corresponding to the th data unit (token) is denoted as , and the position feature corresponding to the th data unit (token) is denoted as . The receiving subunit sends data such as content features and position features to the first matrix multiplication subunit, the geometric weight subunit, the second matrix multiplication subunit, etc.
[0086] Referring to Figure 5 , the first matrix multiplication subunit performs arithmetic processing according to the formula ; among them, is the content feature corresponding to the th data unit, is the content feature corresponding to the th data unit, and are the weight matrices of the first matrix multiplication subunit, and represents matrix transpose operation.
[0087] Referring to Figure 5 , the scaling subunit scales the output by the first matrix multiplication subunit. Specifically, the scaling subunit performs arithmetic processing according to the formula to obtain the content weight . Among them, is the scaling coefficient.
[0088] The combination of the first matrix multiplication subunit and the scaling subunit realizes the calculation of the content weight
[0089]
[0090] between the th data unit and the th data unit among the N data units (tokens). The content weight represents the content similarity between the th data unit and the th data unit. Among them, by dividing by can prevent the dot product value from being too large and causing the subsequent Softmax gradient to disappear.
[0091] Referring to Figure 5 , the geometric weight subunit uses the formula
[0092]
[0093]
[0094] Perform arithmetic processing to obtain geometric weights ; where is the position feature corresponding to the th data unit, is the position feature corresponding to the th data unit, and are the spatial coordinates corresponding to the th data unit (specifically, the coordinates of the th data unit in the training image), is the width corresponding to the th data unit, is the height corresponding to the th data unit, and are the spatial coordinates corresponding to the th data unit (specifically, the coordinates of the th data unit in the training image), is the width corresponding to the th data unit, is the height corresponding to the th data unit, represents the GIOU operation, is the weight matrix of the geometric weight sub-unit, represents encoding.
[0095] The geometric weight sub-unit can calculate the geometric weight between the th data unit and the th data unit. The geometric weight represents the position correlation between the th data unit and the
[0096] Refer to Figure 5 , the weighted Softmax sub-unit further processes the content weight output by the scaling sub-unit and the geometric weight output by the geometric weight sub-unit, and performs arithmetic processing according to the formula
[0097]
[0098] to obtain the total attention weight . The total attention weight fuses the content weight and geometric weights
[0099] Refer to Figure 5 , the second matrix multiplication subunit obtains the total attention weights output by the weighted Softmax subunit and the content features corresponding to the data units output by the receiving subunit , according to the formula
[0100]
[0101] perform arithmetic processing; where is the number of data units corresponding to the input data of the RSwinTransformer module, is the weight matrix of the second matrix multiplication subunit
[0102] Output by the second matrix multiplication subunit , is the th data unit that forms a relationship enhancement representation by integrating the information of other data units
[0103] Refer to Figure 5 , the multi-layer perceptron subunit receives the relationship enhancement representation output by the second matrix multiplication subunit and the content features corresponding to the data units output by the receiving subunit , according to the formula
[0104]
[0105] perform arithmetic processing to obtain the processing result of the relationship-aware self-attention unit ; where represents concatenation processing, represents multi-layer perceptron processing
[0106] In this embodiment, the relationship-aware self-attention unit in the RSwinTransformer module can perform data processing through a relationship-aware self-attention mechanism (Relation-aware Self-Attention, RASA). RASA is an improved self-attention mechanism that can encode geometric priors (such as the relative distance of the target) into the attention weights. By introducing geometric weights, RASA enhances the perception ability of the attention module for local spatial structures and semantic contexts, making the first detection model containing the relationship-aware self-attention unit particularly suitable for structural target detection scenarios, such as defect recognition of types such as underground pipeline cracks and damages
[0107] Refer to Figure 4 , the processing result output by the relationship-aware self-attention unit in stage 1 , after being fused with the input data of the RSwinTransformer module, the obtained data itself is fused with the obtained data after being processed by layer normalization and multi-layer perceptron units in sequence. The obtained data is used as the output data of the RSwinTransformer module in Stage 1, which is also the output data of Stage 1.
[0108] Refer to Figure 3 , the output data of Stage 1 is received by Stage 2 and continues to be processed. The output data of Stage 1 still retains the chunking pattern like the output data of the image partitioning module, that is, the output data of Stage 1 also includes multiple parts, and each part corresponds to a data unit (token) respectively. Therefore, in Stage 2, it is still said that the data to be processed includes multiple data units (tokens), which will not be confused with the data units in Stage 1.
[0109] The image patch fusion module in Stage 2 performs spatial dimensionality reduction on the input data. The spatial dimensionality reduction is similar to the downsampling operation of CNN, which can merge adjacent data units to obtain more abstract and higher semantic level feature representations, while increasing the channel dimension. For example, the image patch fusion module in Stage 2 can take 4 adjacent data units with a size of 2×2 each, concatenate their channels (there are a total of 4×C channels), so as to obtain data of size 4C, and then use a linear transformation to map the data of size 4C to a new dimension (such as 2C or other values), so that the spatial resolution of the data is reduced to half of the original, and the channel dimension is increased. The output tensor shape of the image patch fusion module in Stage 2 becomes (H / 8)×(W / 8)×2C, and can also be divided into multiple data units (tokens) like in Stage 1.
[0110] Refer to Figure 3 , the RswinTransformer module in Stage 2 processes the data output by the image patch fusion module in Stage 2. Its principle and process are the same as those of the RswinTransformer module in Stage 1. The data output by the RswinTransformer module in Stage 2 is received by Stage 3 and processed. The structure of Stage 3 is the same as that of Stage 2, and its data processing process and principle are also the same as those of Stage 2. Similarly, the data output by the RswinTransformer module in Stage 3 is received by Stage 4 and processed.
[0111] Refer to Figure 3 , the output data of the RswinTransformer module in the last stage, that is, Stage 4, is the output data of the backbone network, that is, the image feature data obtained by processing the training image.
[0112] In step P4, the image feature data output by the backbone network is encoded by the Transformer encoding block to obtain encoded feature data. The Transformer encoding block is the encoder in the Transformer structure, which can perform global context modeling on the local features of the image feature data. Through multiple self-attention layers and a feed-forward neural network, it outputs a feature sequence with enhanced global information, that is, the encoded feature data.
[0113] Refer to Figure 2 , in step P5, the encoded feature data output by the Transformer encoding block is input into a multi-scale adjuster for multi-scale adjustment to obtain a feature pyramid. . The feature pyramid includes layers of features.
[0114] Refer to Figure 2 , the first detection model is provided with multiple auxiliary heads such as auxiliary head 1... auxiliary head k. These auxiliary heads can be networks such as Faster-RCNN, ATSS, RetinaNet, FCOS, etc. Specifically, all auxiliary heads can use the same network, or different auxiliary heads can use different networks.
[0115] Refer to Figure 2 , the operation of each auxiliary head is independent, that is, each auxiliary head will have a corresponding one-to-many label assignment, that is, one ground truth box can correspond to multiple positive sample positions.
[0116] In step P6, all auxiliary heads are traversed. The steps performed by one auxiliary head do not affect the steps performed by other auxiliary heads. Taking one of the traversed auxiliary heads (for example, the th auxiliary head) as an example, this auxiliary head performs the following steps:
[0117] P601. Process the feature pyramid and obtain multiple positive sample positions according to the ground truth box;
[0118] P602. Perform position encoding on each positive sample position to obtain customized query information;
[0119] P603. Input the customized query information, learnable query information, and encoded feature data into the Transformer decoding block for decoding to obtain the multi-scale prediction results of the Transformer decoding block;
[0120] P604. Determine the multi-scale aggregation loss function according to the multi-scale prediction results, and train the first detection model according to the multi-scale aggregation loss function.
[0121] In step P601, the th auxiliary head processes the feature pyramid to obtain a prediction result , denotes the defective objects that may exist in the training image predicted by the th auxiliary head. The positive sample positions can be determined according to the ground truth box . Specifically, the label assignment corresponding to the th auxiliary head is . Then, the processing in step P601 can be performed by executing the formula
[0122]
[0123] to obtain the positive sample positions as well as and and other data. Among them, the positive sample position denotes the position of the positive sample (e.g., overlapping or close to the ground truth box ) in the prediction result of the th auxiliary head. The supervision target of the positive sample,
[0124] is the supervision target of the negative sample, where the supervision targets include object classification and regressed offset, etc. Different types of auxiliary heads can have different positive sample position generation methods. For example, if Faster-RCNN is used as the auxiliary head, then the positive proposals output by the auxiliary head can be selected as the positive sample positions; if ATSS is used as the auxiliary head, then the positive anchors output by the auxiliary head can be selected as the positive sample positions.
[0125] In step P602, for the positive sample position obtained by the th auxiliary head, it can be processed by the formula
[0126]
[0127] to obtain the customized query information . Among them, denotes the position encoding, and denotes the linearization.
[0128] In step P603, the customized query information is combined with the learnable query information The encoded feature data output by the Transformer encoding block is input into the Transformer decoding block, which is decoded by the Transformer decoding block to obtain the multi-scale prediction results of the Transformer decoding block.
[0129] The multi-scale prediction results obtained in step P603 are also divided into layers, where the layer prediction result is denoted as .
[0130] In step P604, for the layer prediction result , the Adaptive Threshold Focus Loss (ATFL) can be used to calculate its corresponding loss function , and then a scale weighting coefficient designed according to the large target size is set for this layer . The loss functions of all layers are weighted and summed, that is, according to the formula
[0131]
[0132] the multi-scale aggregation loss function is calculated.
[0133] In this embodiment, the expression of the Adaptive Threshold Focus Loss function used is
[0134]
[0135] where is the adaptive hard sample modulation factor, is the prediction probability (for example, for , can be the corresponding prediction probability, that is, the prediction belongs to the ground truth box).
[0136] In this embodiment, the scale weighting coefficient can be calculated by the formula
[0137]
[0138] where is the average target size corresponding to the layer in the feature pyramid.
[0139] Scale weighting coefficient The principle lies in that in the object detection of multi-scale feature fusion, the receptive fields of feature maps at different layers are different. Generally speaking, the lower layers pay more attention to details, and the higher layers pay more attention to semantics. For detections such as pipeline defect detection, the high-level semantic features are more representative; the scale weighting coefficients calculated in this way increase as the target size in the feature layer increases (that is, closer to the original image space and corresponding to large targets), enabling the first detection model to make full use of the feature information of large targets during training and improving the detection accuracy of large targets.
[0140] In step P604, after calculating the multi-scale aggregation loss function , the AdamW optimizer can be used, with the initial learning rate set to 1e-5, and the learning rate decay is performed in the middle of training, and the batch size is set to 32. The parameters trained when training the first detection model include the weight matrix , , and as well as the learnable query information and so on.
[0141] The above steps P601 - P604 are the training process using one of the auxiliary heads. The training processes of other auxiliary heads are also executed through steps P601 - P604, and the training processes of different auxiliary heads can be independent, training the parameters of the same first detection model.
[0142] The above steps P1 - P6 are the training process using one of the training images. In the case of multiple training images, the multiple training images can be divided into a training set and a validation set. During the training process, the training set and the validation set are loaded simultaneously, and steps P1 - P6 are executed for each training image respectively.
[0143] The first detection model established and trained through steps P1 - P6 has the following characteristics:
[0144] 1. Innovative backbone network structure: The RswinTransformer module is provided in the backbone network. The relationship-aware self-attention mechanism is realized through the relationship-aware self-attention unit. By fusing content attention and the geometric weight matrix, the model's ability to model the spatial structure between targets is enhanced, especially suitable for the structural defect perception task in the underground pipe network scenario;
[0145] Introduce a relationship-aware self-attention mechanism to enhance spatial modeling and target perception capabilities. Different from related structures such as W-MSA window attention, the "relationship-aware self-attention unit" in the RswinTransformer module in this embodiment can introduce geometric information between content and position, making the modeling of the spatial relationship between targets more expressive, and is particularly suitable for defect recognition tasks with non-rigid and complex structures;
[0146] 2. Optimized loss function design: Propose a multi-scale adaptive focal loss function optimized for large target detection, dynamically adjust the loss weight according to the target size of each feature layer, so as to strengthen the training guidance of large-size targets in the high-level semantic feature map and improve the target localization accuracy and recall rate;
[0147] The improved MATFL loss function combines multi-scale feature map information and adaptively adjusts the loss weight according to the target size, so that large targets can obtain stronger supervision signals in the deep semantic feature map, significantly improving the detection recall rate and localization accuracy of large target defects, capable of enhancing the large target detection performance and enhancing the model scale adaptability;
[0148] 3. Cooperative label assignment mechanism: Improve the "backbone + auxiliary head" structure of Co-DETR, further combine the improved decoder positive sample query and label matching mechanism, improve the utilization efficiency of the auxiliary head training supervision, accelerate model convergence, and enhance feature discriminability;
[0149] The first detection model in this embodiment is easy to be compatible with structures such as Co-DETR, easy to deploy and integrate, convenient for the model to be quickly replaced and deployed in engineering systems, and has good scalability and engineering value;
[0150] 4. Engineering deployment and real-time detection capabilities: Accelerate model deployment through ONNX and TensorRT to achieve a low-latency inference speed, meeting the high requirements for real-time performance in underground pipeline network defect detection.
[0151] Therefore, the first detection model trained through steps P1-P6 introduces a relationship-aware self-attention mechanism (RASA) to construct a backbone network based on the improved RSwinTransformer, enhancing the first detection model's ability to model the spatial geometric relationship between targets and improving the first detection model's perception and recognition capabilities in complex structure scenarios such as underground pipeline networks; the improved multi-scale aggregation loss function (MATFL) is designed for large target scenarios, combines a multi-scale weighting mechanism, enhances the response intensity of deep features to large-size targets, and improves target missed detection and localization deviation; the first detection model is superior to related target detection models in terms of accuracy, robustness, and real-time performance, and is suitable for actual detection scenarios such as underground pipeline networks.
[0152] In this embodiment, the trained first detection model can be used to execute the pipeline defect detection method. The pipeline defect detection method includes the following steps:
[0153] S1. Obtain the second detection model;
[0154] S2. Obtain the image to be detected;
[0155] S3. Input the image to be detected into the second detection model for processing;
[0156] S4. Determine the pipeline defect position in the image to be detected according to the processing result of the second detection model.
[0157] In step S1, the second detection model is obtained based on the trained first detection model. Specifically, all the auxiliary heads and multi-scale adjusters in the trained first detection model are discarded to obtain the second detection model. The structure of the second detection model is as Figure 6 shown. The second detection model is applied to the actual underground pipe network defect detection system to realize the deployment, inference, and real-time response of the model.
[0158] When deploying the second detection model, to balance the deployment convenience and inference efficiency, the second detection model can be first exported in ONNX format, and the model graph is optimized and quantized in combination with the TensorRT tool, and finally deployed on the edge computing device with GPU acceleration ability. The latency of the model during the inference process is controlled within 30ms, meeting the real-time detection requirements.
[0159] In the actual detection process, the second detection model can be integrated into the underground pipe network inspection robot system. The robot continuously collects the images in the pipe network during operation as the images to be detected in step S2. Such images to be detected contain the pipeline image content, and step S3 is executed to input them into the second detection model. Referring to Figure 6 , the backbone network processes the image to be detected to obtain image feature data, and the Transformer encoding block encodes the image feature data to obtain encoded feature data. The encoded feature data and the trained learnable query information are input into the Transformer decoding block for processing to obtain the position of the defect in the image to be detected (represented in the form of a bounding box), the defect type (such as structural cracks, foreign object blockages, deformation and collapse, etc.), and the corresponding confidence score output by the Transformer decoding block. The detection results output by the second detection model can be uploaded to the upper-level service module for processing, recording, and alarming.
[0160] In multiple real underground pipeline testing tasks, the second detection model can demonstrate excellent robustness and adaptability. It still has stable detection performance under interference scenarios such as complex lighting conditions, irregular reflections on the pipe wall surface, and blurring. Compared with other object detection models such as the YOLO series and Faster R-CNN, the second detection model has improved in both accuracy and recall, especially when there are multiple medium and large target defects in the image. Visual analysis shows that the second detection model can more accurately locate the defect area and effectively avoid false alarms and missed detections, verifying the effectiveness and practical value of the proposed structural improvement and loss function design in practical engineering applications.
[0161] A computer program for implementing the pipeline defect detection model training method in this embodiment can be written and stored in a computer device or storage medium. When the computer program is read and run, it executes the pipeline defect detection model training method and / or the pipeline defect detection method in this embodiment, thereby achieving the same technical effects as the pipeline defect detection model training method and / or the pipeline defect detection method in the embodiment.
[0162] It should be noted that, unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. In addition, the up, down, left, and right descriptions used in this disclosure are only relative to the mutual positional relationship of the components of this disclosure in the drawings. The singular forms "a", "an", and "the" used in this disclosure are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as commonly understood by those skilled in the technical field of this invention. The terms used in the description of this embodiment are only for describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in this embodiment includes any and all combinations of one or more of the related listed items.
[0163] It should be understood that although terms such as first, second, and third may be used in this disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of this disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as", etc.) provided in this embodiment is only intended to better illustrate the embodiments of the present invention and will not impose a limitation on the scope of the present invention unless otherwise required.
[0164] It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The methods can be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with the computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - in accordance with the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. Additionally, for this purpose the program is capable of running on a programmed application specific integrated circuit.
[0165] In addition, the operations of the processes described in this embodiment can be performed in any suitable order, unless this embodiment otherwise indicates or is otherwise detected to be clearly inconsistent with the context. The processes described in this embodiment (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) commonly executed on one or more processors, by hardware, or a combination thereof. The computer program includes multiple instructions executable by one or more processors.
[0166] Furthermore, the method can be implemented in any type of computing platform operably connected, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, standalone or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it is readable by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to execute the processes described herein. Additionally, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media include instructions or programs that implement the above steps in combination with a microprocessor or other data processor, the invention of this embodiment includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.
[0167] A computer program can be applied to input data to perform the functions of this embodiment, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.
[0168] The above is only a preferred embodiment of the present invention. The present invention is not limited to the above embodiments. As long as it achieves the technical effects of the present invention by the same means, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation manners can have various different modifications and changes.
Claims
1. A method for training a pipeline defect detection model, characterized in that The pipeline defect detection model training method includes: Obtaining a first detection model; the first detection model includes a backbone network, a Transformer encoding block, a Transformer decoding block, a multi-scale adjuster, and at least one auxiliary head, and the backbone network includes multiple stages; Obtaining training images and ground truth boxes; the ground truth boxes represent the real positions of the defective parts included in the training images; Inputting the training images into the backbone network, and each of the stages in the backbone network sequentially processes the training images through a relation-aware self-attention mechanism to obtain image feature data; Inputting the image feature data into the Transformer encoding block for encoding to obtain encoded feature data; Inputting the encoded feature data into the multi-scale adjuster for multi-scale adjustment to obtain a feature pyramid; Traversing all the auxiliary heads; For any traversed auxiliary head, the auxiliary head processes the feature pyramid and obtains multiple positive sample positions according to the ground truth boxes, performs position encoding on each of the positive sample positions to obtain customized query information, inputs the customized query information, learnable query information, and the encoded feature data into the Transformer decoding block for decoding, obtains the multi-scale prediction results of the Transformer decoding block, determines a multi-scale aggregation loss function according to the multi-scale prediction results, and trains the first detection model according to the multi-scale aggregation loss function.
2. The pipeline defect detection model training method according to claim 1, wherein: The backbone network further includes an image partitioning module; The image partitioning module is used to partition the training image into multiple image patches; The first stage includes a linear embedding module and an RSwinTransformer module, and the other stages except the first stage include an image patch fusion embedding module and an RSwinTransformer module; The input data of the first stage is each of the image patches; the input data of any other stage except the first stage is the output data of the previous stage; The linear embedding module is used to perform dimensionality conversion on the input data; The image patch fusion embedding module is used to perform spatial dimensionality reduction on the input data.
3. The pipeline defect detection model training method according to claim 2, wherein The RSwinTransformer module includes a relation-aware self-attention unit and a multi-layer perceptron unit; The relation-aware self-attention unit is used to process the input data of the RSwinTransformer module through a relation-aware self-attention mechanism and then fuse it with the input data of the RSwinTransformer module to obtain the processing result of the relation-aware self-attention unit; The multi-layer perceptron unit is used to process the processing result of the relation-aware self-attention unit and then fuse it with the processing result of the relation-aware self-attention unit to obtain the output data of the RSwinTransformer module.
4. The method for training a pipeline defect detection model according to claim 3, wherein The relationship-aware self-attention unit includes a receiving subunit, a first matrix multiplication subunit, a scaling subunit, a geometric weight subunit, a weighted Softmax subunit, a second matrix multiplication subunit, and a multi-layer perceptron subunit; The receiving subunit is used to receive and obtain multiple data units of the input data of the RSwinTransformer module, and disassemble the data units into content features and position features; The first matrix multiplication subunit is used to perform operation processing according to the formula ; where is the content feature corresponding to the th data unit, is the content feature corresponding to the th data unit, and are the weight matrices of the first matrix multiplication subunit, represents matrix transpose operation; The scaling subunit is used to perform arithmetic processing according to the formula to obtain the content weight ; is the scaling coefficient; The geometric weight subunit is used according to the formula Perform arithmetic processing to obtain geometric weights ; where is the position feature corresponding to the th data unit, is the position feature corresponding to the th data unit, and are the spatial coordinates corresponding to the th data unit, is the width corresponding to the th data unit, is the height corresponding to the th data unit, and are the spatial coordinates corresponding to the th data unit, is the width corresponding to the th data unit, is the height corresponding to the th data unit, represents the GIOU operation, is the weight matrix of the geometric weight sub-unit, represents encoding; The weighted Softmax subunit is used according to the formula Perform arithmetic processing to obtain the total attention weight ; The second matrix multiplication subunit is used according to the formula Perform arithmetic processing; wherein, is the number of data units corresponding to the input data of the RSwinTransformer module, is the weight matrix of the second matrix multiplication subunit; The multi-layer perceptron subunit is used according to the formula Perform arithmetic processing to obtain the processing result of the relationship-aware self-attention unit ; where represents concatenation processing represents multi-layer perceptron processing 5. The method for training a pipeline defect detection model according to any one of claims 1-4, characterized in that Determining the multi-scale aggregation loss function according to the multi-scale prediction result includes: According to the formula Calculate the multi-scale aggregation loss function ; where is the prediction result of the th layer in the multi-scale prediction result, is the scale weighting coefficient corresponding to the th layer, is the adaptive threshold focal loss function, is the total number of layers of the prediction results included in the multi-scale prediction result.
6. The method for training a pipeline defect detection model according to claim 5, wherein: For the average target size corresponding to the layer in the feature pyramid.
7. A pipeline defect detection method, characterized in that, The pipeline defect detection method includes: Obtaining a second detection model; the second detection model is obtained based on the first detection model in any one of claims 1-6; Obtaining an image to be detected; the image to be detected includes pipeline image content; Inputting the image to be detected into the second detection model for processing; Determining the pipeline defect position in the image to be detected according to the processing result of the second detection model.
8. The pipeline defect detection method according to claim 7, wherein The obtaining of the second detection model includes: Obtaining the first detection model; Discarding all the auxiliary heads in the first detection model to obtain the second detection model.
9. A computer device, characterized in that, It includes a memory and a processor. The memory is used to store at least one program, and the processor is used to load at least one program to execute the pipeline defect detection model training method in any one of claims 1-6 and / or the pipeline defect detection method in claim 7 or 8.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to execute the pipeline defect detection model training method in any one of claims 1-6 and / or the pipeline defect detection method in claim 7 or 8.
Citation Information
Patent Citations
Pipeline defect detection method and device based on improved YOLOv8 algorithm
CN119152199A
Pyramid vision Transform-based polyp image segmentation method
CN119399229A