Metal loss detection method of pipeline magnetic leakage signal based on weak annotation
By converting the magnetic flux leakage signal image into an RGB image and combining it with the detection model and classification model, the problem of time-consuming and labor-intensive manual labeling in pipeline magnetic flux leakage signal detection is solved, and the detection accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510863686.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-26
AI Technical Summary
In the existing technology, pipeline magnetic leakage signal detection requires a lot of manual labeling, which makes the labeling process time-consuming and labor-intensive and the detection effect is poor.
A weak labeling-based method is adopted. By converting the leakage magnetic signal image into an RGB image, combining the detection model and classification model, and using heat map and region growing processing, a fitting rectangular frame is generated, which reduces the need for manual labeling and improves detection accuracy.
This reduces the need for manual labeling while improving the accuracy and efficiency of pipeline magnetic leakage signal metal loss detection.
Smart Images

Figure CN120375374B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method for detecting metal loss in pipeline magnetic leakage signals based on weak annotation. Background Art
[0002] Pipeline transportation, as the primary mode of transporting energy sources such as oil and natural gas, offers advantages such as high throughput, low cost, and low energy consumption. However, the risk of leakage from natural corrosion and human damage over long periods of use poses a threat to the economy and the environment. Therefore, regular monitoring of pipeline health is essential. Magnetic flux leakage (MFL) detection technology, which non-destructively acquires physical signatures to analyze pipeline damage, has become a crucial technology for ensuring pipeline integrity.
[0003] To improve detection efficiency, and to address the problems of long pipeline lines and time-consuming and labor-intensive manual inspections, an artificial intelligence detection method based on a deep learning model is adopted. For example, the YOLO series target detection model is used to automatically analyze the leakage magnetic signal, and the metal loss characteristics are identified by constructing a supervised learning model with rectangular box annotations.
[0004] This type of fully supervised model based on detection frame annotation requires a large amount of metal loss sample data to complete training, and the sample annotation process requires manpower, and the annotation detection effect is poor. Summary of the Invention
[0005] The present application provides a pipeline magnetic leakage signal metal loss detection method based on weak labeling to solve the problem that the labeling process needs to rely on manpower and the labeling detection effect is poor.
[0006] This application provides a pipeline magnetic leakage signal metal loss detection method based on weak annotation, including:
[0007] Acquire a magnetic flux leakage signal image, and convert the magnetic flux leakage signal image into an RGB image;
[0008] Using the RGB images as a first training set, and using a preset number of RGB images with rectangular box annotations in the first training set to train a detection model, so that the detection model performs prediction on all the RGB images to generate predicted rectangular boxes;
[0009] Training a classification model using the RGB images with category annotations, so that the classification model obtains category confidence scores for all RGB images;
[0010] generating a heat map based on the classification model;
[0011] Performing region growing processing on the heat map to generate a fitting rectangular box;
[0012] Calculating a fitting score for the RGB image based on an evaluation result of an overlap between the predicted rectangular box and the fitted rectangular box and the category confidence score;
[0013] Based on the fitting score, obtaining a target detection model;
[0014] The target detection model is used to perform metal loss detection on the magnetic leakage signal image to be detected to obtain a detection result.
[0015] In some feasible embodiments, converting the magnetic leakage signal image into an RGB image includes:
[0016] Arrange the magnetic flux leakage signal image into preset rows and preset columns to output a first magnetic flux leakage signal image;
[0017] Based on the preset rows, performing mean filtering and threshold truncation on each row of the first magnetic leakage signal image to output a second magnetic leakage signal image;
[0018] Performing a color scale transformation on the second magnetic leakage signal image to output an RGB image.
[0019] Circumferential noise is suppressed through axial mean filtering, sensor drift is eliminated through dynamic threshold truncation, and positive and negative deviations of magnetic flux are correlated with red and blue scale mapping to enhance the visual recognition of defect features.
[0020] In some feasible embodiments, the RGB images are used as a first training set, and a detection model is trained using a preset number of RGB images with rectangular box annotations in the first training set, so that the detection model performs prediction on all the RGB images to generate predicted rectangular boxes, including:
[0021] Acquire a preset number of RGB images in the first training set, where the preset number of RGB images are RGB images including specific features;
[0022] Marking the preset number of RGB images with rectangular frames to obtain rectangular frame marked samples;
[0023] Dividing the rectangular frame labeled samples into a detection training subset and a detection verification subset;
[0024] Set training parameters;
[0025] Based on the training parameters, performing iterative training on the detection training subset to obtain a detection model;
[0026] Input all the RGB images to the detection model, so that the detection model outputs the predicted rectangular box corresponding to each RGB image.
[0027] Representative sample layer screening combined with pre-training weight transfer and multi-scale anchor frame adaptation to pipeline defect morphology effectively prevent overfitting in small sample training and accelerate model convergence.
[0028] In some feasible embodiments, the classification model includes a feature extraction backbone network, a feature fusion network, and a prediction head module;
[0029] The method of training a classification model using the RGB images with category annotations so that the classification model obtains category confidence scores for all RGB images includes:
[0030] Generate a multi-scale feature map through the feature extraction backbone network;
[0031] fusing the multi-scale feature maps in the feature fusion network to obtain a first depth feature map;
[0032] Processing the first depth feature map by the prediction head module to obtain a first splicing result;
[0033] performing non-maximum suppression processing on the first splicing result, removing detection frames with a repetition rate higher than a threshold in the first splicing result, to output a second splicing result;
[0034] Map the size of the detection box in the second stitching result to the size of the RGB image to output a category confidence score for each detection box.
[0035] The feature pyramid integrates deep and shallow semantics, the four-level prediction head covers full-size defects, and channel splicing integrates multi-scale information, significantly improving the problem of missed detection of micro-pitting targets.
[0036] In some feasible embodiments, the feature extraction backbone network includes a convolutional layer;
[0037] Generating a heat map based on the classification model includes:
[0038] Calculating a response weight factor of the target category to the feature channel in the convolutional layer, wherein the response weight factor represents the contribution of the feature channel to the target category judgment;
[0039] Multiplying the characteristic spectrum of the characteristic channel by the response weight factor to obtain a weighted characteristic spectrum;
[0040] superimposing the weighted feature maps along the feature channels to obtain a superposition result;
[0041] performing linear rectification activation processing on the superposition result to obtain an activation area;
[0042] retaining the positive activation area in the activation area as the initial heat map;
[0043] The spatial resolution of the initial heat map is matched to the size of the RGB image to generate a heat map.
[0044] Gradient weighting focuses on key feature channels, and heat map overlay and activation filtering highlight the discriminant areas, realizing spatial positioning visualization of the classification decision basis.
[0045] In some feasible embodiments, the prediction head module includes a first prediction head, a second prediction head, a third prediction head and a fourth prediction head;
[0046] The processing of the first depth feature map by the prediction head module to obtain a first splicing result includes:
[0047] Processing the first depth feature map by the first prediction head to obtain a second depth feature map;
[0048] Processing the second depth feature map by the second prediction head to obtain a third depth feature map;
[0049] Processing the third depth feature map by the third prediction head to obtain a fourth depth feature map;
[0050] Processing the fourth depth feature map by the fourth prediction head;
[0051] The first depth feature map, the second depth feature map, the third depth feature map, and the fourth depth feature map are spliced in the feature dimension to output a first splicing result, where the first splicing result includes multiple detection boxes.
[0052] Hierarchical upsampling preserves the integrity of high-level features, shallow features are directly connected to avoid detail loss, and multi-resolution features are uniformly spliced to break through the bottleneck of cross-scale fusion.
[0053] In some feasible embodiments, performing a region growing process on the heat map to generate a fitting rectangular frame includes:
[0054] Identifying pixels in the heat map that exceed an activation intensity threshold to determine highlighted pixels;
[0055] Aggregating the spatially adjacent highlighted pixel points to form an initial activation area cluster;
[0056] Calculating the geometric center of gravity of the initial activation area cluster;
[0057] Determining the geometric center of gravity point in the geometric center of gravity position as the starting seed point of region growth;
[0058] Performing multi-directional pixel expansion with the starting seed point as the center to obtain neighborhood pixel points;
[0059] Calculating the feature similarity between the neighborhood pixel points and the starting seed point;
[0060] If the feature similarity is greater than a similarity threshold, the neighborhood pixel points are merged into a growing region to generate a final growing region, where the final growing region includes a plurality of independent regions;
[0061] Performing spatial connected domain labeling on the final growth region and extracting an external contour boundary point set of the independent region;
[0062] Calculating the coordinate span value of the boundary point set in the image coordinate system;
[0063] Based on the coordinate span value, a minimum bounding rectangle is determined, where the minimum bounding rectangle is a fitted rectangular frame.
[0064] Multi-directional region growing constrained by thermal value similarity, combined with contour geometric feature extraction, achieves sub-pixel fitting of defect boundaries under weak supervision conditions.
[0065] In some feasible embodiments, calculating the fitting score of the RGB image by combining the overlap evaluation result between the predicted rectangular box and the fitted rectangular box and the category confidence score includes:
[0066] Calculating an intersection-and-union ratio between the fitted rectangular frame and the predicted rectangular frame, where the intersection-and-union ratio represents a degree of spatial position overlap;
[0067] Obtaining a first weight coefficient and a second weight coefficient;
[0068] A fitting score is generated by using the first weight coefficient, the second weight coefficient, the category confidence score, and the intersection-over-union ratio.
[0069] The intersection-over-union ratio and confidence factor dual-factor coupling evaluation are used, and the weight coefficient dynamically balances the contribution of spatial positioning and feature discrimination to build a high-reliability pseudo-annotation screening mechanism.
[0070] In some feasible embodiments, obtaining the target detection model based on the fitting score includes:
[0071] adding RGB images exceeding a fitting score threshold to the first training set to generate a second training set;
[0072] The detection model is trained using the second training set to obtain a target detection model.
[0073] The fitting score threshold controls the quality of pseudo-annotation, and the full network parameter optimization combined with the early stopping mechanism drives the model's iterative self-reinforcement and improves generalization performance.
[0074] In some feasible embodiments, performing metal loss detection on the magnetic leakage signal image to be detected by the target detection model to obtain a detection result includes:
[0075] Acquire the magnetic flux leakage signal image to be measured;
[0076] Converting the magnetic flux leakage signal image to be measured into an RGB image;
[0077] The RGB image is input into the target detection model to obtain a detection result, which includes the spatial position coordinates of the predicted rectangular box, the target category label, and the category confidence score.
[0078] As can be seen from the above technical solution, the present application provides a pipeline magnetic flux leakage signal metal loss detection method based on weak annotation, comprising acquiring a magnetic flux leakage signal image and converting the magnetic flux leakage signal image into an RGB image; using the RGB image as a first training set, and training a detection model with a preset number of RGB images with rectangular box annotations in the first training set, so that the detection model performs predictions on all the RGB images to generate predicted rectangular boxes; using the RGB images with category annotations to train a classification model, so that the classification model obtains category confidence scores for all RGB images; generating a heat map based on the classification model; performing region growing processing on the heat map to generate a fitted rectangular box; combining the overlap evaluation results of the predicted rectangular box and the fitted rectangular box and the category confidence score to calculate the fitting score of the RGB image; obtaining a target detection model based on the fitting score; and performing metal loss detection on the magnetic flux leakage signal image to be tested using the target detection model to obtain a detection result. The method reduces the need for manual annotation while improving the accuracy of metal loss detection by combining the detection model with the classification model, the region growing guided by the heat map, and the multi-dimensional fitting evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0080] Figure 1 A flow chart of a pipeline magnetic leakage signal metal loss detection method based on weak annotation provided in an embodiment of the present application;
[0081] Figure 2 A schematic diagram of the overall flow of the detection process provided in the embodiment of the present application;
[0082] Figure 3 A schematic diagram of the structure of the detection model and classification model provided in the embodiment of the present application;
[0083] Figure 4 A schematic diagram of the flow of a region growing algorithm for generating fitted rectangular holes from a heat map provided in an embodiment of the present application. DETAILED DESCRIPTION
[0084] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.
[0085] In some embodiments, a pipeline magnetic flux leakage defect detection method based on the fully supervised network YOLOv5 uses axial curve images of magnetic flux leakage data as model input, uses automatic search data augmentation techniques to address the problem of insufficient sample size, and combines multi-scale convolution kernel feature extraction to achieve multi-scale feature fusion. Although this method expands the metal loss dataset through data augmentation, it still requires manual annotation of training samples. In addition, the axial curve images used in this method do not clearly distinguish between defects and surrounding non-defective areas, which affects the effectiveness of the fully supervised network YOLOv5, which uses image features as the detection criterion.
[0086] In other embodiments, a method for detecting ore clumps based on a weakly supervised YOLO model uses transfer learning from a model pre-trained on the public COCO dataset. This method employs an active learning strategy to filter samples with uncertain features and provide them to experts for labeling, significantly reducing the amount of sample labeling required. Furthermore, a feature pyramid network layer is used to integrate multi-scale feature information, enhancing the model's generalization. The ore clump image samples used in this method are similar to the natural scene images in the COCO dataset, making transfer learning easier. However, magnetic leakage signal data images, due to their unique features, are more difficult to implement with this general-purpose weakly supervised learning strategy.
[0087] In summary, the sample labeling process in the above methods relies on manpower, and the labeling detection effect is poor.
[0088] To solve the above problems, Figure 1 , Figure 2 As shown, the embodiment of the present application provides a pipeline magnetic leakage signal metal loss detection method based on weak annotation, including:
[0089] S100: Acquire a magnetic flux leakage signal image and convert the magnetic flux leakage signal image into an RGB image.
[0090] A magnetic flux leakage signal image is a raw two-dimensional data matrix collected by a magnetic sensor array within a pipeline inspection device. The image's rows correspond to the axial distance sampling points along the pipeline, and the columns correspond to the circumferential sensor channel numbers. The image's grayscale values quantify the intensity of magnetic flux leakage, with abnormal fluctuations exceeding the baseline threshold indicating potential metal loss defects. In practical applications, this image is acquired by a magnetic flux leakage detector onboard a pipeline crawler and stored in a single-channel 16-bit TIFF format.
[0091] Because the original magnetic flux leakage signal image features are not distinct enough, the model needs to learn more distinct features for better detection results. The magnetic flux leakage signal image is converted into an RGB image, a three-channel color image. This conversion process preserves the physical properties of the original data but uses a color scale transformation to visualize the magnetic flux variation trend: positive deviations are mapped to the red channel (R), negative deviations are mapped to the blue channel (B), and the baseline area is mapped to the green channel (G). This conversion enhances the human eye's ability to discern weak signal differences and provides a standardized three-dimensional input tensor for the convolutional neural network.
[0092] In some embodiments, a mapping of grayscale values to color space is adopted, and single-channel values are converted into three channels according to preset rules. The grayscale value range is divided into three segments and mapped to RGB channels respectively. For example, the low-value area is mapped to the blue channel (background noise), the middle-value area is mapped to the green channel (normal pipe wall), and the high-value area is mapped to the red channel (defect signal), forming an intuitive cold-warm color scale contrast.
[0093] Single-channel to RGB conversion can be achieved, but the noise sensitivity and weak feature dominance problems unique to magnetic leakage signals are not solved. In some embodiments, the magnetic leakage signal image is set to preset rows and preset columns to output a first magnetic leakage signal image; based on the preset rows, mean filtering and threshold truncation are performed on each row of the first magnetic leakage signal image to output a second magnetic leakage signal image; and color scale conversion is performed on the second magnetic leakage signal image to output an RGB image.
[0094] The collected original magnetic leakage signal images may be multiple, for example, 200 images. During the collection process, the image sizes may be different. Before converting to an RGB image, the magnetic leakage signal image is first set to a preset pixel size, that is, a preset row and a preset column.
[0095] Preset rows and columns are parameters used to normalize the size of the raw magnetic flux leakage signal image. Raw data is collected by the magnetic sensor array of the pipeline inspection instrument, forming a two-dimensional signal matrix. The rows correspond to the axial distance of the pipeline, and the columns correspond to the circumferential angle of the pipeline. By forcibly resetting the image to a fixed size, for example, 2400 rows × 5000 columns, image distortion caused by fluctuations in the inspection equipment's movement speed is eliminated, resulting in a standardized digital image base.
[0096] For example, the raw magnetic flux leakage signal is input as a two-dimensional matrix of irregular dimensions. Using a bilinear interpolation algorithm, it is resampled to a preset standard size of 2400 rows by 5000 columns. This ensures uniform spatial resolution for pipeline inspection data across different sections, establishing a reference coordinate system for subsequent processing.
[0097] That is, the first magnetic flux leakage signal image is a plurality of images of preset rows and preset columns.
[0098] In the row direction (i.e., the channel dimension), mean filtering is performed, with upper and lower thresholds applied to avoid outliers with values that are too large or too small. Mean filtering is a linear smoothing operation performed in the row dimension. A fixed-width sliding window is used, moving pixel by pixel along each row of data. The arithmetic mean of all pixel values within the window's coverage area is calculated and used to replace the original pixel value at the center of the window. This operation suppresses spike noise caused by electromagnetic interference while preserving the continuous, slowly varying characteristics of metal loss defects.
[0099] Exemplarily, the first magnetic leakage signal image is processed line by line, and a sliding window with a width of 5 pixels is set. The window moves from the beginning to the end of each line, sliding 1 pixel each time. The arithmetic mean of the 5 pixels in the window is calculated in real time, and the calculation result is used to overwrite the original pixel value at the center of the window. This process effectively smoothes the circumferential random noise while retaining the axial continuity feature.
[0100] Threshold truncation is a core step in dynamic range compression. First, the absolute maximum value of the filtered image's pixel values is calculated and used as the symmetrical truncation threshold in the positive and negative directions. Pixel values outside the positive threshold are forced to the positive threshold, while pixel values below the negative threshold are forced to the negative threshold. This operation eliminates outliers caused by sensor zero drift or saturation, ensuring that data distribution remains within a reasonable range.
[0101] The upper and lower thresholds are set by the following formula:
[0102] ;
[0103] Where D represents the data after mean filtering, are pairs of adjacent axial pixels located at the same circumferential angle position in the magnetic flux leakage signal image.
[0104] For example, the entire image is scanned to find the maximum positive pixel value and the minimum negative pixel value, and the absolute values of the two are compared. The one with the larger absolute value is taken as the reference threshold, and symmetrical positive and negative cutoff thresholds are set. All pixels are traversed, and pixels exceeding the positive threshold are set as the positive threshold, and pixels below the negative threshold are set as the negative threshold to eliminate errors and standardize the data.
[0105] The second magnetic flux leakage signal image is a single-channel intermediate image generated by size normalization, axial mean filtering, and dynamic threshold truncation of the original single-channel magnetic flux leakage data. It maintains the same spatial resolution as the original data (2400 rows × 5000 columns), but the pixel value range is constrained to a symmetrical interval.
[0106] The magnetic flux baseline value of the intact pipeline area is taken as the zero point, and is truncated upward to the positive threshold (+threshold, taking the positive limit of the absolute maximum value of the data) and downward to the negative threshold (-threshold, taking the negative limit of the absolute maximum value of the data). Abnormal values outside this range are set at the boundary value, thereby eliminating sensor noise and pulse interference while retaining the key magnetic anomaly characteristics of the metal loss area. The positive deviation area corresponds to the enhanced leakage magnetic field caused by thinning of the pipe wall, and the negative deviation area corresponds to the magnetic field distortion caused by magnetic foreign matter or welds. The area near zero represents the intact pipe wall.
[0107] The second magnetic flux leakage signal image is then subjected to a color scale transformation, and the values within the positive and negative threshold ranges are mapped to red to cyan according to the warm and cold color changes, so that the originally monochrome magnetic flux leakage data image is transformed into a color image with a size of (2400×5000×3), which is convenient for the model to detect and classify.
[0108] Exemplarily, a three-channel output image structure is created, red channel generation: for positive pixels, the red intensity is calculated according to the ratio of their value in the interval, and the red channel is set to zero at negative and zero value positions; blue channel generation: for negative pixels, the blue intensity is calculated according to the ratio of their absolute value in the interval, and the blue channel is set to zero at positive and zero value positions; green channel generation: only the pixel positions where the magnetic flux is strictly zero are assigned a medium green value, and the rest of the positions are set to zero. In the final output RGB image, the metal loss area appears red or blue according to the direction of the magnetic anomaly, and the defect-free area appears green.
[0109] In this way, magnetic flux direction is encoded as hue differences, and intensity is encoded as saturation. The human eye and convolutional neural networks are significantly more sensitive to hue differences than to grayscale variations. This makes weak signal features explicit, improving the model's ability to identify minor metal damage and reducing the risk of missed detections. Row-wise filtering preserves axial continuity (metal damage extends along the pipe) while suppressing random noise. Dynamic thresholding eliminates inherent device bias, enhancing the model's generalization across devices and reducing detection fluctuations caused by sensor differences.
[0110] S200: Using RGB images as a first training set, and using a preset number of RGB images with rectangular box annotations in the first training set to train a detection model, so that the detection model performs prediction on all the RGB images to generate predicted rectangular boxes.
[0111] This solution uses a semi-supervised learning strategy. For example, first, 20 samples with a few representative features in the training set of 200 RGB images are accurately annotated with rectangular boxes, and then the categories of the entire dataset of 200 images are annotated.
[0112] The first training set consists of all RGB images converted from magnetic flux leakage signals, known as training samples. Its core characteristic is weak annotation. For example, only 10%-15% of the images contain manually annotated rectangular boxes (the annotated areas indicate metal loss defects), while the remaining 85%-90% contain only image-level category labels (presence / absence of defects).
[0113] In some embodiments, obtaining a preset number of RGB images in the first training set;
[0114] Marking the preset number of RGB images with rectangular frames to obtain rectangular frame marked samples;
[0115] Dividing the rectangular frame labeled samples into a detection training subset and a detection verification subset;
[0116] Set training parameters;
[0117] Based on the training parameters, iterative training is performed on the detection training subset to obtain a detection model, i.e., an initial version of the YOLO detection model;
[0118] Input all the RGB images to the detection model, so that the detection model outputs the predicted rectangular box corresponding to each RGB image.
[0119] The preset number of RGB images are a preset number of RGB images including specific features, for example, 20 RGB images out of 200 RGB images contain metal loss defects, such as pitting corrosion and groove corrosion.
[0120] Specific features refer to the core morphological properties of metal loss defects, which may include size features, morphological features, and intensity features. For example, the size feature is a continuous signal area with an axial length greater than 5mm and less than 50mm; the morphological feature is an elliptical or long strip closed signal area; the intensity feature is a magnetic flux change exceeding ±30% of the baseline. Samples with the above features are judged to be valid annotation objects.
[0121] Rectangular box annotation is a bounding box operation drawn on the RGB image using a manual annotation tool. The box must be parallel to the image coordinate axes on all four sides, completely cover the defect signal area and edge transition zone, and the box boundary must be at least 1 pixel away from the outer edge of the signal fluctuation area. The annotation data is stored as a YOLO format text file.
[0122] Traditional methods require full labeling, which leads to underfitting of small sample scenario models. Screening based on the defect feature library ensures full coverage of key samples.
[0123] The samples annotated by the rectangular boxes are divided into a detection training subset and a detection verification subset, where the training set and the verification set are divided into 7:3 to train the detection model.
[0124] Perform model training and set the network training parameters, namely, training round epoch, batch size batchsize, learning rate and momentum, and early stopping parameter patience. The epoch is set to 600, the batch size is 8, the learning rate is 0.1, the momentum is 0.937, and the early stopping parameter is set to 300.
[0125] The predicted rectangle is a parameterized representation of the bounding box output by the forward propagation of the detection model. Each box contains the center point coordinates (x, y), floating point values (normalized to 0-1), width and height (w, h), the ratio relative to the image size, a confidence score, and a 0-1 probability value.
[0126] Exemplarily, the trained detection model processes all images in the first training set. The input of the detection model is an RGB image (maintaining the original size of 2400×5000), and the output is a set of predicted rectangular boxes for each image.
[0127] Through semi-supervised learning strategies and iterative training, the dependence on sample annotation is reduced while ensuring better training results.
[0128] S300: Training a classification model using the RGB images with category annotations, so that the classification model obtains category confidence scores for all RGB images.
[0129] The classification model is a convolutional neural network focused on image-level discrimination, such as the ResNet50 architecture. The model takes an RGB image as input and outputs a discrete class probability distribution (in the range [0.0, 1.0]). Model training relies only on weakly supervised data with two labels: metal loss or no defect.
[0130] The category-labeled samples after category labeling are divided into a detection training subset and a detection verification subset, among which the training set and the verification set are divided into 7:3 to train the classification model.
[0131] Similar to the training process of the detection model, after generating the training set and validation set, the classification model is obtained by setting the training parameters and iterative training.
[0132] In some embodiments, the classification model includes a feature extraction backbone network, a feature fusion network, and a prediction head module. The feature extraction backbone network is the core architectural component of the classification model and is implemented using a deep convolutional neural network. Its structure comprises multiple cascaded convolutional modules, each consisting of a convolutional layer, a batch normalization layer, and an activation function. The backbone network receives RGB image input, extracts abstract features through layer-by-layer downsampling, and outputs a set of multi-scale feature maps.
[0133] The feature fusion network is the intermediate structure connecting the backbone network and the prediction head. It employs a feature pyramid architecture to fuse information from feature maps of different scales through both top-down and bottom-up pathways. This network enhances the model's ability to perceive objects of various sizes, particularly capturing subtle metal loss features.
[0134] The prediction head module is used to identify metal defects of different sizes.
[0135] In some embodiments, a multi-scale feature map is generated by the feature extraction backbone network;
[0136] fusing the multi-scale feature maps in the feature fusion network to obtain a first depth feature map;
[0137] Processing the first depth feature map by the prediction head module to obtain a first splicing result, where the first splicing result includes a plurality of detection boxes;
[0138] performing non-maximum suppression processing on the first splicing result, removing detection frames with a repetition rate higher than a threshold in the first splicing result, to output a second splicing result;
[0139] Map the size of the detection box in the second stitching result to the size of the RGB image to output a category confidence score for each detection box.
[0140] Exemplarily, the classification model adopts an improved YOLO architecture, the feature extraction backbone is the CSPDarknet53 network, which includes multiple downsampling stages, and the feature fusion network is the FPN and PANet bidirectional pyramid. The top-down path transfers deep semantic features to the shallow layer, and the bottom-up path transfers shallow detail features to the deep layer.
[0141] The first deep feature map is the set of multi-scale feature maps after output fusion. The first concatenation result refers to the concatenation of the feature maps output by the prediction head module along the channel dimension. After resizing the feature maps of different scales to a uniform spatial size, they are concatenated along the channel axis to form a high-dimensional feature tensor, which serves as input for subsequent processing.
[0142] Non-maximum suppression is a post-processing algorithm that eliminates redundant detection frames. By calculating the intersection over union (IoU) between detection frames, it filters out overlapping frames within the same target region, retaining only the detection frames with the highest confidence. For example, all detection frames are sorted in descending order of confidence, the highest-scoring frame is selected, and adjacent frames with an IoU (Intersection over Union) ratio (IoU) greater than 0.5 are removed. This process is repeated until no more frames remain to be processed, and the second stitching result is output.
[0143] Finally, the image is resized to restore the output image to the input size and the coordinate information of the detection box and the category confidence score are output.
[0144] To detect defects of different sizes, in some embodiments, the prediction head module includes a first prediction head, a second prediction head, a third prediction head and a fourth prediction head, wherein the first prediction head processes deep feature maps and is responsible for identifying large-size defects, the second prediction head processes middle-layer feature maps and identifies medium-size defects, the third prediction head processes shallow feature maps and captures small-size defects, and the fourth prediction head is directly connected to the shallowest output of the backbone network and is dedicated to micro-defect detection.
[0145] Processing the first depth feature map through the first prediction head to obtain a second depth feature map; processing the second depth feature map through the second prediction head to obtain a third depth feature map; processing the third depth feature map through the third prediction head to obtain a fourth depth feature map; processing the fourth depth feature map through the fourth prediction head; splicing the first depth feature map, the second depth feature map, the third depth feature map and the fourth depth feature map in the feature dimension to output a first splicing result.
[0146] The first deep feature map is the initial multi-scale feature set output by the feature fusion network. For example, it contains four levels of feature maps: 160×160 shallow detail features, 80×80 mid-level features, 40×40 mid-high-level features, and 20×20 high-level semantic features. Each level of feature map has a different spatial resolution and number of channels (shallow layers have fewer channels, deep layers have more).
[0147] The second deep feature map is the output feature after processing by the first prediction head. By performing a 3×3 convolution operation on the 20×20 high-level feature map, the representation capability of semantic information is enhanced. The output size remains 20×20, but the channel dimension is expanded. This feature map focuses on the global morphological characteristics of large-scale defects.
[0148] The third deep feature map is the output feature after processing by the second prediction head. This second deep feature map is fed into a module consisting of two serial convolutional layers (with kernel sizes of 3×3 and 1×1). This reorganizes the features while preserving spatial information, resulting in an output size of 40×40. This feature map identifies medium-sized defects.
[0149] The fourth depth feature map is the output feature after processing by the third prediction head. A strided convolution operation (stride = 2) is performed on the third depth feature map, increasing the spatial resolution to 80×80. This process restores detailed information through the deconvolution layer, and the output features focus on the local characteristics of small defects.
[0150] The first concatenation result is a fused feature formed by concatenating the four-level feature maps along the channel dimension. For example, the 20×20 feature map is bilinearly upsampled to 80×80, the 40×40 feature map is bilinearly upsampled to 80×80, and the 160×160 feature map is max-pooled downsampled to 80×80. All feature maps are then concatenated along the channel axis.
[0151] In this way, the four-level prediction heads process features of different resolutions respectively and finally splice them at a unified scale. The first prediction head maintains the integrity of high-level semantics, the third prediction head restores spatial details through transposed convolution, and the fourth prediction head retains the original shallow information, which can achieve a balanced expression of micro details and macro semantics.
[0152] The first stitching result includes multiple detection frames, which are used as inputs for non-maximum suppression processing. By calculating the intersection-over-union ratio between the detection frames, overlapping frames of the same target area are screened, and only the detection frame with the highest confidence is retained.
[0153] In this embodiment, the detection model and the classification model are both YOLOv51 model structures, such as Figure 3 As shown, the solid line in the figure represents the original YOLOv5l model structure.
[0154] YOLOv5l consists of three modules: backbone, neck, and head modules. The initial image input is uniformly processed into a size of (3×640×640). The backbone module is the core module of YOLOv5l, which includes CBS, namely convolution CONV, batch normalization BN, activation function SiLU, and a cross-stage local module (CSP (C3)) to avoid gradient disappearance. The output of the CSP module is a feature map. From the shallow to the deep layer of the network, there are multiple CSPs. The output feature map sizes of different CSP modules are different. The deeper the depth, the smaller the size.
[0155] The backbone also includes a Spatial Pyramid Fast Pooling (SSPF) module, which generates a feature map with the smallest size (20×20) but the largest number of channels (1024). The neck module is a feature pyramid network (FPN) and path aggregation network (PAN) architecture based on the feature pyramid. It uses top-down side connections to construct high-level semantic feature maps at all scales. To enhance underlying target information, the PAN adds an upward FPN to enhance positioning information from the bottom up, generating prediction heads at multiple scales. Each prediction head has a size of (1×3×H×W×5+cls), where H and W represent the feature map size, and the last dimension is 5+cls. cls is the number of categories, representing the bounding box coordinates (x, y, w, h) and object confidence. In this solution, cls is set to 1 for metal loss detection, and the last dimension is 5+1=6.
[0156] Because the size of metal loss in magnetic flux leakage signals varies from small to large, and missing small metal loss detection is also a major hazard to pipeline safety, the original YOLOv5l network was improved. The improved structure is shown in the dotted line in the figure. The shallower feature map (128×160×160) of the Backbone module is also input into the Neck module. This is built according to the original structure to obtain prediction head 4, the fourth prediction head. Because prediction head 4 uses a larger feature map (160×160), it can better detect small objects. The four prediction heads are concatenated to a feature dimension of 3×(160×160+80×80+40×40+20×20)=(3×102000). Non-maximum suppression is performed to filter out excessively overlapping detection frames, retaining only the most important ones. Finally, the output image is scaled back to the input size, and the detection frame xy coordinate information and category confidence score are output, completing the task of detecting metal loss of different sizes in actual pipelines.
[0157] For classification models, in addition to YOLOv5, other convolutional network-based classification models can also be used, such as Deformable ConvNets, ResNet-101, and Vgg-19.
[0158] S400: Generate a heat map based on the classification model.
[0159] To automatically generate rectangular boxes, a heatmap is first calculated using a classification model. In some embodiments, a response weight factor is calculated for the target category to the feature channels in the convolutional layer. The response weight factor represents the contribution of the feature channel to the target category judgment. The response weight factor is a parameter vector that quantifies the contribution of the feature channel to the target category judgment. This factor is obtained by calculating the gradient of the target category output score with respect to the feature channel. Each channel is assigned a weight value, and its absolute value represents the degree of influence of the channel feature on the target category decision.
[0160] The feature map of the feature channel is multiplied by the response weight factor to obtain a weighted feature map. The feature map is a three-dimensional tensor data structure output by the convolutional layer. It contains spatial dimensions (width × height) and channel dimensions (number of feature channels). The vector at each spatial position represents an abstract feature expression of the local area. In classification models, deep feature maps contain high-level semantic information, while shallow maps retain detailed texture features. The weighted feature map is a tensor obtained by multiplying the response weight factor by the original feature map channel by channel. This operation amplifies the signal strength of important feature channels, suppresses interference from irrelevant channels, and achieves adaptive selection of feature channels.
[0161] The weighted feature maps are superimposed along the feature channels to obtain a superposition result, which is a two-dimensional matrix of the weighted feature maps superimposed along the channel dimension. By accumulating the weighted feature values of multiple channels at the same spatial location, a spatial activation distribution map is formed, highlighting the key areas related to the target category.
[0162] The superposition result is then subjected to linear rectification activation to obtain the activation region. Linear rectification activation is a mathematical operation that transforms the superposition result through the ReLU function. This function retains the positive values of the input matrix and sets the negative values to zero, thereby filtering out feature regions that contribute positively to the target category prediction.
[0163] The positive activation regions within the activation regions are retained as the initial heatmap. The activation regions are connected regions consisting of positive pixels retained after ReLU activation. These regions appear as bright areas in the heatmap, and their spatial distribution closely matches the visual features underlying the model's decision making. The positive activation region specifically refers to a subset of pixels whose activation intensity exceeds a preset threshold (e.g., intensity greater than 70% of the maximum value across the entire image). This region eliminates weak noise interference and focuses on the areas underlying the model's judgment.
[0164] The spatial resolution of the initial heatmap is matched to the size of the RGB image to generate a heatmap. Spatial resolution matching is the process of restoring the heatmap to the input image size using an interpolation algorithm. The low-resolution heatmap is scaled up using a bilinear interpolation algorithm so that each pixel in the original image has a corresponding thermal value.
[0165] The calculation formula of the heat map is:
[0166] ;
[0167] Among them, c represents the category of the selection area, k represents the kth channel, A represents the feature layer that needs to be visualized for feature maps, and the last convolutional layer in the YOLOv5l model is selected. Represents the weight of category c on the kth channel of feature layer A. is the weight matrix on the kth channel of feature layer A. ReLU is a linear activation function used to make the final output greater than 0 and suppress the weight parts that are not of interest.
[0168] GradCAM's heatmap highlights areas of the input image that contribute to the classification result in red, while suppressing areas that do not contribute in blue. Ideally, the highlighted area should be located at the center of each metal loss detection rectangle, serving as the basis for automatically generated rectangles.
[0169] For example, the final convolutional layer of the classification model (size 8×8×2048) is selected as the feature source. After forward propagation obtains the predicted score for the target category (metal loss), a backward gradient calculation is performed to calculate the partial derivative of the predicted score with respect to the feature map. The gradient values are globally averaged and pooled along the channel dimension to obtain a 2048-dimensional response weight vector.
[0170] The response weight vector is multiplied by the feature map channel by channel, and each channel slice of the feature map is multiplied by the weight factor value corresponding to the channel. The weighted feature map (size 8×8×2048) is output. The weighted feature map is superimposed along the channel dimension, and the channel values at the same spatial position are accumulated and summed to generate a two-dimensional activation intensity map (size 8×8). It is processed by the ReLU function: positive values retain the original values and negative values are reset to zero.
[0171] Calculate the maximum value of the activation intensity map, set the intensity threshold (such as 70% of the maximum value), retain the pixel area exceeding the threshold, and use bilinear interpolation to upsample the 8×8 heat map to the original size of 2400×5000. Normalize the heat value to the range of 0-255, map the low-intensity area to blue, and map the high-intensity area to red and yellow, and overlay it with the original image.
[0172] The feature channel contribution evaluation is established based on the weight factor of the gradient response, and the deep and shallow layer features are weightedly fused to capture global semantics and local details simultaneously. The adaptive activation intensity threshold eliminates noise interference.
[0173] To obtain heatmaps, in addition to GradCAM, other class activation heatmap CAM methods can also be used, such as HiResCAM, LayerCAM, etc.
[0174] S500: performing region growing processing on the heat map to generate a fitting rectangular frame.
[0175] Region growing is an adaptive image segmentation algorithm based on seed points. Pixels in the heat map exceeding an intensity threshold (0.7) are used as seed points. The intensity similarity of pixels in the eight-neighborhood region is calculated (Euclidean distance < 0.05). Connected regions are expanded through multiple rounds of iteration until the similarity of boundary pixels falls below the growing threshold.
[0176] The fitted rectangle is the product of calculating the minimum enclosing rectangle of the region growing result. The algorithm extracts the edge contour points of the growing region and uses the rotating calculus method to solve for the minimum enclosing rectangle. The coordinates of the four vertices of this rectangle are mapped to the original image coordinate system through a perspective transformation.
[0177] To further improve the quality of GradCAM thermal map, the region growing algorithm is used to merge pixels with similar properties, such as Figure 4 shown.
[0178] In some embodiments, pixels in the heat map that exceed an activation intensity threshold are identified to determine highlighted pixels. The activation intensity threshold is a critical parameter for screening significant pixels in the heat map. The threshold is set based on a ratio of the global maximum intensity value of the heat map (e.g., 70% of the maximum value) to distinguish between valid activation areas and background noise. Threshold calculation is a dynamic process that adapts to the intensity distribution characteristics of different images. Highlighted pixels refer to a set of coordinate points in the heat map that exceed the activation intensity threshold. These points form a discrete distribution in space, representing the key feature areas on which the model relies for discrimination, and their aggregation morphology is related to the actual location of the metal loss defect.
[0179] The spatially adjacent highlighted pixels are aggregated to form an initial activation region cluster. The initial activation region cluster is a collection of highlighted pixels aggregated based on the spatial proximity principle. Adjacent highlighted pixels are connected into independent clusters using neighborhood connectivity, with each cluster representing a candidate defect region.
[0180] Calculate the geometric center of gravity of the initial activation cluster. This is the coordinate of the center of mass of the initial activation cluster. This is obtained by taking the arithmetic mean of the coordinates of all pixels within the cluster. This point serves as the starting seed point for the region growing algorithm, ensuring that the growth process begins in the defect core region.
[0181] The geometric center of gravity in the geometric center of gravity position is determined as the starting seed point for region growth. The starting seed point for region growth is the growth origin selected in the geometric center of gravity position. When there are multiple clusters, the centroid of the cluster with the largest area is selected as the main seed point, and the remaining clusters are used as auxiliary seed points to form a multi-center growth model.
[0182] Multi-directional pixel expansion is performed with the starting seed point as the center to obtain neighboring pixels. Multi-directional pixel expansion is a radial growth method centered on the seed point. Each iteration examines the eight neighboring pixels of the current growth area boundary, calculates the feature similarity between the neighboring pixels and the seed point, and decides whether to include them in the growth area.
[0183] The feature similarity between the neighborhood pixels and the starting seed point is calculated. If the feature similarity is greater than a similarity threshold, the neighborhood pixels are incorporated into the growth region to generate a final growth region. The final growth region consists of multiple independent regions, which are stable pixel sets formed after multiple rounds of expansion. This region meets the internal connectivity requirements, has closed boundaries, and has consistent internal pixel similarity, providing complete coverage of metal loss defects.
[0184] The final growth region is spatially connected domain labeling, which analyzes the topological structure of the final growth region. Independent connected regions are identified using a scanline algorithm, each of which is assigned a unique identifier to distinguish between multiple possible defect targets in the image. The external contour boundary point set of each independent region is extracted. The external contour boundary point set is a coordinate sequence describing the outer edge of the connected domain. A boundary tracing algorithm (such as Moore neighborhood tracing) is used to obtain contour points, and an ordered coordinate chain is stored in a clockwise direction to accurately depict the defect geometry.
[0185] Calculate the coordinate span of the boundary point set in the image coordinate system. The coordinate span is the difference between the extreme values of the boundary point set in the image coordinate system. It includes the minimum / maximum horizontal coordinate values (x_min, x_max) and the minimum / maximum vertical coordinate values (y_min, y_max), which are used to determine the bounding rectangle.
[0186] Based on the coordinate span value, a minimum bounding rectangle (MBR) is determined. This MBR is a fitted rectangle. This MBR is an axis-aligned bounding box generated based on the coordinate span value. Its four sides are parallel to the image axes, its upper left corner is at (x_min, y_min), and its lower right corner is at (x_max, y_max), completely covering the object outline.
[0187] For example, we input the heat map generated by GradCAM (size 2400×5000), calculate the maximum heat value V_max of the entire image, set the activation intensity threshold T_act = 0.7×V_max, and perform binarization. Pixels with heat values ≥ T_act are set to 1, and the rest are set to 0.
[0188] Initial cluster generation: Scan the binary image, mark all pixels with a value of 1, use the eight-neighborhood connectivity rule to mark the region, output several independent connected regions (initial activation region clusters), and calculate the geometric center of gravity of each cluster, x c = the average value of the horizontal coordinates of all pixels in the cluster, y c = the average vertical coordinate of all pixels in the cluster, and the centers of the first K clusters with the largest areas are selected as seed points (K=3).
[0189] Perform independent growth on each seed point and add the seed point to the growth queue. The iterative expansion includes taking the current pixel point from the queue, examining the position of its eight neighboring pixels, and calculating the neighboring pixel thermal value H. neighbor and the seed point heat value H seed The absolute difference ΔH, if ΔH≤0.15×V max , the neighborhood pixels are included in the growth area and added to the queue. The termination condition is that the queue is empty or the area of the growth area reaches the preset upper limit.
[0190] For each final growth area, starting from the bottom left pixel, traverse the contour points in a clockwise direction, and calculate the coordinate extreme value, x min is the minimum horizontal coordinate of the contour point, x max is the maximum horizontal coordinate of the contour point, y min is the minimum ordinate of the contour point, y max The maximum vertical coordinate of the contour point. Generate an enclosing rectangle, the upper left corner coordinate (x min ,y min) , the lower right corner coordinate (x max ,y max ), output the fitted rectangle parameter list.
[0191] This embodiment uses an adaptive region growing algorithm to achieve precise positioning under weak supervision conditions. The centroid positioning based on thermal distribution ensures the reliability of the growth starting point. The thermal value difference threshold controls the boundary expansion accuracy. The contour tracking algorithm retains the geometric characteristics of the defect.
[0192] S600: Calculating a fitting score of the RGB image by combining an evaluation result of an overlap between the predicted rectangular box and the fitted rectangular box and the category confidence score.
[0193] In some embodiments, the intersection over union (IoU) of the fitted rectangle and the predicted rectangle is first calculated. This IoU represents the degree of spatial overlap and is calculated by dividing the area of the intersection of the fitted rectangle and the predicted rectangle by the area of their union. This value ranges from 0 to 1, with larger values indicating greater spatial overlap. A higher IoU indicates greater consistency between the approximated rectangle and the model-generated rectangle.
[0194] Obtain the first and second weight coefficients. The first weight coefficient is an adjustment factor for the Intersection-Over-Union (IoU) parameter. This coefficient is a preset constant (default value 0.6) that controls the weight of spatial overlap in the final evaluation. A larger value indicates greater importance for location overlap. The second weight coefficient is an adjustment factor for the category confidence score. This coefficient is a preset constant (default value 0.4), and its sum with the first weight coefficient is 1. Its value reflects the emphasis on the reliability of the model judgment.
[0195] The fitting score is generated by combining the first weight coefficient, the second weight coefficient, the category confidence score, and the intersection over union ratio. In this embodiment, the fitting score AP for discriminating high-confidence samples is designed by combining the category confidence score P and the IoU score, and is calculated by the following formula:
[0196] ;
[0197] in, and It is a weight coefficient, which represents the different contributions of IOU and category confidence score P to the final fitting, and serves as a hyperparameter for adjusting the detection model.
[0198] A higher AP indicates that the box generated by the model after reasoning is more reliable and correctly classified, and can be used to expand the detection training samples.
[0199] S700: Obtaining a target detection model based on the fitting score.
[0200] The object detection model is the final deployed model after iterative optimization. Its architecture is identical to the initial detection model, but the training set is expanded with pseudo-annotated rectangles automatically generated for samples with high fit scores (AP>0.75). The model output includes the spatial location, size, and confidence score of metal loss defects.
[0201] The detection model is retrained using the expanded data and the samples that originally contained accurate rectangular boxes to obtain the second version of the detection model, namely the target detection model, which will have better detection results.
[0202] In some embodiments, RGB images exceeding a fitting score threshold are added to the first training set to generate a second training set; and the detection model is trained using the second training set to obtain a target detection model.
[0203] The fit score threshold is a critical parameter for selecting high-quality pseudo-annotated samples. This threshold is a preset constant (e.g., 0.75). When the calculated fit score for an RGB image exceeds this value, the prediction result for that sample is considered reliable. The threshold setting must strike a balance between recall and precision. The optimal value is typically determined using the validation set ROC curve.
[0204] The second training set is an expanded and optimized dataset. Based on the first training set, it adds RGB images filtered by fitting scores and their corresponding fitted rectangle pseudo-annotations. The object detection model is the final deployed metal loss identification network. It shares the same architecture as the initial detection model (both using YOLOv5), but is reinforced with the second training set to enable the model to learn a more comprehensive representation of defect features.
[0205] S800: Performing metal loss detection on the magnetic leakage signal image to be detected by using the target detection model to obtain a detection result.
[0206] The detection result is the quantitative output of the target detection model processing. The detection result includes the spatial position coordinates of the predicted rectangular box, the target category label, and the category confidence score. For example, the axial position of the defect center point, the circumferential position, the equivalent rectangle size, the defect type code, and the confidence score.
[0207] In some embodiments, a magnetic flux leakage signal image is acquired. This image is the raw data obtained during the actual inspection task. This image is acquired using a magnetic induction sensor onboard the in-pipeline inspection equipment, recording the magnetic flux density distribution on the pipeline wall at a millisecond sampling frequency. The data is stored as a single-channel grayscale matrix, with the row dimension corresponding to axial distance and the column dimension corresponding to circumferential angle.
[0208] The magnetic flux leakage signal image to be measured is converted into an RGB image. RGB image conversion is the process of converting the single-channel data to be measured into a three-channel color image. The RGB image is input into the target detection model to obtain a detection result.
[0209] For example, the pipeline crawler travels at a constant speed of 0.5 m / s, and the magnetic sensor array collects magnetic flux at a frequency of 1 kHz. The data is transmitted in real time to the ground processing system for storage. The raw signal matrix (size M × N) is received and preprocessed by normalization, bilinear interpolation to 2400 × 5000, row-wise 5-pixel window mean filtering, dynamic threshold truncation, and color scale mapping, outputting a standardized RGB image.
[0210] Model prediction execution, loading the target detection model weight file, inputting the preprocessed RGB image, forward propagation process, backbone network extracting multi-scale features, FPN, PANet fusion features, prediction head output prediction results, detection results include coordinate parameters center point (x, y), width and height (w, h), confidence level of metal loss probability, and category fixed label Metal Loss After post-processing, non-maximum suppression (IoU threshold 0.5) and confidence filtering (threshold 0.6) are performed to generate the results.
[0211] The method provided in this embodiment uses only a small number of accurate rectangular box annotations and easily labeled category labels. It obtains approximate rectangular boxes by using a weighted category activation heat map combined with a region growing algorithm. High-confidence sample screening ensures the reliability of iterative detection model training, ultimately resulting in a model with excellent detection results. This reduces the large amount of labor required for traditional fully supervised target detection, and performs input image feature enhancement based on RGB images, which is beneficial for model training. Compared with existing technologies, this method is more suitable for the application of YOLO-based target detection models in the field of pipeline magnetic leakage detection, thereby improving detection accuracy.
[0212] Similar parts between the embodiments provided in this application can be referenced to each other. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods expanded based on the scheme of this application without expending creative work shall fall within the scope of protection of this application.
Claims
1. A pipeline magnetic leakage signal metal loss detection method based on weak annotation, characterized in that: include: Acquire a magnetic flux leakage signal image, and convert the magnetic flux leakage signal image into an RGB image; The RGB images are used as a first training set, and a detection model is trained using a preset number of RGB images with rectangular box annotations in the first training set, so that the detection model performs prediction on all the RGB images to generate predicted rectangular boxes, and the detection model is an improved YOLO model; Training a classification model using the RGB images with category annotations, so that the classification model obtains category confidence scores for all RGB images; generating a heat map based on the classification model; Performing region growing processing on the heat map to generate a fitting rectangular box; Calculating a fitting score for the RGB image based on an evaluation result of an overlap between the predicted rectangular box and the fitted rectangular box and the category confidence score; Based on the fitting score, obtaining a target detection model; The target detection model is used to perform metal loss detection on the magnetic leakage signal image to be detected to obtain a detection result.
2. The pipeline magnetic flux leakage signal metal loss detection method based on weak annotation according to claim 1 is characterized in that: The converting the magnetic leakage signal image into an RGB image includes: Arrange the magnetic flux leakage signal image into preset rows and preset columns to output a first magnetic flux leakage signal image; Based on the preset rows, performing mean filtering and threshold truncation on each row of the first magnetic leakage signal image to output a second magnetic leakage signal image; Performing a color scale transformation on the second magnetic leakage signal image to output an RGB image.
3. The pipeline magnetic flux leakage signal metal loss detection method based on weak annotation according to claim 1 is characterized in that: The method of using the RGB images as a first training set and using a preset number of RGB images with rectangular box annotations in the first training set to train a detection model, so as to enable the detection model to perform prediction on all the RGB images to generate predicted rectangular boxes, includes: Acquire a preset number of RGB images in the first training set, where the preset number of RGB images are RGB images including specific features; Marking the preset number of RGB images with rectangular frames to obtain rectangular frame marked samples; Dividing the rectangular frame labeled samples into a detection training subset and a detection verification subset; Set training parameters; Based on the training parameters, performing iterative training on the detection training subset to obtain a detection model; Input all the RGB images to the detection model, so that the detection model outputs the predicted rectangular box corresponding to each RGB image.
4. The pipeline magnetic flux leakage signal metal loss detection method based on weak annotation according to claim 1 is characterized in that: The classification model includes a feature extraction backbone network, a feature fusion network and a prediction head module; The method of training a classification model using the RGB images with category annotations so that the classification model obtains category confidence scores for all RGB images includes: Generate a multi-scale feature map through the feature extraction backbone network; fusing the multi-scale feature maps in the feature fusion network to obtain a first depth feature map; Processing the first depth feature map by the prediction head module to obtain a first splicing result; performing non-maximum suppression processing on the first splicing result, removing detection frames with a repetition rate higher than a threshold in the first splicing result, to output a second splicing result; Map the size of the detection box in the second stitching result to the size of the RGB image to output a category confidence score for each detection box.
5. The pipeline magnetic flux leakage signal metal loss detection method based on weak annotation according to claim 4 is characterized in that: The feature extraction backbone network includes a convolutional layer; Generating a heat map based on the classification model includes: Calculating a response weight factor of the target category to the feature channel in the convolutional layer, wherein the response weight factor represents the contribution of the feature channel to the target category judgment; Multiplying the characteristic spectrum of the characteristic channel by the response weight factor to obtain a weighted characteristic spectrum; superimposing the weighted feature maps along the feature channels to obtain a superposition result; performing linear rectification activation processing on the superposition result to obtain an activation area; retaining the positive activation area in the activation area as the initial heat map; The spatial resolution of the initial heat map is matched to the size of the RGB image to generate a heat map.
6. The pipeline magnetic flux leakage signal metal loss detection method based on weak annotation according to claim 4 is characterized in that: The prediction head module includes a first prediction head, a second prediction head, a third prediction head and a fourth prediction head; The processing of the first depth feature map by the prediction head module to obtain a first splicing result includes: Processing the first depth feature map by the first prediction head to obtain a second depth feature map; Processing the second depth feature map by the second prediction head to obtain a third depth feature map; Processing the third depth feature map by the third prediction head to obtain a fourth depth feature map; Processing the fourth depth feature map by the fourth prediction head; The first depth feature map, the second depth feature map, the third depth feature map, and the fourth depth feature map are spliced in the feature dimension to output a first splicing result, where the first splicing result includes multiple detection boxes.
7. The pipeline magnetic flux leakage signal metal loss detection method based on weak annotation according to claim 1 is characterized in that: The performing region growing processing on the heat map to generate a fitting rectangular frame includes: Identifying pixels in the heat map that exceed an activation intensity threshold to determine highlighted pixels; Aggregating the spatially adjacent highlighted pixels to form an initial activation area cluster; Calculating the geometric center of gravity of the initial activation area cluster; Determining the geometric center of gravity point in the geometric center of gravity position as the starting seed point of region growth; Performing multi-directional pixel expansion with the starting seed point as the center to obtain neighborhood pixel points; Calculating the feature similarity between the neighborhood pixel points and the starting seed point; If the feature similarity is greater than a similarity threshold, the neighborhood pixel points are merged into the growing region to generate a final growing region, where the final growing region includes a plurality of independent regions; Performing spatial connected domain labeling on the final growth region and extracting an external contour boundary point set of the independent region; Calculating the coordinate span value of the boundary point set in the image coordinate system; Based on the coordinate span value, a minimum bounding rectangle is determined, where the minimum bounding rectangle is a fitted rectangular frame.
8. The pipeline magnetic flux leakage signal metal loss detection method based on weak annotation according to claim 1 is characterized in that: The calculating the fitting score of the RGB image by combining the overlap evaluation result between the predicted rectangular box and the fitted rectangular box and the category confidence score includes: Calculating an intersection-and-union ratio between the fitted rectangular frame and the predicted rectangular frame, where the intersection-and-union ratio represents a degree of spatial position overlap; Obtaining a first weight coefficient and a second weight coefficient; A fitting score is generated by using the first weight coefficient, the second weight coefficient, the category confidence score, and the intersection-over-union ratio.
9. The pipeline magnetic flux leakage signal metal loss detection method based on weak annotation according to claim 1, characterized in that: The obtaining of the target detection model based on the fitting score includes: adding RGB images exceeding a fitting score threshold to the first training set to generate a second training set; The detection model is trained using the second training set to obtain a target detection model.
10. The pipeline magnetic flux leakage signal metal loss detection method based on weak annotation according to claim 1, characterized in that: The performing metal loss detection on the magnetic leakage signal image to be detected by the target detection model to obtain a detection result includes: Acquire the magnetic flux leakage signal image to be measured; Converting the magnetic flux leakage signal image to be measured into an RGB image; The RGB image is input into the target detection model to obtain a detection result, which includes the spatial position coordinates of the predicted rectangular box, the target category label, and the category confidence score.
Citation Information
Patent Citations
Infrared small target detection and classification method based on multi-scale feature fusion
CN118038152A
Automatic optical detection method and system based on gasket surface defects
CN119643448A