A method, device, equipment and medium for identifying defects in overhead pipelines
By enhancing and optimizing the overhead pipeline defect images, a deep learning model for overhead pipeline defect recognition is built, which solves the problem of inefficient manual inspection in the existing technology, and realizes efficient and accurate pipeline defect recognition to adapt to the rapidly growing detection needs.
Patent Information
- Application Number
- CN202411688576.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-25
AI Technical Summary
In the prior art, overhead pipeline defect detection relies on manual inspection, which is inefficient, has safety hazards, and is difficult to adapt to the growing number of pipelines, resulting in increased detection pressure.
A method for identifying overhead pipeline defects is adopted, including the model training stage and the model usage stage. In the model training stage, data augmentation of the original defect image, classification branches and representation learning branch sub-models are constructed, and the overall loss function is optimized to obtain the defect recognition model. During the model usage phase, the image to be tested is input into the model to generate defect recognition results.
It significantly improves the recognition ability of pipeline defects under long tail distribution, improves the recognition accuracy of balanced or unbalanced pipeline defect images, and the model speed is faster and can more efficiently adapt to the growing detection needs.
Smart Images

Figure CN119625395B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of defect identification, and in particular to a method, device, equipment and medium for identifying defects in overhead pipelines. Background Art
[0002] As a key part of modern urban infrastructure, overhead pipelines are responsible for urban water supply, gas supply and transportation of various important materials. However, with the passage of time and the continuous influence of the external environment, overhead pipelines have gradually exposed a series of safety hazards such as aging, corrosion, and mechanical damage. If these safety hazards are not discovered and properly handled in time, it will not only lead to the degradation of pipeline performance, but also may trigger serious safety accidents, causing adverse effects on urban operations and residents' lives. Therefore, it is particularly important to strengthen the maintenance and management of overhead pipelines to ensure their safe and stable operation.
[0003] In the past, manual inspection was often used as the main method for detecting defects in overhead pipelines. Inspectors evaluate pipelines by visual observation, handheld inspection equipment such as telescopes, and using climbing equipment such as climbing ropes to approach overhead pipelines. However, manual inspection is not only inefficient and has safety hazards, but is also easily affected by subjective factors, resulting in missed inspections or misjudgments. In addition, with the acceleration of urbanization, the number of overhead pipelines has increased dramatically, and the pressure on manual inspections has also increased. Therefore, there is an urgent need for an efficient and accurate detection method to meet the growing demand for overhead pipeline inspections. Summary of the invention
[0004] In view of this, the purpose of the present application is to provide a method, device, equipment and medium for identifying defects in overhead pipelines to overcome the problems in the prior art.
[0005] In a first aspect, an embodiment of the present application provides a method for identifying defects in an overhead pipeline, the method comprising: a model training phase and a model use phase;
[0006] The model training phase includes:
[0007] Performing data enhancement on the original defect images in the data set to obtain multiple groups of enhanced defect images; wherein the enhanced defect images include a first partial image and a second partial image;
[0008] Using the first feature vector of the first partial image to construct a first loss function of the classification branch sub-model;
[0009] Using the second feature vector of the second partial image to construct a second loss function representing the learning branch sub-model; wherein the first loss function and the second loss function are used to construct an overall loss function;
[0010] By adjusting the parameters of the overall loss function, a defect recognition model is obtained;
[0011] The model usage stage includes:
[0012] Input the test image of the overhead pipeline to be tested into the defect recognition model to obtain the defect recognition result output by the defect recognition model.
[0013] In some technical solutions of this application, the above-mentioned original defect images in the dataset are subjected to data augmentation to obtain multiple groups of enhanced defect images; wherein, the enhanced defect images include a first part of the image and a second part of the image, including:
[0014] Use the first preprocessing method to perform data augmentation on the original defect images in the dataset to obtain the first part of the image;
[0015] Use the second preprocessing method and the third preprocessing method to perform data augmentation on the original defect images in the dataset respectively to obtain the second part of the image.
[0016] In some technical solutions of this application, the above-mentioned defect recognition model includes a backbone network; wherein, the backbone network is used to extract basic features to obtain the first basic feature map of the first part of the image and extract the second basic feature map of the second part of the image.
[0017] In some technical solutions of this application, the above-mentioned classification branch sub-model includes: an adaptive perception module, a global attention pooling module, and a classification layer;
[0018] The adaptive perception module is used to dynamically adjust the first basic feature map of the first part of the image, and fuse the adjusted features with the original features through a residual connection to obtain a fused first adjusted feature map;
[0019] The global attention pooling module is used to perform a pooling operation on the fused first adjusted feature map, dynamically assign weights to different feature regions, and thus generate the first feature vector of the first part of the image;
[0020] The classification layer is used to map the first feature vector to obtain a first mapping vector, and use the first mapping vector to construct the first loss function of the classification branch sub-model.
[0021] In some technical solutions of this application, the above-mentioned adaptive perception module is used to dynamically adjust the first basic feature map of the first part of the image, and fuse the adjusted features with the original features through a residual connection to obtain a fused first adjusted feature map, including:
[0022] Use the first volume integration branch and the second volume integration branch to perform convolution operations on the first basic feature map respectively to obtain a first convolution map and a second convolution map; wherein, the first convolution map and the second convolution map are used to be fused to obtain a first fusion map;
[0023] Perform a pooling operation and a compression operation on the first fusion map in sequence to obtain a fourth feature vector;
[0024] Process the fourth feature vector through an attention mechanism to obtain a first weight of the first volume integration branch and a second weight of the second volume integration branch;
[0025] Fuse the first convolution map and the second convolution map according to the first weight and the second weight to obtain the preliminary adjusted feature map.
[0026] Perform element-wise addition fusion on the preliminary adjusted feature map and the first basic feature map through a residual connection to obtain the finally output first adjusted feature map.
[0027] In some technical solutions of the present application, the above-mentioned second part of the image includes a first sub-image and a second sub-image;
[0028] The representation learning branch sub-model includes: a cross-space feature fusion module, a multi-layer perceptron, and a clustering module;
[0029] The cross-space feature fusion module is used to perform enrichment processing on the second basic feature map of the first sub-image and the third basic feature map of the second sub-image to obtain a second feature vector of the first sub-image and a third feature vector of the second sub-image;
[0030] The multi-layer perceptron is used to map the second feature vector and the third feature vector to obtain a second mapping vector of the second feature vector and a third mapping vector of the third feature vector;
[0031] The clustering module is used to construct a second loss function of the representation learning branch sub-model based on the vMF mixture distribution for the second mapping vector and the third mapping vector.
[0032] In some technical solutions of the present application, the above-mentioned cross-space feature fusion module enriches the second basic feature map of the first sub-image in the following manner to obtain the second feature vector of the first sub-image:
[0033] Use multiple branches to process the second basic feature map to obtain processing results of multiple branches;
[0034] Fuse the processing results of multiple branches to obtain the second feature vector of the first sub-image.
[0035] In a second aspect, an apparatus for identifying defects in overhead pipelines provided by an embodiment of the present application includes:
[0036] a model training module and a model using module;
[0037] The model training module is used for:
[0038] performing data augmentation on the original defect images in the dataset to obtain multiple groups of augmented defect images; wherein, the augmented defect images include a first part image and a second part image;
[0039] constructing a first loss function of the classification branch sub-model using the first feature vector of the first part image;
[0040] constructing a second loss function of the representation learning branch sub-model using the second feature vector of the second part image; wherein, the first loss function and the second loss function are used to construct an overall loss function;
[0041] obtaining a defect recognition model by adjusting the parameters of the overall loss function;
[0042] The model using module is used for:
[0043] inputting the test image of the overhead pipeline to be tested into the defect recognition model to obtain the defect recognition result output by the defect recognition model.
[0044] In a third aspect, an electronic device provided by an embodiment of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for identifying defects in overhead pipelines are implemented.
[0045] In a fourth aspect, a computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is run by a processor, the steps of the above-mentioned method for identifying defects in overhead pipelines are executed.
[0046] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0047] The method of this application includes that the model training stage includes: performing data augmentation on the original defect images in the dataset to obtain multiple groups of augmented defect images; wherein, the augmented defect images include a first part of the image and a second part of the image; constructing a first loss function of the classification branch sub-model using the first feature vector of the first part of the image; constructing a second loss function of the representation learning branch sub-model using the second feature vector of the second part of the image; wherein, the first loss function and the second loss function are used to construct an overall loss function; adjusting the parameters of the overall loss function to obtain a defect recognition model; the model usage stage includes: inputting the test image of the overhead pipeline to be tested into the defect recognition model to obtain the defect recognition result output by the defect recognition model.
[0048] By optimizing the pipeline defect recognition classification module, the method of this application significantly improves the recognition ability of pipeline defects under the long-tail distribution, has high recognition accuracy for both balanced and unbalanced pipeline defect images, and the model is faster.
[0049] To make the above objects, features, and advantages of this application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of this application, so they should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.
[0051] Figure 1 Shows a schematic flowchart of a method for identifying overhead pipeline defects provided by an embodiment of this application;
[0052] Figure 2 Shows a schematic overall framework diagram provided by an embodiment of this application;
[0053] Figure 3 Shows a schematic diagram of a branch network provided by an embodiment of this application;
[0054] Figure 4 Shows a schematic diagram of an adaptive perception module provided by an embodiment of this application;
[0055] Figure 5 Shows a schematic diagram of a cross-space feature fusion module provided by an embodiment of this application;
[0056] Figure 6 Shows a schematic diagram of a device for identifying overhead pipeline defects provided by an embodiment of this application;
[0057] Figure 7 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific implementation manners
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and the steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0059] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0060] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the subsequently stated features, but does not exclude adding other features.
[0061] As a key part of modern urban infrastructure, overhead pipelines shoulder the important responsibilities of urban water supply, gas supply, and transportation of various important materials. However, over time and under the continuous influence of the external environment, a series of potential safety hazards such as aging, corrosion, and mechanical damage have gradually emerged in overhead pipelines. If these potential safety hazards cannot be discovered and properly handled in a timely manner, it will not only lead to the degradation of pipeline performance, but may also trigger serious safety accidents, having an adverse impact on urban operation and residents' lives.
[0062] The formation of defects in overhead pipelines can be affected by the combined action of multiple factors, such as long-term ultraviolet radiation, continuous impact of harsh climate, erosion by chemical substances, and accumulation of mechanical stress, etc. Under the long-term action of these factors, the performance of the pipeline gradually degrades, and the material is damaged, thus increasing the risk of defects. For example, pipeline corrosion is one of the most common defects in overhead pipelines. Especially in a humid environment or under the action of chemical corrosive media, oxidation and corrosion are likely to occur on the surface of metal pipelines, which not only seriously weakens the structural strength of the pipeline but also threatens its stability. In addition, overhead pipelines may also encounter mechanical impacts or external damages, such as accidental touches during building construction, traffic accidents, wind disasters and other unexpected events. These unforeseen events may cause the pipeline to deform, rupture or even suffer more serious structural damages, thereby affecting the normal operation and safety of the pipeline. In view of this, timely and accurate identification of defects in overhead pipelines and taking corresponding repair and maintenance measures are of crucial significance for ensuring the stable operation of the urban supply system and the safety of residents' lives.
[0063] In previous defect detections of overhead pipelines, manual inspection was often used as the main method. Inspectors evaluated the pipelines by visual observation, using handheld detection equipment such as telescopes, and approaching the overhead pipelines with climbing equipment such as climbing ropes. However, manual detection is not only inefficient and poses safety hazards but is also easily affected by subjective factors, resulting in missed detections or misjudgments. In addition, with the acceleration of urbanization, the number of overhead pipelines has increased sharply, and the pressure of manual inspection has also increased accordingly. Therefore, there is an urgent need for an efficient and accurate detection method to meet the growing detection requirements of overhead pipelines.
[0064] In recent years, with the development of computer vision and deep learning technologies, intelligent detection means have gradually been introduced into the defect identification of overhead pipelines. Devices such as drones and high-definition cameras are used for real-time monitoring to automatically collect data, and the collected image data is transmitted to a remote server with strong data processing capabilities for defect identification and processing in combination with deep learning algorithms. In the urban overhead pipeline image dataset, some defect types may appear with a higher frequency, while some relatively rare defect types may show a long-tailed distribution. However, most deep learning algorithms are based on relatively balanced datasets. For real data, the data volume of different categories usually does not have an ideal uniform distribution but an imbalanced data distribution. For data with a long-tailed distribution, it is equally important to pay attention to those rare defect types. Rare defects may represent potential safety hazards and require special attention and identification. The more the number of low-frequency classes, the greater the gap between the number of samples and the number of samples in high-frequency classes, and the lower the defect identification accuracy, which poses a challenge to traditional supervised learning algorithms.
[0065] Based on this, the embodiments of the present application provide a method, device, equipment and medium for identifying overhead pipeline defects, which will be described below through embodiments.
[0066] Figure 1 The flowchart of a method for identifying overhead pipeline defects provided by the embodiments of the present application is shown. Among them, the method includes step S101, the model training stage, and S102, the model usage stage; specifically:
[0067] S101. The model training stage includes:
[0068] Performing data augmentation on the original defect images in the dataset to obtain multiple groups of enhanced defect images; among them, the enhanced defect images include a first part of the image and a second part of the image;
[0069] Using the first feature vector of the first part of the image to construct the first loss function of the classification branch sub-model;
[0070] Using the second feature vector of the second part of the image to construct the second loss function of the representation learning branch sub-model; among them, the first loss function and the second loss function are used to construct the overall loss function;
[0071] By adjusting the parameters of the overall loss function, a defect recognition model is obtained;
[0072] S102. The model usage stage includes:
[0073] Inputting the test image of the overhead pipeline to be tested into the defect recognition model to obtain the defect recognition result output by the defect recognition model.
[0074] The method of the present application significantly improves the recognition ability of pipeline defects under the long-tail distribution by optimizing the pipeline defect recognition classification module, has high recognition accuracy for both balanced and unbalanced pipeline defect images, and the model is faster.
[0075] Some embodiments of the present application will be described in detail below. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0076] In the model training stage, the embodiments of the present application need to first determine the dataset for model training. Specifically, as Figure 2 shown, collect image data and construct a dataset of urban overhead pipeline defects: use a drone equipped with a high-definition camera to collect actual images of the pipeline, ensuring that various defect types that may occur in the pipeline and different environmental conditions are covered. Manually screen the original images that can truly reflect the pipeline state, especially the images containing potential defects.
[0077] Data augmentation is performed on the original defect images in the dataset to obtain multiple groups of augmented defect images; the collected images are preprocessed, that is, data augmentation is performed. The first preprocessing method performs conventional augmentation on the images: random cropping, random horizontal flipping, and normalization; the second preprocessing method performs enhanced augmentation on the images: randomly cropping to a specified size, random horizontal flipping, randomly changing the brightness, contrast, saturation, and hue of the images, converting the images to grayscale with a probability of 0.2, Gaussian blurring, and normalization.
[0078] Specifically, the first preprocessing method is used to perform data augmentation on the original defect images in the dataset to obtain the first part of the images; the second preprocessing method and the third preprocessing method are used to perform data augmentation on the original defect images in the dataset respectively to obtain the second part of the images.
[0079] For example, for the same defect image, three different image preprocessing operations are applied to obtain the first part of the images and the second part of the images (augmented images x 1 、x 2 and x 3 ). Among them, the first part of the images is the augmented image x 1 , which is used for the classification task, and the second part of the images is the augmented images x 2 and x 3 , which are used for the contrastive learning task.
[0080] As Figure 3 shown, after data augmentation, a dual-branch learning network based on a linear classifier's classification branch and a projection head's representation branch is constructed, and the network is improved in combination with the overhead pipeline defect features. For the improved defect recognition dual-branch learning model, the model includes: a backbone network, a classification branch, and a representation learning branch. The backbone network is used to extract basic features to obtain the first basic feature map of the first part of the images and extract the second basic feature map of the second part of the images.
[0081] In the classification branch, the network first extracts basic features, and then dynamically adjusts the basic features through an adaptive perception module while retaining the original feature information, and fuses them in a residual connection manner. This module can dynamically adapt to the optimal perception scale, refine feature extraction while maintaining the supplement and balance of the original global information. Subsequently, through the global attention pooling module, the network can dynamically assign different weights to different feature regions, thereby further enhancing the feature expression of key regions and generating the class feature vector f 1 of the image x 1This design effectively alleviates possible information loss in deep feature extraction through residual connections, enabling the network to more precisely focus on fine-grained features for head classes while expanding the receptive field for tail classes to capture more extensive context information, thus significantly improving the classification performance. Subsequently, the one-dimensional feature vector f 1 is passed to the fully-connected classification layer. During this process, the refined features are mapped to a new feature space to generate the logits s required for classification, where s represents the confidence prediction for each class and is a key step in the classification process, directly affecting the generation of the final classification result. After obtaining the logits, the softmax function is applied to normalize them into a probability distribution. Finally, the probabilities obtained from the softmax function are used to calculate the cross-entropy loss L CE :
[0082]
[0083] where y i is the true label and v i is the predicted probability. This loss function measures the deviation between the predicted probability distribution and the actual label distribution. By minimizing this loss, the weights of the fully-connected classification layer can be further adjusted to reduce the classification error.
[0084] The improved representation learning branch can achieve multi-scale perception and cross-space learning. By fusing spatial features of different scales and regions, it promotes feature complementarity, thus effectively extracting transferable and strongly discriminative features suitable for long-tail image tasks. First, the backbone network extracts the basic feature vectors f′ 2 and f′ 3 of the images x 2 and x 3 . Subsequently, these features are further processed by the cross-space feature fusion module to generate a richer multi-scale class feature representation, denoted as f 2 and f 3 . Next, a multi-layer perceptron MLP with hidden layers is used to map the feature representations f 2 , f 3 to another feature space to obtain the representations z 2 , z 3 for contrastive learning. The MLP consists of multiple fully-connected networks and can learn more abstract and high-order feature representations. All the representations used for contrastive learning here are normalized using the L 2 norm to ensure that the feature space is a unit hypersphere.
[0085] To address the class imbalance problem caused by the long-tail distribution, a von Mises-Fisher (vMF) mixture distribution is introduced to model the normalized feature vectors in contrastive learning. The vMF distribution can enhance the compactness of feature clusters and alleviate the class imbalance in long-tail classification through a uniform distribution in the hypersphere space. Based on this, a new loss function \(L\) corresponding to the representation learning branch is derived. R This loss function reduces the dependence on large-scale sample sampling by generating sufficient contrast pairs and improves the overall performance of the model in long-tail data.
[0086] In the whole framework, the representation learning branch and the classification branch cooperate closely, aiming to minimize the overall loss function through the following coordinated training process:
[0087] \(L = \alpha L_{r}+\beta L_{c}\) R +\(\beta L_{c}\) CE
[0088] where \(L\) represents the overall loss function of the model, which is a weighted combination of the classification branch loss \(L_{c}\) CE and the representation learning branch loss \(L_{r}\), R ensuring the dual optimization of classification performance and feature learning. During the optimization process, the two branches are synchronously trained using the stochastic gradient descent algorithm. With this design, the recognition model we designed demonstrates stronger generalization ability and higher classification accuracy under long-tail pipeline defect image data.
[0089] Specifically, the classification branch sub-model includes: an adaptive perception module, a global attention pooling module, and a classification layer;
[0090] The adaptive perception module is used to dynamically adjust the first basic feature map of the first part of the image and fuse the adjusted features with the original features through a residual connection to obtain a fused first adjusted feature map;
[0091] The global attention pooling module is used to perform a pooling operation on the fused first adjusted feature map, dynamically assign weights to different feature regions, and thus generate a first feature vector of the first part of the image; the classification layer is used to map the first feature vector to obtain a first mapping vector and construct a first loss function of the classification branch sub-model using the first mapping vector.
[0092] The adaptive perception module is used to dynamically adjust the first basic feature map of the first part of the image and fuse the adjusted features with the original features through a residual connection to obtain a fused first adjusted feature map, including:
[0093] Use the first volume integration branch and the second volume integration branch to perform convolution operations on the first basic feature map respectively to obtain a first convolution map and a second convolution map; wherein, the first convolution map and the second convolution map are used to be fused to obtain a first fusion map;
[0094] Perform a pooling operation and a compression operation on the first fusion map in sequence to obtain a fourth feature vector;
[0095] Process the fourth feature vector through an attention mechanism to obtain a first weight of the first volume integration branch and a second weight of the second volume integration branch;
[0096] Fuse the first convolution map and the second convolution map according to the first weight and the second weight to obtain the preliminary adjusted feature map.
[0097] Perform element-wise addition fusion on the preliminary adjusted feature map and the first basic feature map through a residual connection to obtain the finally output first adjusted feature map.
[0098] Specifically as Figure 4 shown, in a traditional convolutional neural network, the receptive field of a certain layer is fixed and cannot be dynamically adjusted according to the input, which limits the ability to capture multi-scale information. Here, it is designed that through an adaptive perception module, the network can provide multiple receptive field selections for neurons and adjust the optimal output through a dynamic weight mechanism to capture the multi-scale features of the target object. The design of the residual connection further ensures the robustness and information integrity of the features, thereby enhancing the representation ability of the final feature map.
[0099] Assume that the first basic feature map of the input first part of the image is X∈R. The first basic feature map is first fed into two branches (the first volume integration branch and the second volume integration branch): the first volume integration branch applies a 3×3 convolutional kernel, and the second volume integration branch applies a 5×5 convolutional kernel (to improve efficiency, the 5×5 convolution is replaced by a dilated convolution with a dilation rate of 2, which can both maintain a large receptive field and reduce the amount of calculation). Each branch generates its own output feature map (the first convolution map and the second convolution map), which are respectively denoted as the first convolution map (output of 3×3 convolution) and the second convolution map (output of 5×5 convolution).
[0100]
[0101] wherein, F 3 and F 5 respectively represent the 3×3 and 5×5 convolution operations.
[0102] Fuse the first convolution map and the second convolution map through element-wise addition to obtain a first fusion map U:
[0103]
[0104] Then, a pooling operation is performed on the first fusion graph through global average pooling to generate a global channel-level statistical vector s, which summarizes the global information of each channel.
[0105]
[0106] To further compress the features and reduce the computational amount, a fully connected layer is used to perform dimensionality reduction on the global information, generating a compact fourth feature vector z for guiding the subsequent selection process.
[0107] z = σ(Ws)
[0108] where σ is the activation function and W is the weight matrix of the fully connected layer.
[0109] This step fuses the information from different convolutional kernel branches to generate global information to assist in the selection of subsequent branches, ensuring that information at different scales can be fused and provide a basis for the next selection.
[0110] The weights a and b for each branch are generated using the vector z through the softmax attention mechanism, corresponding to the 3×3 and 5×5 convolutional kernel branches respectively:
[0111]
[0112] where A and B are weight matrices, ensuring a + b = 1.
[0113] According to the generated weights a and b, weighted summation is performed on the output feature maps of the two branches to generate a preliminary adjusted feature map V:
[0114]
[0115] Finally, through residual connection, element-wise addition fusion is performed between the input first basic feature map X and the adjusted feature map V to obtain the final first adjusted feature map:
[0116] V final = V + X
[0117] This module provides multiple receptive field selections (such as 3×3 and 5×5 dilated convolutional kernels), dynamically selects the optimal feature representation, and combines residual connection to retain the original feature information, further enhancing the feature refinement and expression ability. The representation learning branch sub-model includes: a cross-space feature fusion module, a multi-layer perceptron, and a clustering module;
[0118] The cross - space feature fusion module is used to enrich the second basic feature map of the first sub - image and the third basic feature map of the second sub - image, obtaining the second feature vector of the first sub - image and the third feature vector of the second sub - image;
[0119] The multi - layer perceptron is used to map the second feature vector and the third feature vector to obtain the second mapped vector of the second feature vector and the third mapped vector of the third feature vector;
[0120] The clustering module is used to construct the second loss function representing the learning branch sub - model for the second mapped vector and the third mapped vector based on the vMF mixture distribution.
[0121] The cross - space feature fusion module enriches the second basic feature map of the first sub - image in the following way to obtain the second feature vector of the first sub - image:
[0122] Use multiple branches to process the second basic feature map to obtain the processing results of multiple branches;
[0123] By fusing the processing results of multiple branches, the second feature vector of the first sub - image is obtained.
[0124] Specifically, as Figure 5 shown, the core of this module is the design of parallel sub - networks. First, the global feature extraction sub - network extracts global dependency information from the horizontal and vertical directions respectively through average pooling (AvgPool) operations in the X - direction and Y - direction.
[0125] The formula for the two - dimensional global pooling operation is as follows:
[0126]
[0127] Therefore, the output at height h of the c - th channel can be expressed as:
[0128]
[0129] Similarly, the output of the c - th channel with width W can be written as:
[0130]
[0131] After pooling, these two one - dimensional feature vectors are concatenated along the channel dimension to produce a pair of direction - aware feature maps.
[0132] Immediately afterwards, a 1×1 convolution operation is used to reduce the dimension of these concatenated features, which not only reduces the number of channels but also promotes the in - depth fusion of spatial information.
[0133] Subsequently, the generated attention map is normalized using the sigmoid function (denoted as σ), compressing the values between 0 and 1, thereby assigning different weights to the features of each channel. This process significantly enhances the multi-scale perception ability of the model, ensuring that the model can evenly focus on the feature distribution in different directions and providing a basis for subsequent feature reweighting.
[0134] Among them, represents element-wise multiplication. By applying the attention map to the original input features, weighted features are generated, thus achieving cross-space feature reweighting, which helps the model to focus more on the information-dense areas in different spatial regions of the image.
[0135] Meanwhile, the local feature enhancer sub-network captures fine-grained spatial detail features through 3×3 convolutions, enhancing the network's ability to express local features.
[0136] By performing multi-scale pooling and softmax operations on the weighted features r 1 and r 3 of the two branches, two corresponding spatial attention weights W 1 and W 3 are generated to capture multi-scale features. The multi-scale attention mechanism ensures that the model evenly focuses on feature information at different spatial scales, further enriching the hierarchical structure of the features.
[0137] Furthermore, through matrix multiplication, the multi-scale weights W 1 and W 3 are respectively interacted with the feature matrices r 3 and r 1 of the other branch to generate cross-space dimension attention maps M. These attention maps not only encapsulate the local information of each pixel point but also reveal the global relationship between the pixel point and other spatial positions.
[0138] Subsequently, the outputs of the two weight matrices are fused through element-wise addition ⊕ and normalized using the Sigmoid activation function to generate the final weight map. This weight map is then combined with the original input feature map r through element-wise multiplication to achieve feature reweighting. Then, the reweighted feature map is compressed into a compact one-dimensional feature vector through global average pooling (GAP). This step ensures that the model can highlight key features and suppress irrelevant information, generating fused features with rich spatial expression ability, thereby enhancing the spatial information expression effect of the model.
[0139] In an alternative embodiment, the embodiments of the present application further introduce the vMF distribution and derive the loss function of the representation branch. The mixture modeling of the normalized feature vectors with the vMF distribution: There are two reasons for using the von Mises-Fisher (vMF) distribution to model the normalized vectors.
[0140] The representation branch designed in the embodiments of the present application learns features by comparing the similarities and differences between different samples, aiming to map similar samples into a close representation space and dissimilar samples into a distant representation space. On the one hand, in contrastive learning, the normalized vectors are usually restricted to the unit hypersphere (i.e., their norms are 1). To effectively model the distribution of these vectors, the embodiments of the present application design to use a mixture of multiple von Mises-Fisher (vMF) distributions to model the normalized feature vectors in contrastive learning. The vMF distribution is a probability distribution commonly used to model data on the unit hypersphere and can be regarded as a generalization of the normal distribution on the unit hypersphere, capable of capturing the concentration and direction of the normalized features. After the features are normalized, assuming that these features follow the von Mises-Fisher distribution, and by mixing different distributions, this helps the model learn more complex and effective feature representations, making similar samples closer in the feature space. On the other hand, the parameters of the vMF distribution can be efficiently estimated by maximum likelihood estimation in different mini-batches, which is suitable for dealing with the problem of data imbalance. In this way, the expected contrast loss function can be effectively calculated, avoiding the computational overhead brought by directly sampling a large number of contrast pairs.
[0141] By modeling the vMF mixture distribution of the feature vectors, the loss function of the representation branch is derived: In the derivation process, it is first assumed that the normalized feature vectors in contrastive learning follow the vMF mixture distribution, and then the parameters of this distribution are obtained by maximum likelihood estimation. Based on these parameters, a closed-form solution of the expected loss is derived, replacing the loss obtained by sampling in traditional contrastive learning. The key to this reasoning process lies in introducing the vMF distribution and utilizing its good statistical properties, allowing the model to optimize the contrast loss without explicit sampling, avoiding the computational overhead of relying on large batch sampling in traditional contrastive learning. The specific derivation process is as follows:
[0142] Definition of the vMF distribution: In contrastive learning, after the features are normalized, they are restricted to the unit hypersphere, so the conventional normal distribution cannot be directly used for modeling. To adapt to this situation, it is assumed that these features follow the von Mises-Fisher (vMF) distribution.
[0143] The probability density function of the vMF distribution is:
[0144]
[0145] where: z is a unit vector representing the eigenvector, satisfying ||z|| = 1; μ is the mean direction, satisfying ||μ|| = 1; k ≥ 0 is the concentration parameter, representing the concentration degree of the feature in the mean direction; C p (k) is the normalization constant, given in the following form:
[0146]
[0147] where is the first kind of modified Bessel function, used to ensure that the integral of the probability density function is 1, defined in the following form:
[0148]
[0149] Feature distribution assumption: In the representation learning branch, it is assumed that the feature distribution of each class can be modeled by the vMF distribution, and the overall feature distribution in the dataset is a mixture of multiple vMF distributions. Then the total probability distribution of the features can be expressed as:
[0150]
[0151] where, is the prior probability of class y i ; and are the mean direction and concentration parameter of class y i respectively, is the modified Bessel function. Through maximum likelihood estimation, these parameters can be dynamically updated during the training process.
[0152] Estimation of vMF distribution parameters: To derive the new loss function, first, it is necessary to estimate the parameters of the vMF distribution of each class, that is, the mean direction and the concentration parameter Suppose we have a set of N independent unit vector samples {z i}, and these samples come from the vMF distribution of class y i For class y i , we can obtain the estimated values of and through maximum likelihood estimation.
[0153] According to maximum likelihood estimation, the mean direction is:
[0154]
[0155] where is the sample mean, is the norm (length) of the sample mean.
[0156] Lumped parameters Approximately equal to:
[0157]
[0158] Where and are the modified Bessel functions.
[0159] Denote the derivation of the branch loss function: The derivation of the loss function, the key idea is to construct infinitely many contrast pairs through the parameter estimation of the vMF distribution and derive a closed - form expected contrast loss function.
[0160] According to the definition of the supervised contrast loss, we have:
[0161]
[0162] Where denotes the number of samples in the batch that have the same class y i as the sample x i , that is, the number of positive samples of class y i . z i and z p are the normalized feature vectors of the sample x i and the positive sample x p ; z a denotes the normalized feature vector of the sample a in class j; τ is the temperature parameter; A(y i ) is the set of positive samples that have the same class y i as the sample x i . A(j) denotes the set of samples with label j in the batch. K is the total number of classes in the batch.
[0163] The above formula can be further processed to obtain:
[0164]
[0165] When the number of samples N→∞, at this time denotes the sampling frequency of class j. In this case, the constant term logN can be ignored, and the following simplified loss formula can be obtained:
[0166]
[0167] According to the expectation formula of the vMF distribution and the properties of the moment - generating function:
[0168]
[0169] Finally, the closed - form expression of the contrast learning branch loss is obtained:
[0170]
[0171] wherein: is the corrected lumped parameter; π j is the frequency of class j; is the prior probability of class y i ; C p (k) is the corrected Bayesian function related to the vMF distribution.
[0172] The loss L of the contrastive learning branch R has the core advantage of generating sufficient positive and negative sample pairs by estimating the feature distribution, which is particularly suitable for learning long-tail classes. With the vMF mixture distribution model, L R can generate an infinite number of sample pairs at the expected level, ensuring that each class is treated fairly, reducing the dependence on large-scale sample sampling, alleviating the problem of scarce tail-class samples, and improving the learning effect of tail classes. In addition, L R significantly improves the training speed and efficiency through closed-form solution optimization.
[0173] Read the test image using the trained and improved defect recognition model as the test network, and output the result of classifying the pipeline defect recognition: Divide a part of the data collected from the drone as the validation set to ensure that the validation set covers different defect types and environmental scenarios. The data of the validation set follows the same preprocessing process as the data of the training set to ensure data consistency. Input all samples of the validation set into the trained model, perform forward pass one by one, and obtain the prediction result of each sample. Use the prediction result and the true label output by the model to calculate various metrics on the validation set, such as accuracy, precision, recall, F1 score, etc., to evaluate the classification ability of the model. At the same time, a confusion matrix can be used to analyze the recognition effect of the model on different defect types.
[0174] Organize the recognition results into a report, including defect types, confidence scores, and recommended maintenance measures, etc., to help the maintenance team make efficient decisions.
[0175] After the defect recognition model is trained, input the test image of the overhead pipeline to be measured into the defect recognition model to obtain the defect recognition result output by the defect recognition model.
[0176] Different from the prior art, traditional single classification models often result in poor recognition ability for long-tail categories when dealing with class imbalance. In this application, through a dual-branch structure, on the one hand, a classification branch is used for defect recognition classification, and on the other hand, the feature vector representation is further optimized through a representation learning branch. At the same time, different improvements are made to the classification branch and the representation learning branch respectively according to the pipeline defect characteristics. Here, in particular, deep characterization is performed on minority class samples, which enables the classification model to accurately identify uncommon pipeline defect categories even under long-tail distributions. In addition, a vMF mixture distribution model is introduced to optimize the contrastive learning loss, ensuring that similar samples are closer in the feature space, so that the model can maintain a high recognition accuracy in different scenarios.
[0177] In traditional methods, the model often captures features through layer-by-layer convolution operations. In particular, capturing global and local features often requires a large amount of computing resources. In this application, through cross-space feature fusion, a multi-branch convolutional network is designed (for example, 1×1 convolution is used to retain global information, and 3×3 convolution is used to capture local features), enabling the model to process spatial information at different scales in parallel, reducing the computing time of the traditional model's layer-by-layer operations, improving the overall processing speed, and being suitable for real-time applications. Cross-space interaction further extracts long-range dependencies through matrix multiplication and other methods, realizing efficient interaction between spatial dimensions, thereby reducing computational complexity, making the overall model's processing speed faster, and being suitable for real-time applications.
[0178] Traditional convolutional neural networks often adopt a fixed receptive field in design, that is, the receptive field size of neurons in each layer is static and unchanged in the network. This method lacks flexibility in dealing with multi-scale features and may lead to the inability to simultaneously consider the feature extraction of small-scale and large-scale targets. This application introduces an adaptive perception mechanism, selects convolution operations of different scales (such as 3×3 and 5×5 dilated convolutions) through multi-convolution kernel branches, and combines a soft attention mechanism to adaptively generate the weight distribution of each convolution kernel according to the spatial scale of the input feature map, thereby dynamically selecting the optimal convolution kernel for feature extraction. In addition, the dynamically adjusted features are fused with the original features through residual connections, which not only enhances the flexibility of multi-scale feature extraction but also retains the original global information, improving the integrity of feature expression. While ensuring computational efficiency, this mechanism effectively expands the model's perception ability for multi-scale features, enabling the model to more accurately capture different scale information of the target when dealing with complex image scenarios, thus significantly enhancing the model's expression ability and adaptability.
[0179] Traditional contrastive learning usually requires sampling a large number of samples to ensure that the model learns the relative relationships of different classes in the feature space. This approach has high requirements for the data scale. In this application, by assuming that the normalized feature vectors follow a von Mises-Fisher (vMF) mixture distribution and using the maximum likelihood estimation method to model these feature vectors, the feature representation in contrastive learning is optimized, and the form of the expected loss function that does not require a large number of samples is derived. This modeling method reduces the dependence on large-scale data sets, enabling the model to maintain the effectiveness of contrastive learning even under small-batch data training, especially performing well under long-tailed data distributions. This improvement makes the model more flexible in practical applications and adaptable to scenarios with scarce or unbalanced data, especially suitable for dealing with long-tailed data distributions.
[0180] Figure 6 The structural schematic diagram of a device for identifying overhead pipeline defects provided by an embodiment of this application is shown. The device includes: a model training module and a model using module;
[0181] The model training module is used for:
[0182] Performing data augmentation on the original defect images in the data set to obtain multiple groups of augmented defect images; wherein, the augmented defect images include a first part of the image and a second part of the image;
[0183] Using the first feature vector of the first part of the image to construct the first loss function of the classification branch sub-model;
[0184] Using the second feature vector of the second part of the image to construct the second loss function of the representation learning branch sub-model; wherein, the first loss function and the second loss function are used to construct the overall loss function;
[0185] Adjusting the parameters of the overall loss function to obtain a defect recognition model;
[0186] The model using module is used for:
[0187] Inputting the test image of the overhead pipeline to be measured into the defect recognition model to obtain the defect recognition result output by the defect recognition model.
[0188] The performing data augmentation on the original defect images in the data set to obtain multiple groups of augmented defect images; wherein, the augmented defect images include a first part of the image and a second part of the image, includes:
[0189] Performing data augmentation on the original defect images in the data set using a first preprocessing method to obtain the first part of the image;
[0190] The original defective images in the dataset are respectively data-augmented using the second preprocessing method and the third preprocessing method to obtain the second part of the images.
[0191] The defect recognition model includes a backbone network; wherein, the backbone network is used to extract basic features to obtain the first basic feature map of the first part of the images and extract the second basic feature map of the second part of the images.
[0192] The classification branch sub-model includes: an adaptive perception module, a global attention pooling module, and a classification layer;
[0193] The adaptive perception module is used to dynamically adjust the first basic feature map of the first part of the images, and fuse the adjusted features with the original features through a residual connection to obtain the first adjusted feature map after fusion;
[0194] The global attention pooling module is used to perform a pooling operation on the first adjusted feature map after fusion, dynamically assign weights to different feature regions, and thus generate the first feature vector of the first part of the images;
[0195] The classification layer is used to map the first feature vector to obtain a first mapping vector, and use the first mapping vector to construct the first loss function of the classification branch sub-model.
[0196] The adaptive perception module is used to dynamically adjust the first basic feature map of the first part of the images, and fuse the adjusted features with the original features through a residual connection to obtain the first adjusted feature map after fusion, including:
[0197] Use the first convolutional branch and the second convolutional branch to respectively perform convolutional operations on the first basic feature map to obtain a first convolutional map and a second convolutional map; wherein, the first convolutional map and the second convolutional map are used to be fused to obtain a first fused map;
[0198] Perform a pooling operation and a compression operation on the first fused map in sequence to obtain a fourth feature vector;
[0199] Process the fourth feature vector through an attention mechanism to obtain the first weight of the first convolutional branch and the second weight of the second convolutional branch;
[0200] Fuse the first convolutional map and the second convolutional map according to the first weight and the second weight to obtain the preliminary adjusted feature map.
[0201] Fuse the preliminary adjusted feature map and the first basic feature map through a residual connection by element-wise addition to obtain the finally output first adjusted feature map.
[0202] The second part of the image includes a first sub-image and a second sub-image;
[0203] The representation learning branch sub-model includes: a cross-space feature fusion module, a multi-layer perceptron, and a clustering module;
[0204] The cross-space feature fusion module is used to enrich the second base feature map of the first sub-image and the third base feature map of the second sub-image to obtain a second feature vector of the first sub-image and a third feature vector of the second sub-image;
[0205] The multi-layer perceptron is used to map the second feature vector and the third feature vector to obtain a second mapped vector of the second feature vector and a third mapped vector of the third feature vector;
[0206] The clustering module is used to construct a second loss function of the representation learning branch sub-model based on the vMF mixture distribution for the second mapped vector and the third mapped vector.
[0207] The cross-space feature fusion module enriches the second base feature map of the first sub-image in the following manner to obtain a second feature vector of the first sub-image:
[0208] Use multiple branches to process the second base feature map to obtain the processing results of multiple branches;
[0209] By fusing the processing results of multiple branches, a second feature vector of the first sub-image is obtained.
[0210] As Figure 7 shown, an embodiment of the present application provides an electronic device for executing the method for identifying overhead pipeline defects in the present application. The device includes a memory, a processor, a bus, and a computer program stored on the memory and executable on the processor. Wherein, when the above-mentioned processor executes the above-mentioned computer program, the steps of the above-mentioned method for identifying overhead pipeline defects are implemented.
[0211] Specifically, the above-mentioned memory and processor can be general memory and processor, which are not specifically limited here. When the processor runs the computer program stored in the memory, it can execute the above-mentioned method for identifying overhead pipeline defects.
[0212] Corresponding to the method for identifying overhead pipeline defects in the present application, an embodiment of the present application also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the above-mentioned method for identifying overhead pipeline defects are executed.
[0213] Specifically, the storage medium can be a general storage medium, such as a removable disk, a hard disk, etc. When the computer program on the storage medium is run, it can execute the above-mentioned method for identifying overhead pipeline defects.
[0214] In the embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the system or unit can be in electrical, mechanical or other forms.
[0215] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0216] In addition, each functional unit in the embodiments provided in the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0217] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0218] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0219] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for identifying defects in an overhead pipeline, characterized in that: The method comprises: a model training phase and a model use phase; The model training phase includes: Performing data enhancement on the original defect images in the data set to obtain multiple groups of enhanced defect images; wherein the enhanced defect images include a first partial image and a second partial image; Using the first feature vector of the first partial image to construct a first loss function of the classification branch sub-model; Using the second feature vector of the second partial image to construct a second loss function representing the learning branch sub-model; wherein the first loss function and the second loss function are used to construct an overall loss function; By adjusting the parameters of the overall loss function, a defect recognition model is obtained; The model use phases include: Inputting the image of the overhead pipeline to be tested into the defect recognition model to obtain a defect recognition result output by the defect recognition model; The classification branch sub-model includes: an adaptive perception module, a global attention pooling module, and a classification layer; The adaptive perception module is used to dynamically adjust the first basic feature map of the first partial image, and fuse the adjusted features with the original features through a residual connection to obtain a fused first adjusted feature map; The global attention pooling module is used to perform a pooling operation on the fused first adjusted feature map, dynamically assign weights to different feature regions, and thus generate a first feature vector of the first partial image; The classification layer is used to map the first feature vector to obtain a first mapping vector, and use the first mapping vector to construct a first loss function of the classification branch sub-model; The second partial image includes a first sub-image and a second sub-image; The representation learning branch sub-model includes: a cross-space feature fusion module, a multi-layer perceptron and a clustering module; The cross-space feature fusion module is used to enrich the second basic feature map of the first sub-image and the third basic feature map of the second sub-image to obtain the second feature vector of the first sub-image and the third feature vector of the second sub-image; The multilayer perceptron is used to map the second eigenvector and the third eigenvector to obtain a second mapping vector of the second eigenvector and a third mapping vector of the third eigenvector; A clustering module is used to construct a second loss function representing a learning branch sub-model for the second mapping vector and the third mapping vector based on a vMF mixed distribution.
2. The method according to claim 1, characterized in that The data enhancement is performed on the original defect images in the data set to obtain a plurality of enhanced defect images; wherein the enhanced defect images include a first partial image and a second partial image, including: Performing data enhancement on the original defect image in the data set using a first preprocessing method to obtain the first partial image; The second preprocessing method and the third preprocessing method are used to perform data enhancement on the original defect image in the data set to obtain the second part of the image.
3. The method according to claim 1, characterized in that The defect recognition model includes a backbone network; wherein the backbone network is used to extract basic features to obtain a first basic feature map of the first partial image and to extract a second basic feature map of the second partial image.
4. The method according to claim 1, characterized in that: The adaptive perception module is used to dynamically adjust the first basic feature map of the first partial image, and fuse the adjusted features with the original features through a residual connection to obtain a fused first adjusted feature map, including: Using a first convolution branch and a second convolution branch to perform a convolution operation on the first basic feature map respectively, to obtain a first convolution map and a second convolution map; wherein the first convolution map and the second convolution map are used to fuse to obtain a first fusion map; Performing a pooling operation and a compression operation on the first fusion image in sequence to obtain a fourth eigenvector; Processing the fourth eigenvector through an attention mechanism to obtain a first weight of the first convolution branch and a second weight of the second convolution branch; Fusing the first convolution map and the second convolution map according to the first weight and the second weight to obtain a preliminary adjusted feature map; The preliminary adjusted feature map is fused with the first basic feature map by element-by-element addition through a residual connection to obtain the first adjusted feature map that is finally output.
5. The method according to claim 1, characterized in that The cross-space feature fusion module enriches the second basic feature map of the first sub-image in the following manner to obtain a second feature vector of the first sub-image: Using multiple branches to process the second basic feature map to obtain processing results of multiple branches; The second feature vector of the first sub-image is obtained by fusing the processing results of multiple branches.
6. A device for identifying defects in overhead pipelines, characterized in that: The device comprises: a model training module and a model use module; The model training module is used to: Performing data enhancement on the original defect images in the data set to obtain multiple groups of enhanced defect images; wherein the enhanced defect images include a first partial image and a second partial image; Using the first feature vector of the first partial image to construct a first loss function of the classification branch sub-model; Using the second feature vector of the second partial image to construct a second loss function representing the learning branch sub-model; wherein the first loss function and the second loss function are used to construct an overall loss function; By adjusting the parameters of the overall loss function, a defect recognition model is obtained; The model uses modules for: Inputting the image of the overhead pipeline to be tested into the defect recognition model to obtain a defect recognition result output by the defect recognition model; The classification branch sub-model includes: an adaptive perception module, a global attention pooling module, and a classification layer; The adaptive perception module is used to dynamically adjust the first basic feature map of the first partial image, and fuse the adjusted features with the original features through a residual connection to obtain a fused first adjusted feature map; The global attention pooling module is used to perform a pooling operation on the fused first adjusted feature map, dynamically assign weights to different feature regions, and thus generate a first feature vector of the first partial image; The classification layer is used to map the first feature vector to obtain a first mapping vector, and use the first mapping vector to construct a first loss function of the classification branch sub-model; The second partial image includes a first sub-image and a second sub-image; The representation learning branch sub-model includes: a cross-space feature fusion module, a multi-layer perceptron and a clustering module; The cross-space feature fusion module is used to enrich the second basic feature map of the first sub-image and the third basic feature map of the second sub-image to obtain the second feature vector of the first sub-image and the third feature vector of the second sub-image; The multilayer perceptron is used to map the second eigenvector and the third eigenvector to obtain a second mapping vector of the second eigenvector and a third mapping vector of the third eigenvector; A clustering module is used to construct a second loss function representing a learning branch sub-model for the second mapping vector and the third mapping vector based on a vMF mixed distribution.
7. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the method for identifying overhead pipeline defects as described in any one of claims 1 to 5 are performed.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for identifying defects of an overhead pipeline as claimed in any one of claims 1 to 5 are executed.
Citation Information
Patent Citations
Semi-supervised product surface defect detection method and system based on heavy balance
CN114998258A
Face recognition model training method and device, face recognition method and device and storage medium
CN116978100A