Bolt Defect Detection Method Based on Semi-Supervised Learning and Prior Knowledge Embedding Strategy
Through the variational autoencoder network model of semi-supervised learning and prior knowledge embedding strategy, the problem of sample imbalance in the detection of bolt defects in transmission line is solved, efficient bolt defect detection is achieved, and labeling costs are reduced and detection accuracy and speed are improved.
Patent Information
- Application Number
- CN202210378734.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-04-12
AI Technical Summary
In the detection of bolt defects of transmission lines, traditional manual inspection efficiency is low, deep learning models rely on a large amount of labeled data and unbalanced samples, resulting in insufficient detection accuracy and generalization capabilities, especially in the long-tail distribution, it is difficult to effectively utilize information without labeled data.
Using a method based on semi-supervised learning and prior knowledge embedding strategy, a variational autoencoder network model is constructed, combined with batch normalization units, graph convolutional neural networks and convolution units, and synergistic training is used to enhance the generalization and feature extraction capabilities of the model.
Under low proportion of labeled data, the model performance is close to the supervised learning model, which improves the accuracy of detection of bolt defects under long tail distribution, reduces the cost of data labeling, and improves detection speed and accuracy.
Smart Images

Figure CN114708518B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and more particularly, to a bolt defect detection method based on semi-supervised learning and prior knowledge embedding strategy. Background Art
[0002] As a vital line in the transmission of power systems, the stable operation of transmission lines has a crucial impact on the safety of the power grid. The traditional manual inspection method can no longer meet the needs of today's society. Therefore, how to use computer vision technology to automatically and accurately locate and detect defects from the aerial inspection images of drones has become a key technical issue. The recognition and classification efficiency of classical machine learning algorithms for transmission line bolts are relatively low, and the processing methods based on machine learning require manual extraction of bolt defect features, resulting in poor generality of such methods. Currently, the mainstream object detection algorithms are Faster R-CNN and Yolo series. Training models using the above two deep learning algorithms rely on a large amount of labeled data. However, due to the time-consuming and laborious labeling work, it costs a high price to obtain a large amount of labeled data in actual production. In addition, the imbalance of training sample category data samples will also cause extremely poor final training effects of deep learning models.
[0003] In the prior art, according to whether labeled data is used during the training process, machine learning models can be divided into three types: "supervised learning", "unsupervised learning", and "semi-supervised learning". The key to whether the object detection algorithm based on supervised learning can achieve good performance lies in whether there is sufficient labeled sample data. However, in many scenarios such as power production, due to reasons such as the small total amount of crisis defect sample data and the high cost of image annotation, it is difficult to meet the requirement of obtaining a large amount of labeled data sets. Unsupervised learning does not use labeled data at all, and by mining the internal features of the data, the relationship between samples is found. However, due to the lack of prior knowledge in unsupervised learning algorithms, it is impossible to predict the accuracy of the output results, and it is difficult to be applied to actual production. Semi-supervised learning uses both unlabeled data and labeled data for collaborative training. The unlabeled data will be assigned a label in the semi-supervised learning model and then used as labeled data, which plays a role in expanding the data set and thus improving the quality of the model. However, in actual work, the quality of unlabeled data is difficult to control, and improper use of unlabeled data will have a great impact on performance.
[0004] In response to the above problems, there are currently methods for target detection based on semi-supervised learning. For example, Huazhong University of Science and Technology has disclosed a method for automatically labeling images of industrial product appearance defects based on semi-supervised learning (application number: CN202010804831.X). This method uses a trained deep convolutional network classification model to label unknown label data and label it with a pseudo-label. Afterwards, a predetermined amount of data is extracted from the pseudo-label data and incorporated into the known label data to form a new training set to iteratively train the automatic labeling model. Since the labeling results of this method largely depend on the accuracy of the deep learning model trained with the previously labeled samples, if the pseudo-label carries too much noise, the accuracy of the final model of the semi-supervised training cannot be guaranteed. Ping An Technology (Shenzhen) Co., Ltd. has also disclosed a target detection method, device and storage medium based on semi-supervised learning (application number: 202011288652.1). This method uses the data expanded by operations such as horizontal flipping and data enhancement of the labeled data as new training samples to improve the detection generalization ability of the model. However, this approach severely limits the availability of unlabeled samples and fails to fully utilize the information contained in a large number of unlabeled samples to help improve model accuracy.
[0005] Secondly, deep learning models are currently unable to learn the difference between category A samples and category B samples, the connection and dependency information between category C samples and category B samples, and it is difficult to mine the potential knowledge information behind different category samples. In the current target detection task, for a certain sample, the category label is either category A, then category B, or category C, and so on. However, this annotation method does not have other semantic information except for the category information. Each category label is regarded as the basis on the axes orthogonal to each other in the Cartesian coordinate system, which means that the Euclidean distance between each category is consistent, which is seriously inconsistent with the actual category meaning in reality. In other words, the model believes that bolts are normal, bolts are missing pins, and bolts are missing nuts. All are equivalent categories, but obviously, the features of normal bolts and bolts missing pins are very similar, while normal bolts and bolts missing nuts are obviously different.
[0006] In addition, there are currently only two ways to deal with the long-tail distribution problem: one based on resampling and weighting and the other based on balancing different sample subsets.
[0007] On the one hand, the method based on resampling and weighting can only solve the problem that the head classes occupy more gradient backpropagation than the tail classes during gradient backpropagation. In nature, defective samples have the typical long-tail distribution characteristic, and different defect classes have different numbers of instances. More importantly, defects of the same class do not all appear in one image, so the problem of few-shot learning for tail features is not solved. On the other hand, in nature, the distribution of different class labels is usually correlated with other classes and does not have the characteristic of independent and identical distribution, so the problem that it is difficult to identify features due to fewer tail sample data cannot be solved. Therefore, there is an urgent need for a technical solution to fundamentally solve the problem of unstable model detection ability caused by sample imbalance by enhancing the feature extraction ability of tail classes. Summary of the Invention
[0008] To overcome the deficiencies of the prior art, the present invention proposes a method for bolt defect detection that comprehensively utilizes semi-supervised learning and prior knowledge embedding strategy, and co-trains multiple related tasks together, enabling the sharing of parameters and data between each task, thereby increasing the generalization ability of the model. Specifically, this method is a bolt defect detection method based on semi-supervised learning and prior knowledge embedding strategy.
[0009] To achieve the above technical objectives, the present invention provides a bolt defect detection method based on semi-supervised learning and prior knowledge embedding strategy, including the following steps:
[0010] Collect bolt images of different components and establish a bolt detection data set for different types of defects;
[0011] Based on the prior knowledge embedding strategy, perform feature processing on the bolt detection data set to generate a sample data set with sample correlation features, where the sample correlation features are used to indicate that the features between samples are correlated;
[0012] Construct a variational autoencoder network model, which consists of a batch normalization unit, a graph convolutional neural network unit, and a convolutional unit;
[0013] Based on the variational autoencoder network model, through semi-supervised learning, input the sample data set into the variational autoencoder network model for training to construct a bolt defect detection model, which is used to identify the bolt defect types of the bolt images to be detected.
[0014] Preferably, during the process of collecting bolt images of different components, the bolt images include normal bolt images, bolt images without washers, bolt images without pins, and bolt images without nuts.
[0015] Preferably, in the process of constructing the bolt defect detection model, based on the variational autoencoder network model, for the sample data set, feature extraction is performed through the CNN backbone module to obtain image feature information;
[0016] The image feature information is converted into a feature sequence through the encoder module, and the position information encoding is incorporated to obtain a latent space feature vector;
[0017] Under the guidance of prior knowledge, the latent space feature vector is transformed into several intermediate features, and through the feed-forward neural network FNN, it is decomposed into the target coordinates and the classification label;
[0018] According to the target coordinates and the classification label, the sample is labeled to generate labeled data;
[0019] According to the labeled data and the unlabeled data, the variational autoencoder network model is trained to construct the bolt defect detection model.
[0020] Preferably, in the process of constructing the bolt defect detection model, the constructed bolt defect detection model further includes a first decoder module and a second decoder module;
[0021] The first decoder module is used to integrate and obtain the sample correlation feature according to the position information encoding and the intermediate feature, where the sample correlation feature includes the position feature and the class-specific feature information;
[0022] The second decoder module is used to remove the position feature through the deconvolution operation to obtain the class-specific feature information, and the class-specific feature information is used to obtain the feature matching degree between the intermediate feature and the corresponding class-specific feature information;
[0023] Preferably, in the process of training the variational autoencoder network model, a training set is generated according to the labeled data and the unlabeled data;
[0024] The training set is input into the variational autoencoder network model, and a forward propagation is performed once, and then the model is trained through the backpropagation algorithm to obtain the predicted class, the boundary prediction value, the true class label, and the boundary.
[0025] Preferably, in the process of training the model through the backpropagation algorithm, the model is trained with the labeled data to obtain the classification loss and the bounding box regression loss;
[0026] The equation expression of the classification loss is:
[0027]
[0028] The equation expression of the bounding box regression loss is:
[0029]
[0030] p i represents the probability that the i-th anchor is predicted as the true label. It is 1 for positive samples and 0 for negative samples, t i represents the bounding box regression parameters for predicting the i-th anchor. represents the true bounding box regression parameters corresponding to the i-th anchor, and R is the LOSS CIOU Loss function.
[0031] Preferably, during the training of the model by the backpropagation algorithm, LOSS CIOU The equation expression of the loss function is:
[0032]
[0033]
[0034]
[0035] where IOU is the ratio of the intersection and union of the predicted bounding box and the true bounding box, b, b gt are the center points of the predicted bounding box and the true bounding box respectively, ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the predicted bounding box and the true bounding box, c represents the diagonal distance of the smallest closed region that can simultaneously contain the predicted bounding box and the true bounding box, α is a parameter for balancing the ratio, and v is used to measure the ratio consistency between the anchor box and the predicted box.
[0036] Preferably, during the training of the model by the backpropagation algorithm, by limiting the distance between the output features of each convolutional layer, unlabeled data is used for model training, where the feature matching loss between the original image and the corresponding synthesized image is used to constrain the model training process:
[0037]
[0038] where, N cls represents the number of samples in a batch, N reg represents the number of anchor positions.
[0039] Preferably, during the process of using unlabeled data for model training, taking the feature information Q corresponding to the original image as the input, through the first decoder module, the position feature and the category-specific feature information can be obtained. The second decoder module uses deconvolution to integrate the two features of the first decoder module to obtain the feature information. According to the duality, taking the obtained feature information as the input, the position feature and the category-specific feature information can also be obtained through the first and second decoder modules. Since the classification and bounding box regression losses cannot be directly calculated for unlabeled data samples, the present invention utilizes the duality of feature extraction of unlabeled data to design the following loss:
[0040]
[0041] where T represents the number of layers for feature extraction in the discriminator, D i represents the extracted feature, and N i represents the number of features extracted by the discriminator network at the i-th layer.
[0042] Preferably, a bolt defect detection system for implementing the bolt defect detection method includes:
[0043] A data acquisition module for acquiring bolt images of different components and establishing a bolt detection data set for different types of defects;
[0044] A data processing module for performing feature processing on the bolt detection data set based on the prior knowledge embedding strategy to generate a sample data set with sample correlation features, where the sample correlation features are used to represent the correlation of features between samples;
[0045] A defect recognition module for constructing a variational autoencoder network model, which consists of a batch normalization unit, a graph convolutional neural network unit, and a convolutional unit; based on the variational autoencoder network model, through semi-supervised learning, the sample data set is input into the variational autoencoder network model for training to construct a bolt defect detection model, and the bolt defect detection model is used to identify the bolt defect types of the bolt images to be detected.
[0046] The present invention discloses the following technical effects:
[0047] The present invention proposes a bolt defect detection method based on semi-supervised learning and prior knowledge embedding strategy. On the one hand, a unique variational autoencoder network is designed to transform the object detection problem into a set of dual problems of image transformation, and the dual relationship between them is used as a constraint. Two learning models are trained simultaneously, and the performance of the two models promotes each other. Finally, the model can reach or approach the performance result of the supervised learning model under the condition of a low proportion of labeled data, thus greatly reducing the data labeling cost;
[0048] On the other hand, make full use of the correlation and dependence of different categories of samples to improve the accuracy of bolt defect detection under long-tailed distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a flowchart of the bolt defect detection method according to the embodiment of the present invention;
[0051] Figure 2 It is a schematic structural diagram of the variational autoencoder according to the embodiment of the present invention;
[0052] Figure 3 It is a schematic diagram of the correlation relationship between samples according to the embodiment of the present invention;
[0053] Figure 4 It is a schematic diagram of a physical sample according to the embodiment of the present invention;
[0054] Figure 5 It is a comparison chart of the detection effects of different models according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0056] As Figures 1-5 shown, the present invention provides a bolt defect detection method based on a semi-supervised learning and prior knowledge embedding strategy, including the following steps:
[0057] Collect bolt images of different components and establish a bolt detection data set for different types of defects;
[0058] Based on the prior knowledge embedding strategy, the bolt detection data set is processed for features to generate a sample data set with sample correlation features, where the sample correlation features are used to represent that the features between samples are correlated;
[0059] Build a variational autoencoder network model, which consists of a batch normalization unit, a graph convolutional neural network unit, and a convolutional unit;
[0060] Based on the variational autoencoder network model, through semi-supervised learning, the sample data set is input into the variational autoencoder network model for training to build a bolt defect detection model, which is used to identify the bolt defect types of the bolts in the images to be detected.
[0061] Further preferably, during the process of collecting bolt images of different components, the bolt images include normal bolt images, bolt images without gaskets, bolt images without pins, and bolt images without nuts.
[0062] Further preferably, during the process of building the bolt defect detection model, based on the variational autoencoder network model, the sample data set is subjected to feature extraction through the CNN backbone module to obtain image feature information;
[0063] The image feature information is converted into a feature sequence through the encoder module, and the position information encoding is incorporated to obtain a latent space feature vector;
[0064] Under the guidance of prior knowledge, the latent space feature vector is transformed into several intermediate features, and decomposed into target coordinates and classification labels through a feed-forward neural network FNN;
[0065] According to the target coordinates and classification labels, the samples are labeled to generate labeled data;
[0066] According to the labeled data and unlabeled data, the variational autoencoder network model is trained to build a bolt defect detection model.
[0067] Further preferably, preferably, during the process of building the bolt defect detection model, the bolt defect detection model further includes a first decoder module and a second decoder module;
[0068] The first decoder module is used to integrate and obtain sample correlation features according to the position information encoding and intermediate features, where the sample correlation features include position features and category-specific feature information;
[0069] The second decoder module is used to remove the position features through deconvolution operations to obtain class-specific feature information, and the class-specific feature information is used to obtain the feature matching degree between the intermediate features and the corresponding class-specific feature information.
[0070] The first decoder module is used to integrate the position information encoding and the intermediate feature Q (the intermediate feature Q refers to the feature obtained by the original image through the feature extraction network) to obtain the feature information H. The feature information H includes the position feature L and the class-specific feature information X, Y, Z, and W.
[0071] The second decoder module takes the feature information H integrated by the first decoder module and removes the position encoding information through deconvolution operations to obtain the feature information R.
[0072] The feature information R refers to the class feature information of the image (that is, the position information is removed from the feature information H). The purpose of obtaining the feature information R is to calculate the feature matching degree between the intermediate feature Q of the original image and the corresponding synthesized image feature R.
[0073] Further preferably, during the training of the variational autoencoder network model, a training set is generated according to the labeled data and the unlabeled data;
[0074] The training set is input into the variational autoencoder network model, and a forward propagation is performed once, and then the model is trained through the backpropagation algorithm to obtain the predicted class, the boundary prediction value, the true class label, and the boundary.
[0075] Further preferably, during the training of the model through the backpropagation algorithm, the model is trained with the labeled data to obtain the classification loss and the bounding box regression loss;
[0076] The equation expression of the classification loss is:
[0077]
[0078] The equation expression of the bounding box regression loss is:
[0079]
[0080] p i represents the probability that the i-th anchor is predicted as the true label, which is 1 for positive samples and 0 for negative samples, and t i represents the bounding box regression parameter for predicting the i-th anchor, represents the true bounding box regression parameter corresponding to the i-th anchor, and R is the LOSS CIOU loss function.
[0081] Further preferably, during the process of training the model through the backpropagation algorithm, LOSS CIOU The equation expression of the loss function is:
[0082]
[0083]
[0084]
[0085] Among them, IOU is the ratio of the intersection and union of the predicted bounding box and the true bounding box, b, b gt are the center points of the predicted bounding box and the true bounding box respectively, ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the predicted bounding box and the true bounding box, c represents the diagonal distance of the smallest closed region that can simultaneously contain the predicted bounding box and the true bounding box, α is a parameter used to balance the ratio, and v is used to measure the ratio consistency between the anchor box and the predicted box.
[0086] Further preferably, during the process of training the model through the backpropagation algorithm, by limiting the distance between the output features of each convolutional layer, unlabeled data is used for model training; among them, the feature matching loss between the original image and the corresponding synthetic image is used to constrain the model training process:
[0087]
[0088] Among them, N cls represents the number of samples in a batch, N reg represents the number of anchor positions.
[0089] Further preferably, during the process of using unlabeled data for model training, the process of model training is expressed as:
[0090]
[0091] Among them, T represents the number of layers for extracting features in the discriminator, D i represents the extracted features, N i represents the number of features extracted by the i-th layer discriminator network.
[0092] Further preferably, a bolt defect detection system for implementing the bolt defect detection method includes:
[0093] A data acquisition module for collecting bolt images of different components and establishing a bolt detection data set for different types of defects;
[0094] A data processing module, which is used to perform feature processing on the bolt detection data set based on the prior knowledge embedding strategy to generate a sample data set with sample correlation features, where the sample correlation features are used to indicate that the features between samples are correlated;
[0095] A defect recognition module, which is used to build a variational autoencoder network model. The variational autoencoder network model consists of a batch normalization unit, a graph convolutional neural network unit, and a convolutional unit; based on the variational autoencoder network model, through semi-supervised learning, the sample data set is input into the variational autoencoder network model for training to build a bolt defect detection model, and the bolt defect detection model is used to identify the bolt defect types of the bolts to be detected in the images.
[0096] The present invention proposes a bolt defect detection method based on semi-supervised learning and prior knowledge embedding strategy, as Figure 1 shown in the overall flowchart of the embodiment of the present invention, which includes the following steps:
[0097] Step 1: Establish a bolt detection data set for different types of defects. The inspection of the transmission line is carried out by a drone to collect bolt images on various components. The bolt images are divided into 4 categories: bolt-normal (ls-zc), bolt-missing gasket (ls-qdp), bolt-missing pin (ls-qxd), and bolt-missing nut (ls-qlm), with a total of 1,970 images. Schematic diagrams of various samples are shown in the following figures.
[0098] Step 2: The prior knowledge embedding strategy. The prior knowledge embedding strategy refers to using the effectively learned and marked samples to capture and fuse the correlation and dependence between different categories of samples and between categories in the natural scene, using the graph embedding vector to replace the original one-hot encoding to represent different category information, increasing the semantic information contained in the category label, and improving the reasoning ability between different categories of samples. The advantage of the graph embedding vector is that when judging the category of a sample, when it is uncertain from the features of a single sample alone, the features of related samples can be aggregated and supplemented on the features of the current sample, thereby improving the discrimination ability of a single sample. The graph embedding is to map the graph model to a low-dimensional vector space, and the represented vector form should also try to retain the structural information and potential characteristics of the graph model. The association relationship between samples is shown in the appendix Figure 3 , where the node IDx represents bolt samples of different categories, and the connection lines between nodes indicate that the features between samples are correlated. The correlation relationships between different samples can be transmitted through the edges in the graph.
[0099] Step 3: Build a variational autoencoder network model. The structural schematic of the variational autoencoder is shown in the appendix Figure 2As shown, where Sync BN represents batch normalization, GCN represents graph convolutional neural network, and Conv represents convolution.
[0100] Specifically, first, the image is subjected to feature extraction through the CNN backbone module to obtain the feature information of the image. Second, the encoder module converts the extracted features into a 1D sequence and incorporates positional information encoding to obtain the latent space feature vector. Then, the first decoder module converts the latent space feature vector into N intermediate features under the guidance of prior knowledge. After that, the second decoder module converts the N intermediate features output by the first decoder module into the feature vector of the corresponding image under the guidance of positional information encoding. Finally, the N intermediate features output by the first decoder module are decoded into the target coordinates and classification labels through the feed-forward neural network FNN, and a joint training model is established. This process can be carried out in a semi-supervised learning manner to reduce the dependence on labeled data.
[0101] Step 4: Use the labeled data and unlabeled data for collaborative training. The collaborative training refers to using a small number of labeled samples and a large number of unlabeled samples. Through the dual optimization strategy, the images in one domain are transformed into another domain, and at the same time, the transformed images can be transformed back to the original domain.
[0102] Divide the bolts obtained in Step 1 into a training set and a validation set; for the labeled samples in the training set, in each epoch during the training process, the same group of labeled samples are input into the variational autoencoder network constructed in Step 3 for one forward propagation, and then the model is trained through the backpropagation algorithm. The predicted category and the predicted bounding box value are compared with the true category label and the bounding box, and the classification loss is calculated by formula (2) and the bounding box regression loss is calculated by formula (3):
[0103]
[0104]
[0105]
[0106] In formula (4), p i represents the probability that the i-th anchor is predicted as the true label, which is 1 when it is a positive sample and 0 when it is a negative sample. t i represents the bounding box regression parameter for predicting the i-th anchor, represents the "true bounding box" regression parameter corresponding to the i-th anchor. N cls represents the number of samples in a batch. N regIndicates the number of anchor positions (not the number of anchors), and R is the LOSS CIOU Loss function. Among them, LOSS CIOU The specific calculation method is shown in Formula (5).
[0107]
[0108]
[0109]
[0110] In Formula (5), IOU is the ratio of the intersection and union of the "predicted bounding box" and the "ground truth bounding box", b, b gt are the center points of the "predicted bounding box" and the "ground truth bounding box" respectively, ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the "predicted bounding box" and the "ground truth bounding box", c represents the diagonal distance of the smallest closed region that can simultaneously contain the "predicted bounding box" and the "ground truth bounding box", α is a parameter used to balance the ratio, and v is used to measure the ratio consistency between the anchor box and the predicted box.
[0111] For the unlabeled samples in the training set, in the case where the classification label is not available, the synthetic image corresponding to the original image is obtained by passing the same image through the variational autoencoder network model defined in Step 3. Then, the bolt detection model is trained through the backpropagation algorithm, and the training process of the bolt detection model is constrained by using the feature matching loss between the original image and the corresponding synthetic image. Specifically, it is to use a multi-layer discriminator to extract features from the original image and the synthetic image, and then calculate the L1 distance between the convolutional output features of each layer:
[0112]
[0113] Among them, T represents the number of layers for extracting features in the discriminator, D i represents the extracted features, N i represents the number of features extracted by the i-th layer discriminator network.
[0114] Finally, the test set is used to verify the training effect of the bolt defect detection model to obtain the trained bolt detection model.
[0115] Step 5: Collect the bolt images to be detected and input them into the trained bolt defect detection model to detect the defect conditions of the bolts. The defect conditions are the defect condition labels and bounding box information with the highest confidence in the feature vector representing the bolt defect conditions output by the bolt defect detection model.
[0116] The images in the bolt defect dataset used in the method proposed by the present invention and the method of the comparative experiment were all taken during the actual inspection by the unmanned aerial vehicle, and the sample data was labeled with reference to the labeling scheme of the transmission line equipment of the State Grid. Each image in the labeled samples has a corresponding XML labeling file, and the XML file contains the name of the image, the category of the target, and the coordinate information of the bounding box. The dataset includes 4 categories of samples: bolt-normal (ls-zc), bolt-missing gasket (ls-qdp), bolt-missing pin (ls-qxd), and bolt-missing nut (ls-qlm), with a total of 1970 pictures. The schematic diagrams and quantity distributions of various samples are as Figure 4 shown in Table 1.
[0117] Table 1
[0118]
[0119] The present invention uses the mean Average Precision (mAP) and the inference speed (Frame Per Second, FPS) as the evaluation indicators for the accuracy and processing speed of object detection. Among them, mAP is the average of the accuracies of all categories and is an indicator to measure the overall accuracy of the object detection model. FPS represents the number of pictures that can be processed per second and can effectively measure the processing speed of the algorithm.
[0120] The experimental results of the bolt defect detection method based on semi-supervised learning and prior knowledge embedding strategy proposed by the present invention are compared with those of the Faster R-CNN and YOLOv5 models. Among them, for Faster-RCNN and YOLOv5, after labeling on the entire labeled bolt dataset, the data used by the bolt defect detection method based on semi-supervised learning and prior knowledge embedding strategy proposed by the present invention is 40% labeled data and 60% unlabeled data. Set the training Epoch to 100 and the Batch size to 8. During the training process, the learning rate gradually increases from 10-6 to 10-3 in the first 10 Epochs, the learning rate is 10-2 from the 10-19th Epoch to the 20th Epoch, and the learning rate drops to 10-4 from the 20th to the 100th Epoch. To prevent the model from falling into a local optimum, the SGD optimizer is used, with the momentum coefficient set to 0.9 and the decay coefficient set to 0.005. The comparison results of the detection performances of different models are shown in Table 2.
[0121] Table 2
[0122]
[0123] The results are shown in Table 2. When the IOU threshold is taken as 0.5, compared with the Faster R-CNN model, the mAP of the method proposed in the present invention is increased by 2.8%, and the FPS is increased by 3.3, improving the detection accuracy and detection speed. Compared with the YOLOv5 model, the method proposed in the present invention improves the detection accuracy of bolt defects on transmission lines, but there is still room for improvement in the detection speed.
[0124] From Figure 5 (a) and Figure 5 (b) and Figure 5 (c) comparison effects, it can be seen that the confidence levels of the Faster R-CNN and YOLOv5 models for the detection of small targets are both lower than the method proposed in the present invention, and there is a situation of missed detection in Faster R-CNN. Relatively speaking, the method proposed in the present invention can effectively detect the bolt targets on the hanging parts of the transmission line and can avoid misdetection of small targets. The experimental results show that in the case of unbalanced data samples, the method proposed in the present invention can effectively improve the detection performance of the model. At the same time, the method proposed in the application only uses 60% of the entire data set during initialization, and the final detection result is close to the model using the entire data set. To a certain extent, it proves that the method proposed in the present invention can effectively reduce the number of labeled pictures and thus reduce the labeling cost of images.
[0125] The present invention has been described in detail above. The above description is only a preferred embodiment of the present invention, and the scope of implementation of the present invention cannot be limited. That is, all equal changes and modifications made according to the scope of the present invention should still fall within the scope covered by the present invention. As used herein, unless otherwise specified, the use of ordinal numbers "first", "second", "third", etc. to describe ordinary objects only represents different instances of similar objects and does not intend to imply that the objects so described must have a given order in terms of time, space, sorting, or any other way.
[0126] Although the present invention has been described based on a limited number of embodiments, those skilled in the art in this technical field will understand that other embodiments can be envisioned within the scope of the present invention described herein. In addition, it should be noted that the language used in this specification is mainly selected for readability and teaching purposes, rather than for the purpose of explaining or limiting the subject matter of the present invention. Therefore, many modifications and changes are obvious to those of ordinary skill in the art in this technical field without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative, not restrictive, and the scope of the present invention is defined by the appended claims.
Claims
1. A bolt defect detection method based on a semi-supervised learning and prior knowledge embedding strategy, characterized in that, It includes the following steps: Collect bolt images of different components. During the process of collecting bolt images of different components, the bolt images include normal bolt images, bolt images without washers, bolt images without pins, and bolt images without nuts, and establish a bolt detection data set for different types of defects; Based on the prior knowledge embedding strategy, perform feature processing on the bolt detection data set to generate a sample data set with sample correlation features, where the sample correlation features are used to represent that the features between samples are correlated; Construct a variational autoencoder network model, which is composed of a batch normalization unit, a graph convolutional neural network unit, and a convolutional unit; Based on the variational autoencoder network model, in a semi-supervised learning manner, input the sample data set into the variational autoencoder network model for training to construct a bolt defect detection model. During the process of constructing the bolt defect detection model, based on the variational autoencoder network model, for the sample data set, perform feature extraction through the CNN backbone module to obtain image feature information; convert the image feature information into a feature sequence through the encoder module, and integrate the position information encoding to obtain a latent space feature vector; under the guidance of prior knowledge, convert the latent space feature vector into several intermediate features, and decompose them into target coordinates and classification labels through a feed-forward neural network FNN; according to the target coordinates and the classification labels, label the samples to generate labeled data; according to the labeled data and unlabeled data, train the variational autoencoder network model to construct the bolt defect detection model; the bolt defect detection model is used to identify the bolt defect types of the bolt images to be detected.
2. The bolt defect detection method based on semi-supervised learning and prior knowledge embedding strategy according to claim 1, wherein: During the process of constructing the bolt defect detection model, the construction of the bolt defect detection model further includes a first decoder module and a second decoder module; The first decoder module is used to integrate the position information encoding and the intermediate features to obtain the sample correlation features, where the sample correlation features include position features and class-specific feature information; The second decoder module is used to remove the position features through deconvolution operations to obtain the class-specific feature information, and the class-specific feature information is used to obtain the feature matching degree between the intermediate features and the corresponding class-specific feature information.
3. The bolt defect detection method based on semi-supervised learning and prior knowledge embedding strategy according to claim 1, wherein: During the process of training the variational autoencoder network model, generate a training set according to the labeled data and the unlabeled data; Input the training set into the variational autoencoder network model, perform a forward propagation once, and then train the model through the backpropagation algorithm to obtain the predicted class, boundary prediction values, true class labels, and boundaries.
4. The bolt defect detection method based on semi - supervised learning and prior knowledge embedding strategy according to claim 3, characterized in that: During the process of training the model through the backpropagation algorithm, the model is trained with the labeled data to obtain the classification loss and the bounding box regression loss; The equation expression of the classification loss is: The equation expression of the bounding box regression loss is: p i Indicates the probability that the i-th anchor is predicted as the true label. It is 1 for positive samples and 0 for negative samples, t i Indicates the bounding box regression parameters for predicting the i-th anchor. Indicates the true bounding box regression parameters corresponding to the i-th anchor. R is the LOSS CIOU Loss function.
5. The bolt defect detection method based on semi - supervised learning and prior knowledge embedding strategy according to claim 4, characterized in that: During the process of training the model by the backpropagation algorithm, LOSS CIOU The equation expression of the loss function is: Among them, IOU is the ratio of the intersection and union of the predicted bounding box and the ground truth bounding box, b, b gt are the center points of the predicted bounding box and the ground truth bounding box respectively, ρ 2 (b, b gt ) represents the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box, c represents the diagonal distance of the smallest closed region that can simultaneously contain the predicted bounding box and the ground truth bounding box, α is a parameter used to balance the ratio, and v is used to measure the ratio consistency between the anchor box and the predicted box.
6. The bolt defect detection method based on semi - supervised learning and prior knowledge embedding strategy according to claim 5, characterized in that: During the process of training the model through the backpropagation algorithm, by limiting the distance between the output features of each convolutional layer, the unlabeled data is used for model training, where the feature matching loss between the original image and the corresponding synthesized image is used to constrain the model training process: Among them, N cls represents the number of samples in a batch, and N reg represents the number of anchor positions.
7. The bolt defect detection method based on semi - supervised learning and prior knowledge embedding strategy according to claim 6, characterized in that: During the process of using the unlabeled data for model training, the process of model training is expressed as: Among them, T represents the number of layers for feature extraction in the discriminator, and D i represents the extracted feature, and N i represents the number of features extracted by the i-th layer discriminator network.
8. The bolt defect detection method based on semi - supervised learning and prior knowledge embedding strategy according to claim 7, characterized in that: A bolt defect detection system for implementing the bolt defect detection method, comprising: A data acquisition module, configured to collect bolt images of different components and establish a bolt detection data set for different types of defects; A data processing module, configured to perform feature processing on the bolt detection data set based on the prior knowledge embedding strategy to generate a sample data set with sample - correlation features, where The sample - correlation features are used to represent that the features between samples are correlated; A defect recognition module, configured to construct a variational auto - encoder network model, which is composed of a batch normalization unit, a graph convolutional neural network unit, and a convolutional unit; based on the variational auto - encoder network model, through semi - supervised learning, the sample data set is input into the variational auto - encoder network model for training to construct a bolt defect detection model, and the bolt defect detection model is used to identify the bolt defect types of the bolt images to be detected.
Citation Information
Patent Citations
Method for automatically labeling appearance defect images of industrial products based on semi-supervised learning
CN111899254A
Target detection method, device and storage medium based on semi-supervised learning
CN112580684B
Active crowdsourcing image learning method based on semi-supervised variational auto-encoder
CN112990385A