Insulator defect detection method based on deep neural network and saliency
Through deep neural networks and significance detection methods, combined with image preprocessing, feature fusion and multi-scale feature pyramids, the accuracy and universality of insulator defect detection are solved, and efficient and accurate identification and positioning of insulator defects are achieved.
Patent Information
- Application Number
- CN202510325869.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the detection effect of insulator defects is limited by the accuracy of feature extraction operators, which is poor in popularity, and the detection accuracy of deep learning methods is greatly affected by training data and model network architecture.
Deep neural networks and significance detection methods are adopted, including image denoising preprocessing, standardized preprocessing, visual significance feature construction, visible light image and significance feature fusion, multi-scale feature pyramid construction, insulator defect classification recognition and positioning regression, and model training is improved using cross-entropy loss function.
It realizes fast, automated and accurate insulator defect detection, improves the detection generalization ability and recognition efficiency, and adapts to the detection effect under different scales and complex environments.
Smart Images

Figure CN120259222A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to image recognition technology, and in particular to an insulator defect detection method based on deep neural network and significance. Background Art
[0002] Transmission lines are important links in power transmission. The safe and normal operation of transmission lines is related to the safety of power grids. Since transmission lines are usually long and the environment they are in is complex and changeable, insulators and other components in the lines are exposed to the natural environment for a long time, which is prone to failures such as insulator self-explosion and string loss. In terms of the detection of insulator defects in transmission lines, the traditional method is to conduct line inspections manually, which is labor-intensive, inefficient and the effect cannot be guaranteed. With the rapid development of drone technology and computer vision technology, insulator defect detection based on drone aerial images has become a key research direction. The existing technology proposes to construct feature extraction operators by using image processing technologies such as mathematical morphology and edge detection to detect insulator missing and burst faults. The effect of this method is limited by the accuracy of the feature extraction operator and its universality is poor. The existing technology proposes to use deep learning methods to train a large amount of data to realize automatic detection of insulator defects. The accuracy of this method is greatly affected by the training data and the model network architecture. Summary of the invention
[0003] The technical problem to be solved by the present invention is: to provide an insulator defect detection method based on deep neural network and significance, so as to solve the technical problems that the insulator defect detection effect of the prior art is limited by the accuracy of feature extraction operators and its universality is poor; a large amount of data training is carried out using deep learning methods to realize automatic detection of insulator defects, and the accuracy of this method is greatly affected by the training data and the model network architecture.
[0004] Technical solution of the present invention:
[0005] An insulator defect detection method based on deep neural network and significance, the method comprising:
[0006] Step 1: preprocessing the input inspection transmission line image, including image denoising preprocessing and image standardization preprocessing;
[0007] Step 2: construct visual saliency features on the preprocessed image;
[0008] Step 3: Fusion of visible light image and salient features;
[0009] Step 4: Construct a convolutional neural network for bottom-up feature extraction;
[0010] Step 5: Classification and identification of insulator defects;
[0011] Step 6, insulator defect location regression.
[0012] The methods for preprocessing image denoising include:
[0013] Assume the original input image is f(x, y). First, define the convolution filter template as:
[0014]
[0015] Perform a convolution operation on the original input image f(x, y) using the convolution filter template G(x, y) to obtain the image h(x, y) after convolution denoising:
[0016] h(x, y) = f(x, y) * G(x, y);
[0017] The methods for preprocessing image normalization include:
[0018] Take the image h(x, y) obtained after preprocessing denoising as the input image, and calculate the normalized gray histogram p(r k ):
[0019] p(r k ) = n k / MN
[0020] In the formula, r k represents different gray levels in the image, n k represents the total number of pixels with gray level r k in the image, and M and N represent the number of rows and columns of the image matrix; calculate the cumulative histogram c(r k ) of the image based on the normalized gray histogram p(r k ):
[0021]
[0022] Perform histogram equalization processing based on the obtained cumulative histogram c(r k ), that is, transform the histogram of the image to be approximately evenly distributed in the entire gray level interval to enhance the contrast of the image:
[0023]
[0024] In the formula, s k represents the gray level of the image after equalization, T() represents the equalization transformation function, r k represents the k-th gray level value, the value range of k is 0, 1,..., L - 1, and (L - 1) represents the maximum gray level value of the image, n k represents the number of pixels with gray level equal to r k in the image;
[0025] Through histogram equalization transformation, the gray level of the image changes from the original r k to s k ;
[0026] Select an image with moderate brightness and contrast, use the histogram of the selected image as the specified histogram, and perform the same histogram equalization processing:
[0027] v q = B(z q )
[0028] In the formula, v q represents the gray level of the output equalized image, B() represents the equalization transformation function, and z q represents the gray value of the k-th level of this image;
[0029] Since:
[0030] s k = v q
[0031] Therefore:
[0032] v q = B(z q )
[0033] Performing an inverse transformation on the above formula gives:
[0034] z q = B -1 (v q ) = B -1 (S k ) = B -1 (T(r k )))
[0035] Using the above method, taking histogram equalization as the intermediate result, the relationship mapping between the original gray level r k and the specified gray level z q is completed, realizing the transformation of the original image histogram into the specified histogram form and achieving the standardized preprocessing of the image.
[0036] The methods for constructing visual saliency features for the preprocessed image include:
[0037] Taking the preprocessed image as the input image, and obtaining a Gaussian image pyramid by downsampling the three channels of the image:
[0038]
[0039] Wherein, r(x, y), g(x, y), and b(x, y) respectively represent the three channels of the input image, G(x, y, σ) represents the Gaussian template, σ is the corresponding scale range, and the value range is 0 - 8. By performing convolution operations with Gaussian templates of 9 scales, each channel obtains a 9-layer Gaussian pyramid; r(x, y, σ), g(x, y, σ), and b(x, y, σ) respectively represent the images of different channels of the obtained Gaussian pyramid;
[0040] Set the combination method of pyramid images at different levels as (p, q), where p ∈ {2, 3, 4},
[0041] q = p + δ, δ ∈ {3, 4};
[0042] A total of six different scale combinations of {2, 5}, {2, 6}, {3, 6}, {3, 7}, {4, 7}, {4, 8} are obtained;
[0043] Based on the Gaussian pyramid, construct the image grayscale pyramid:
[0044]
[0045] Wherein, I(x, y, σ) represents the grayscale pyramid image;
[0046] Perform scale combination on the grayscale pyramid according to the six combination methods (p, q), and extract the corresponding grayscale features, obtaining a total of 6 grayscale visual feature maps:
[0047]
[0048] Among them, I(x, y, p) represents the grayscale pyramid image of the p-th layer, I(x, y, q) represents the grayscale pyramid image of the q-th layer, represents cross-scale operation, and I(p, q) represents the obtained grayscale visual feature map;
[0049] Based on the Gaussian pyramid, construct the image color pyramid:
[0050]
[0051] Among them, R(x, y, σ), G(x, y, σ), B(x, y, σ), and Y(x, y, σ) respectively represent the pyramid images of the red, green, blue, and yellow channels;
[0052] Perform cross-scale operation combination on the pyramid images of each color channel in the following manner:
[0053]
[0054] For each (p, q) combination, two feature maps are generated. Since (p, q) represents six different scale combinations, a total of 12 color visual feature maps are obtained;
[0055] Based on the Gaussian pyramid, an image orientation pyramid is constructed:
[0056] O(σ,θ) = f(x,y) * G(x,y,σ,θ)
[0057] where θ ∈ {0°, 45°, 90°, 135°}, and G(x,y,σ,θ) represents the Gabor function, and its specific expression is as follows:
[0058]
[0059] Since θ has four values, for each (p, q) combination, four feature maps are obtained. Since (p, q) represents six different scale combinations, a total of 24 orientation visual feature maps can be obtained:
[0060]
[0061] Salience map generation:
[0062] Normalize all input low-level visual feature maps to a unified range, as follows:
[0063] I'(p,q) = I(p,q) / 255
[0064] RG'(p,q) = RG(p,q) / 255
[0065] BY'(p,q) = BY(p,q) / 255
[0066] O'(p,q,θ) = O(p,q,θ) / 255
[0067] Fuse the low-level visual feature maps by cross-scale addition:
[0068]
[0069] In the formula: represents cross-scale addition, represents the fused grayscale visual feature, represents the fused color visual feature, represents the fused orientation visual feature. Cross-scale addition is all completed on the basis of scale 4; then fuse the fused grayscale visual feature, color visual feature, and orientation visual feature to obtain the salience map:
[0070]
[0071] Among them, S(x, y) represents the saliency map.
[0072] The method of fusing the visible light image and the saliency features includes: superimposing the saliency map as the fourth channel on the three channels of the original visible light image:
[0073]
[0074] In the formula, f R (x, y), f G (x, y) and f B (x, y) respectively represent the three channels of the original input image f(x, y), represents multi-channel superimposed fusion, and img represents the image result after fusing the original visible light image and the saliency feature image.
[0075] The method of constructing a convolutional neural deep network for bottom-up feature extraction includes: The parameters of the convolutional neural deep network model include:
[0076]
[0077]
[0078] Based on the convolutional neural deep network, top-down connections and lateral connections are made to the feature maps. A 1*1 convolution is performed on the Conv5 layer to obtain the M5 layer, and the M5 layer is added to the 1*1 convolution result of the Conv4 layer to obtain the M4 layer; similarly, the M4 layer is added to the 1*1 convolution result of the Conv3 layer to obtain the M3 layer; the M3 layer is added to the 1*1 convolution result of the Conv2 layer to obtain the M2 layer; then 3*3 convolutions are performed on the M2, M3, M4, and M5 layers respectively to obtain the P2, P3, P4, and P5 layers, which are the final feature pyramid;
[0079] At each layer of the feature pyramid, regression boxes with three aspect ratios of 1:2, 2:1, and 1:1 are set. For each aspect ratio of the regression box, three different area sizes are set respectively to cover targets of various scales. Therefore, 9 regression boxes are obtained for each layer of the feature pyramid.
[0080] The method for classifying and identifying insulator defects includes: constructing a sub-network for classifying and identifying insulator defects, connecting it to the P2, P3, P4, and P5 layers of the feature pyramid respectively, performing convolution operations on the input feature map using 4 convolutional kernels of 3*3*256, activating each convolution using the Relu activation function, and obtaining a feature map of size W*H*256; then performing convolution operations again using a convolutional kernel of size 3*3*(9*2), where 9 represents the number of regression boxes and 2 represents the number of recognition categories, here only including foreground and background, and finally using the sigmoid function to output the classification results of each regression box.
[0081] The method for regression positioning of insulator defects includes: constructing a sub-model for regression positioning of insulator defects to regression and position each regression box, connecting the regression sub-model to the P2, P3, P4, and P5 layers of the feature pyramid respectively, first performing convolution operations on the input feature map using 4 convolutional kernels of 3*3*256, and activating using the Relu function after each convolution to obtain a feature map of size W*H*256; then performing convolution operations again using a convolutional kernel of size 3*3*(9*4), where 9 represents the number of regression boxes and 4 represents the offset of the four corner points of the regression box.
[0082] The beneficial effects of the present invention:
[0083] The present invention eliminates the influence of factors such as image noise and inconsistent image brightness by performing image denoising preprocessing and image normalization preprocessing. By constructing visual saliency features to generate a saliency map, the position of the insulator is initially determined. The visible light image and the saliency features are fused by means of multi-channel superposition. The fused image not only has the advantages of rich feature information such as texture and color of the visible light image but also has the advantage of prominent insulator targets in the saliency map. By constructing a multi-scale feature pyramid, it can effectively cope with the problem of multi-scale changes in images caused by inconsistent shooting distances in insulator defect detection. In addition, a sub-network for classifying and identifying insulator defects and a sub-network for regression positioning of insulator defects are respectively constructed and connected to the feature pyramid to achieve the classification and identification of insulator defects and the regression positioning of the image position. To avoid the situation where the training accuracy of the model is very high due to unbalanced positive and negative samples but the actual test accuracy is low, the cross-entropy loss function is improved, and then the saliency features are used as a supplement to the visible light image and jointly input into the deep learning model for iterative training, and finally the recognition result of the insulator defect image is obtained. Compared with the existing manual visual interpretation or defect recognition methods based on traditional image processing, this method is faster, has a higher degree of automation, is more accurate and complete in recognition, has stronger generalization ability, and higher efficiency, and can still maintain a good recognition effect for insulator defect image data of different scales and complex environments.
[0084] The problems in the prior art that the detection effect of insulator defects is limited by the accuracy of feature extraction operators and its universality is poor; and that the accuracy of the method of using deep learning to train a large amount of data to achieve automatic detection of insulator defects is greatly affected by training data and model network architecture are solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 It is a schematic diagram of the feature pyramid network structure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0086] A method for detecting insulator defects based on a deep neural network and saliency, which comprises the following steps:
[0087] Step 1, image preprocessing:
[0088] By establishing an image preprocessing model, the present invention can automatically preprocess the input images of the power transmission line inspected by the unmanned aerial vehicle. The model mainly includes two parts: image denoising preprocessing and image normalization preprocessing. The specific principle is as follows.
[0089] (1) Image denoising preprocessing
[0090] Due to the influence of factors such as the shooting sensor device and the external environmental conditions, the original images of the power transmission line inspected by the unmanned aerial vehicle will inevitably have some noise problems. In order to facilitate the subsequent better realization of automatic detection of insulator defects on the power transmission line, the original images are subjected to denoising preprocessing.
[0091] Assume that the original input image is f(x, y). First, define the convolutional filtering template as follows:
[0092]
[0093] Use the convolutional filtering template G(x, y) to perform a convolution operation on the original input image f(x, y) to obtain the image h(x, y) after convolution denoising:
[0094] h(x, y) = f(x, y) * G(x, y)
[0095] Through the above denoising preprocessing, the influence of Gaussian noise existing in the images can be effectively eliminated, and it can be better applied to various original inspection images.
[0096] (2) Image normalization preprocessing
[0097] Affected by factors such as the shooting sensor device, shooting angle, shooting illumination conditions, and other external shooting conditions, the original UAV inspection transmission line images may have situations such as inconsistent image brightness or large differences in brightness between different images, which may have an adverse impact on subsequent construction of a deep neural network model for insulator defect recognition. To eliminate the influence of this adverse factor, an image standardization preprocessing model is constructed.
[0098] Take the image h(x,y) obtained after denoising preprocessing as the input image, and calculate its normalized gray histogram p(r k ):
[0099] p(r k ) = n k / MN
[0100] In the formula, r k represents different gray levels in the image, n k represents the total number of pixels with gray level r k in this image, M and N represent the number of rows and columns of the image matrix, and p(r k ) represents the frequency of occurrence of gray level r k in the image. The abscissa of the normalized gray histogram describes the gray values of each level, and the ordinate is the statistical frequency of occurrence of each level of gray value in the image.
[0101] Based on the normalized gray histogram p(r k ), calculate the cumulative histogram c(r k ):
[0102]
[0103] The cumulative histogram c(r k ) can specifically describe the cumulative distribution of gray levels in the image, that is, statistically calculate the frequency of less than or equal to this gray value.
[0104] Based on the obtained cumulative histogram c(r k ), perform histogram equalization processing, that is, transform the histogram of the image to be approximately evenly distributed in the entire gray interval to enhance the contrast of the image:
[0105]
[0106] In the formula, s k represents the gray level of the equalized image, T() represents the equalization transformation function, r k represents the k-th level gray value, the value range of k is 0, 1,..., L - 1, (L - 1) represents the maximum gray level value of the image, and n k represents the number of pixels in the image with gray level equal to r kThe number of pixels.
[0107] Through histogram equalization transformation, the gray level of the image changes from the original r k to s k .
[0108] In addition, select an image with moderate brightness and contrast. The image selection criteria are as follows: 1) Avoid the histogram tending to the high gray level side, resulting in an over-bright image; 2) Avoid the histogram tending to the low gray level side, resulting in an over-dark image; 3) Avoid the histogram being narrow and concentrated in the middle of the gray level, resulting in a low-contrast image.
[0109] Use the histogram of the selected image as the specified histogram and perform the same histogram equalization processing on it:
[0110] v q = B(z q )
[0111] In the formula, v q represents the gray level of the output equalized image, B() represents the equalization transformation function, and z q represents the gray value of the k-th level of this image.
[0112] Since:
[0113] s k = v q
[0114] Therefore:
[0115] v q = B(z q )
[0116] Performing an inverse transformation on the above formula gives:
[0117] z q = B -1 (v q ) = B -1 (S k ) = B -1 (T(r k ))
[0118] Using histogram equalization as an intermediate result in the above way, complete the relationship mapping between the original gray level r k and the specified gray level z q to achieve the transformation of the original image histogram to the specified histogram shape and realize the standardized preprocessing of the image.
[0119] Step 2: Construction of visual saliency features:
[0120] Based on the preprocessing results of the UAV inspection images in the above steps, visual saliency features are constructed for the preprocessed images.
[0121] (1) Visual feature construction
[0122] Taking the image preprocessed by the above steps as the input image, Gaussian image pyramids are obtained by downsampling the three channels of this image:
[0123]
[0124] In the formula, r(x,y), g(x,y), and b(x,y) respectively represent the three channels of the input image, G(x,y,σ) represents the Gaussian template, σ is the corresponding scale range, and its value range is 0 - 8. By performing convolution operations with Gaussian templates of 9 scales, 9 layers of Gaussian pyramids can be obtained for each channel. r(x,y,σ), g(x,y,σ), and b(x,y,σ) respectively represent the images of different channels of the obtained Gaussian pyramids.
[0125] Set the combination method of pyramid images at different levels as (p,q), where p ∈ {2, 3, 4} and q = p + δ, δ ∈ {3, 4}.
[0126] A total of six different scale combinations of {2,5}, {2,6}, {3,6}, {3,7}, {4,7}, {4,8} can be obtained.
[0127] 1) Grayscale visual feature extraction
[0128] Based on the Gaussian pyramid, an image grayscale pyramid is constructed:
[0129]
[0130] In the formula, I(x,y,σ) represents the grayscale pyramid image.
[0131] According to the six combination methods (p,q) described above, the grayscale pyramid is combined in scale, and the corresponding grayscale features are extracted. A total of 6 grayscale visual feature maps can be obtained:
[0132]
[0133] Among them, I(x,y,p) represents the image of the p-th layer grayscale pyramid, I(x,y,q) represents the image of the q-th layer grayscale pyramid, represents cross-scale operation, and I(p,q) represents the obtained grayscale visual feature map.
[0134] 2) Color visual feature extraction
[0135] Based on the Gaussian pyramid, an image color pyramid is constructed:
[0136]
[0137] Among them, R(x, y, σ), G(x, y, σ), B(x, y, σ), and Y(x, y, σ) represent the pyramid images of the red, green, blue, and yellow channels respectively.
[0138] Perform cross-scale operation combination on the pyramid images of each color channel in the following manner:
[0139]
[0140] For each (p, q) combination, the above two feature maps are generated. Since (p, q) represents 6 different scale combinations, a total of 12 color visual feature maps can be obtained.
[0141] 3) Direction visual feature extraction
[0142] Based on the Gaussian pyramid, construct an image direction pyramid:
[0143] O(σ, θ) = f(x, y) * G(x, y, σ, θ)
[0144] Among them, θ ∈ {0°, 45°, 90°, 135°}, G(x, y, σ, θ) represents the Gabor function, and its specific expression is as follows:
[0145]
[0146] Since θ has four values, for each (p, q) combination, 4 feature maps will be obtained. Since (p, q) represents 6 different scale combinations, a total of 24 direction visual feature maps can be obtained:
[0147]
[0148] (2) Salience map generation
[0149] Normalize all the input low-level visual feature maps to a unified range, as follows:
[0150] I'(p, q) = I(p, q) / 255
[0151] RG'(p, q) = RG(p, q) / 255
[0152] BY'(p, q) = BY already(p, q) / 255
[0153] O'(p, q, θ) = O(p, q, θ) / 255
[0154] Fuse the underlying visual feature maps by cross-scale addition:
[0155]
[0156] Among them, denotes cross-scale addition, represents the fused grayscale visual feature, represents the fused color visual feature, represents the fused orientation visual feature. The cross-scale addition is all completed on the basis of scale 4. Then fuse the fused grayscale visual feature, color visual feature and orientation visual feature to obtain the saliency map:
[0157]
[0158] Among them, S(x, y) represents the saliency map.
[0159] Step 3: Fuse the visible light image and the saliency feature:
[0160] Fuse the visible light image and the saliency feature by multi-channel superposition, that is, use the saliency map as the fourth channel and superimpose it with the three channels of the original visible light image:
[0161]
[0162] In the formula, f R (x, y), f G (x, y) and f B (x, y) respectively represent the three channels of the original input image f(x, y), represents multi-channel superposition fusion, and img represents the image result after fusing the original visible light image and the saliency feature image.
[0163] Step 4: Construct a multi-scale feature pyramid:
[0164] First, construct a convolutional neural deep network for bottom-up feature extraction. The specific parameters of the model are shown in the following table:
[0165]
[0166]
[0167] Then, based on the convolutional neural deep network, connect its feature maps top-down and laterally to construct a feature pyramid network:
[0168] Perform a 1*1 convolution on the Conv5 layer to obtain the M5 layer, add the M5 layer to the 1*1 convolution result of the Conv4 layer to obtain the M4 layer. Similarly, add the M4 layer to the 1*1 convolution result of the Conv3 layer to obtain the M3 layer. Add the M3 layer to the 1*1 convolution result of the Conv2 layer to obtain the M2 layer. Then perform 3*3 convolutions on the M2, M3, M4, and M5 layers respectively to obtain the P2, P3, P4, and P5 layers, which are the final feature pyramids.
[0169] Set regression boxes with three aspect ratios of 1:2, 2:1, and 1:1 on each layer of the feature pyramid. For each ratio of regression boxes, set three different area sizes respectively, so as to cover targets of various scales. Therefore, 9 regression boxes can be obtained for each layer of the feature pyramid.
[0170] Step 5, Insulator defect classification and recognition:
[0171] Construct an insulator defect classification and recognition sub-network, connect it to the P2, P3, P4, and P5 layers of the feature pyramid respectively, perform convolution operations on the input feature map using 4 convolutional kernels of 3*3*256, and activate each convolution using the Relu activation function to obtain a feature map of size W*H*256. Then perform convolution operations again using a convolutional kernel of size 3*3*(9*2), where 9 represents the number of regression boxes and 2 represents the number of recognition categories. Here, it only includes foreground (insulator defects) and background. Finally, use the sigmoid function to output the classification results of each regression box. The structure of the classification sub-network is as follows:
[0172]
[0173] Step 6, Insulator defect location regression:
[0174] Construct an insulator defect location regression sub-model for regressing and locating the position of each regression box. The regression sub-model is connected to the P2, P3, P4, and P5 layers of the feature pyramid respectively. First, perform convolution operations on the input feature map using 4 convolutional kernels of 3*3*256, and activate using the Relu function after each convolution to obtain a feature map of size W*H*256. Then perform convolution operations again using a convolutional kernel of size 3*3*(9*4), where 9 represents the number of regression boxes and 4 represents the offset of the four corner points of the regression box. The structure of the insulator defect location regression sub-network is as follows:
[0175]
[0176] Step 7, Model loss function setting:
[0177] The common form of the cross-entropy loss function is as follows:
[0178] CE(p, y) = CE(p t ) = -ln(p t )
[0179] where p t is the probability predicted by the model to be true, and y is the class label.
[0180] To avoid the situation where the model training accuracy is very high due to the imbalance between positive and negative samples (the prediction of the model for negative samples overwhelms the prediction for positive samples), but the actual test accuracy is relatively low, the cross-entropy loss function is improved as follows:
[0181] FL(p t ) = -α t (1 - p t ) γ ln(p t )
[0182] where α t represents the balance factor, whose value range is [0, 1], and γ represents the adjustment factor, whose value range is [0, 5].
Claims
1. An insulator defect detection method based on deep neural network and saliency, characterized in that: The method includes: Step 1: Preprocess the input inspection transmission line image, including image denoising preprocessing and image normalization preprocessing; Step 2: Construct visual saliency features for the preprocessed image; Step 3: Fuse the visible light image with the saliency features; Step 4: Construct a convolutional neural deep network for bottom-up feature extraction; Step 5: Classify and identify insulator defects; Step 6: Locate and regress insulator defects.
2. The method for detecting insulator defects based on deep neural network and saliency according to claim 1, wherein: The method for image denoising preprocessing includes: Assume the original input image is f(x, y), and first define the convolutional filtering template as: Use the convolutional filtering template G(x, y) to perform a convolution operation on the original input image f(x, y) to obtain the denoised image h(x, y) after convolution: h(x, y) = f(x, y) * G(x, y); The method for image normalization preprocessing includes: Take the image h(x, y) obtained after denoising preprocessing as the input image, and calculate the normalized gray histogram p(r k ): p(r k ) = n k / MN where r k represents different gray levels in the image, and n k represents the total number of pixels with gray level r k in the image, and M and N represent the number of rows and columns of the image matrix; based on the normalized gray histogram p(r k ), the cumulative histogram c(r k ) of the image is calculated as follows: Based on the obtained cumulative histogram c(r k ), histogram equalization is carried out, that is, the histogram of the image is transformed to be approximately evenly distributed in the entire gray level interval to enhance the contrast of the image: where s k represents the gray level of the equalized image, T() represents the equalization transformation function, r k represents the k-th gray level value, the value range of k is 0, 1,..., L - 1, and (L - 1) represents the maximum gray level value of the image, n k represents the number of pixels in the image whose gray level is equal to r k ; Through histogram equalization transformation, the gray level of the image changes from the original r k to s k ; Select an image with moderate brightness and contrast, use the histogram of the selected image as the specified histogram, and perform the same histogram equalization processing: v q = B(z q ) where, v q represents the gray level of the output equalized image, B() represents the equalization transformation function, and z q represents the gray value of the k-th level of this image; Since: s k = v q Therefore: v q = B(z q ) Performing an inverse transformation on the above formula gives: z q = B -1 (v q ) = B -1 (S k ) = B -1 (T(r k )) Using histogram equalization as an intermediate result in the above manner to complete the original gray level r k and the gray level z after specification q The relationship mapping between them is realized, the transformation of the original image histogram to the specified histogram form is realized, and the standardized preprocessing of the image is realized.
3. The method for detecting insulator defects based on deep neural network and saliency according to claim 1, wherein: The method for constructing visual saliency features for the preprocessed image includes: Use the preprocessed image as the input image, and obtain a Gaussian image pyramid by downsampling the three channels of the image: Where r(x, y), g(x, y), and b(x, y) respectively represent the three channels of the input image, G(x, y, σ) represents the Gaussian template, σ is the corresponding scale range, and the value range is 0 - 8. By performing a convolution operation with Gaussian templates of 9 scales, each channel obtains 9 layers of Gaussian pyramids; r(x, y, σ), g(x, y, σ), and b(x, y, σ) respectively represent the images of different channels of the obtained Gaussian pyramid; Set the combination method of pyramid images at different levels as (p, q), where p ∈ {2, 3, 4}, q = p + δ, δ ∈ {3, 4}; A total of six different scale combinations of {2, 5}, {2, 6}, {3, 6}, {3, 7}, {4, 7}, {4, 8} are obtained; Based on the Gaussian pyramid, construct an image gray pyramid: Where I(x, y, σ) represents the gray pyramid image; Perform scale combination on the gray pyramid according to the six combination methods (p, q), and extract the corresponding gray features, obtaining a total of 6 gray visual feature maps; Among them, I(x, y, p) represents the grayscale pyramid image of the p-th layer, and I(x, y, q) represents the grayscale pyramid image of the q-th layer. represents cross-scale operation, and I(p, q) represents the obtained grayscale visual feature map; Based on the Gaussian pyramid, construct an image color pyramid: Where R(x, y, σ), G(x, y, σ), B(x, y, σ), and Y(x, y, σ) respectively represent the pyramid images of the red, green, blue, and yellow channels; Perform cross-scale operation combination on the pyramid images of each color channel in the following manner: For each (p, q) combination, two feature maps will be generated. Since (p, q) represents 6 different scale combinations, a total of 12 color visual feature maps are obtained; Based on the Gaussian pyramid, construct an image orientation pyramid: O(σ,θ) = f(x,y) * G(x,y,σ,θ) where θ ∈ {0°, 45°, 90°, 135°}, and G(x,y,σ,θ) represents the Gabor function, and its specific expression is as follows: Since θ has four values, for each (p,q) combination, 4 feature maps will be obtained. Since (p,q) represents 6 different scale combinations, a total of 24 orientation visual feature maps can be obtained: Salience map generation: Normalize all input low-level visual feature maps to a unified range, as follows: I'(p,q) = I(p,q) / 255 RG'(p,q) = RG(p,q) / 255 BY'(p,q) = BY(p,q) / 255 O'(p,q,θ) = O(p,q,θ) / 255 Fuse the low-level visual feature maps by cross-scale addition: In the formula: represents cross-scale addition, represents the fused grayscale visual feature, represents the fused color visual feature, represents the fused orientation visual feature. The cross-scale addition is all completed on the basis of scale 4; then the fused grayscale visual feature, color visual feature, and orientation visual feature are fused to obtain the saliency map: where S(x,y) represents the salience map.
4. A method for detecting insulator defects based on a deep neural network and saliency according to claim 1, characterized in that: The method of fusing the visible light image and the salience feature includes: using the salience map as the fourth channel to stack with the three channels of the original visible light image: where f R (x, y), f G (x, y) and f B (x, y) represent the three channels of the original input image f(x, y) respectively. represents multi-channel superposition and fusion, and img represents the image result after fusing the original visible light image and the saliency feature image.
5. A method for detecting insulator defects based on a deep neural network and saliency according to claim 1, characterized in that: The method of constructing a convolutional neural deep network for bottom-up feature extraction includes: the model parameters of the convolutional neural deep network include: Based on the convolutional neural deep network, perform top-down connections and lateral connections on the feature maps. Perform 1*1 convolution on the Conv5 layer to obtain the M5 layer, add the result of the 1*1 convolution of the M5 layer and the Conv4 layer to obtain the M4 layer; similarly, add the result of the 1*1 convolution of the M4 layer and the Conv3 layer to obtain the M3 layer; add the result of the 1*1 convolution of the M3 layer and the Conv2 layer to obtain the M2 layer; then perform 3*3 convolution on the M2, M3, M4, and M5 layers respectively to obtain the P2, P3, P4, and P5 layers, which is the final feature pyramid; Set three aspect ratio regression boxes of 1:2, 2:1, and 1:1 for each layer of the feature pyramid. For each aspect ratio of the regression box, set three different area sizes to cover targets of various scales. Therefore, 9 regression boxes are obtained for each layer of the feature pyramid.
6. The insulator defect detection method based on deep neural network and saliency according to claim 5, characterized in that: The method for insulator defect classification and recognition includes: constructing an insulator defect classification and recognition sub-network, connecting it to the P2, P3, P4, and P5 layers of the feature pyramid respectively, performing convolution operations on the input feature map using 4 3*3*256 convolutional kernels, activating each convolution using the Relu activation function to obtain a feature map of size W*H*256; then performing convolution operations again using a 3*3*(9*2) convolutional kernel, where 9 represents the number of regression boxes and 2 represents the number of recognition categories, here only including foreground and background, and finally using the sigmoid function to output the classification results of each regression box.
7. A method for detecting insulator defects based on a deep neural network and saliency according to claim 6, characterized in that: The method for insulator defect location regression includes: constructing an insulator defect location regression sub-model for regressing and locating the position of each regression box. The regression sub-model is respectively connected to the P2, P3, P4, and P5 layers of the feature pyramid. First, the input feature map is convolved using four 3*3*256 convolutional kernels, and after each convolution, the Relu function is used for activation to obtain a feature map of size W*H*256; then, convolution operation is performed again using a convolutional kernel of size 3*3*(9*4), where 9 represents the number of regression boxes and 4 represents the offsets of the four corner points of the regression box.
Citation Information
Cited By
Insulating material surface defect identification method and system based on visual inspection
CN121305183A