Insulator detection method, model training method, device, equipment and storage medium
By using the target detection model to automatically detect insulators in transmission lines, the problems of high cost and low efficiency of manual inspection are solved, and more efficient and economical insulator detection is achieved.
Patent Information
- Application Number
- CN202411872449.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems of high labor costs and low detection efficiency when conducting manual inspection of insulators in transmission lines.
An insulator detection method is adopted to obtain the image to be detected and input the object detection model, and use the backbone network, the neck network and the head network for feature extraction and fusion to determine the prediction box and detection results of the insulator to be detected.
The small object detection capability of insulators is improved, the ability to distinguish insulators in the detected image is enhanced, the labor cost in the insulator detection process is reduced, and the detection efficiency is improved.
Smart Images

Figure CN120013857A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of insulator defect detection, and in particular relates to an insulator detection method, a model training method, a device, equipment and a storage medium. Background Art
[0002] Insulators are widely used basic components in transmission lines. In transmission lines, the role of insulators is to be placed between different parts to withstand mechanical action and voltage. The safety and stability of transmission lines are closely related to the quality of insulators. In the application scenarios where loads with large dynamic fluctuations in power quality such as electric vehicles, charging piles, and evolving energy storage equipment are connected to the power network, higher requirements are placed on insulator defect detection.
[0003] At present, it is generally necessary to manually inspect the insulators in the transmission lines to find defects in the insulators in the transmission lines and ensure the quality and availability of the insulators in the transmission lines.
[0004] However, most high-voltage transmission lines are usually built in deep mountains, rivers and uninhabited areas, resulting in high labor costs and low detection efficiency in the manual inspection operation mode in related technologies. Summary of the invention
[0005] The embodiments of the present invention provide an insulator detection method, a model training method, an apparatus, a device and a storage medium to solve the problems of high labor cost and low detection efficiency in the process of manual inspection of insulators in transmission lines in the related art.
[0006] In order to solve the above problems, in a first aspect, an embodiment of the present invention discloses an insulator detection method, the method comprising:
[0007] Acquire an image to be detected, and input the image to be detected into a target detection model; the image to be detected includes an insulator to be detected; the target detection model includes a backbone network, a neck network and a head network connected in sequence;
[0008] Utilizing the first single-channel downsampler, the spatial pyramid pooling layer, and the second single-channel downsampler in the backbone network, obtaining a first extracted feature map corresponding to the image to be detected;
[0009] Using the neck network, performing feature fusion processing on the first extracted feature map to obtain a first fused feature map;
[0010] Determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the first fused feature map by using an upsampling layer, a similarity-aware activation layer, and a convolutional normalization layer in the head network;
[0011] The first prediction frame and the first detection result output by the target detection model are obtained, and the target detection result of the insulator to be detected is determined according to the first prediction frame and the first detection result.
[0012] In a second aspect, an embodiment of the present invention discloses a model training method, the method comprising:
[0013] Acquire an image data set corresponding to the insulator; the image data set includes training image samples, target frames corresponding to the training image samples, and label data corresponding to the target frames;
[0014] In each round of training, the training image sample is used to train the detection model to be trained, and a second prediction box corresponding to the training image sample and a second detection result corresponding to the second prediction box are obtained;
[0015] Determine a loss value corresponding to the to-be-trained detection model according to the second prediction box and the second detection result, as well as the target box and the label data;
[0016] The model parameters of the detection model to be trained are adjusted according to the loss value, and the next round of training is performed until the loss value meets the preset conditions, thereby obtaining the target detection model as described in any of the above items.
[0017] In a third aspect, an embodiment of the present invention discloses an insulator detection device, the device comprising:
[0018] A first acquisition module is used to acquire an image to be detected and input the image to be detected into a target detection model; the image to be detected includes an insulator to be detected; the target detection model includes a backbone network, a neck network and a head network connected in sequence;
[0019] A second acquisition module is used to acquire a first extracted feature map corresponding to the image to be detected by using the first single-channel downsampler, the spatial pyramid pooling layer and the second single-channel downsampler in the backbone network;
[0020] An extraction module, used to perform feature fusion processing on the first extracted feature map using the neck network to obtain a first fused feature map;
[0021] A first determination module, configured to determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the first fused feature map by using an upsampling layer, a similarity-aware activation layer, and a convolutional normalization layer in the head network;
[0022] The second determination module is used to obtain the first prediction frame and the first detection result output by the target detection model, and determine the target detection result of the insulator to be detected based on the first prediction frame and the first detection result.
[0023] In a fourth aspect, an embodiment of the present invention discloses a model training device, the device comprising:
[0024] A third acquisition module is used to acquire an image data set corresponding to the insulator; the image data set includes a training image sample, a target frame corresponding to the training image sample, and label data corresponding to the target frame;
[0025] A training module, used to train the to-be-trained detection model using the training image samples in each round of training, to obtain a second prediction box corresponding to the training image samples and a second detection result corresponding to the second prediction box;
[0026] A third determination module is used to determine the loss value corresponding to the to-be-trained detection model according to the second prediction box and the second detection result, the target box and the label data;
[0027] The training module is also used to adjust the model parameters of the detection model to be trained according to the loss value, and perform the next round of training until the loss value meets the preset conditions to obtain the target detection model as described in any of the above items.
[0028] In a fifth aspect, an embodiment of the present invention further discloses an electronic device, comprising a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions enable the processor to execute the insulator detection method as described above, or execute the model training method as described above.
[0029] An embodiment of the present invention further discloses a readable storage medium. When instructions in the readable storage medium are executed by a processor of an electronic device, the processor is enabled to execute the insulator detection method as described above, or to execute the model training method as described above.
[0030] The embodiments of the present invention include the following advantages:
[0031] The insulator detection method provided by the embodiment of the present invention obtains an image to be detected including the insulator to be detected, and inputs the image to be detected into a target detection model. In the process of obtaining a first extracted feature map corresponding to the image to be detected by using the first single-channel downsampler, the spatial pyramid pooling layer and the second single-channel downsampler in the backbone network, the image to be detected can be downsampled without losing information of the image to be detected based on the first single-channel downsampler and the second single-channel downsampler, thereby improving the detection capability of small target insulators in the image to be detected. Multi-scale context information can be obtained from the image to be detected based on the spatial pyramid pooling layer, thereby improving the accuracy of the first extracted feature value obtained by the backbone network. In the process of obtaining a first extracted feature map corresponding to the image to be detected by using the upsampling layer, the spatial pyramid pooling layer and the second single-channel downsampler in the head network, the image to be detected can be downsampled without losing information of the image to be detected. In the process of determining the first prediction box and the first detection result based on the first fused feature map obtained by the neck network by the similarity-aware activation layer and the convolution normalization layer, the global context information of the first fused feature map is calculated by global attention weighting on the first fused feature map based on the similarity-aware activation layer, thereby improving the head network's ability to distinguish insulators in the image to be detected, and improving the accuracy of the first prediction box and the first detection result determined by the head network, thereby improving the accuracy and reliability of the target detection result of the insulator to be detected determined by the embodiment of the present invention according to the first prediction box and the first detection result. Moreover, the embodiment of the present invention can be automatically performed in a scenario without human participation, thereby reducing the labor cost in the insulator detection process and improving the efficiency of insulator detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0033] Figure 1 is a flowchart of the steps of an insulator detection method provided by an embodiment of the present invention;
[0034] Figure 2 is a schematic diagram of the structure of a target detection model provided by an embodiment of the present invention;
[0035] Figure 3 is a schematic diagram of the structure of a spatial pyramid pooling layer provided by an embodiment of the present invention;
[0036] Figure 4 is a structural schematic diagram of a similarity perception activation layer provided by an embodiment of the present invention;
[0037] Figure 5 is a flowchart of the steps of a model training method provided by an embodiment of the present invention;
[0038] Figure 6 It is a logic block diagram of a model training provided by an embodiment of the present invention;
[0039] Figure 7 is a schematic diagram of the relationship between a second prediction box and a target box provided by an embodiment of the present invention;
[0040] Figure 8 is a logic block diagram of an insulator detection device provided by an embodiment of the present invention;
[0041] Fig. 9 It is a logic block diagram of a model training device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] Method Embodiment
[0044] Reference Figure 1 , shows a flowchart of the steps of an insulator detection method provided by an embodiment of the present invention, and the method may specifically include the following steps:
[0045] Step S101, obtaining an image to be detected, and inputting the image to be detected into a target detection model; the image to be detected includes an insulator to be detected; the target detection model includes a backbone network, a neck network and a head network connected in sequence.
[0046] Step S102: Utilize the first single-channel down-sampler, the spatial pyramid pooling layer and the second single-channel down-sampler in the backbone network to obtain a first extracted feature map corresponding to the image to be detected.
[0047] Step S103: using the neck network, perform feature fusion processing on the first extracted feature map to obtain a first fused feature map.
[0048] Step S104: using the upsampling layer, the similarity-aware activation layer and the convolution normalization layer in the head network, based on the first fused feature map, determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box.
[0049] Step S105: Acquire the first prediction frame and the first detection result output by the target detection model, and determine the target detection result of the insulator to be detected according to the first prediction frame and the first detection result.
[0050] The insulator detection method provided in the embodiment of the present invention can be applied to a first electronic device installed with a target detection model. The first electronic device may include but is not limited to a smart terminal, a computer, a personal digital assistant (PDA), a tablet computer, a vehicle-mounted device, etc.
[0051] In an embodiment of the present invention, the first electronic device can acquire the image to be detected according to a preset acquisition cycle, and input the image to be detected into the target detection model, so as to use the target detection model to obtain the first prediction frame of the image to be detected and the first detection result corresponding to the first prediction frame through steps S102 to S104, and finally determine the target detection result of the insulator to be detected according to the first prediction frame and the first detection result output by the target detection model.
[0052] The insulator to be detected is any insulator in the transmission line; the transmission line includes an image acquisition device, which is used to obtain an image of the insulator to be detected in the transmission line according to a preset acquisition period and transmit the image to the first electronic device; the first electronic device determines the image including the insulator to be detected transmitted by the image acquisition device as the image to be detected.
[0053] It is understandable that the insulator to be inspected in the image to be inspected may be a normal insulator without defects or a defective insulator with defects; the defect types of the insulator may include but are not limited to: polymer dirt, pollution-flashover, glass loss, broken disc, etc. The insulator to be inspected includes the insulator body to be inspected and / or defects on the insulator body to be inspected.
[0054] In addition, the image to be detected includes a complex background in addition to the insulator to be detected. The existence of the complex background in the image to be detected will affect the accuracy of the target detection result of the insulator to be detected.
[0055] In order to improve the target detection result of the insulator to be detected obtained in step S105, the embodiment of the present invention uses the YOLOv7 model as the basic model, and improves the YOLOv7 model to obtain the target detection model in the embodiment of the present invention, and uses the various components in the target detection model to process the image to be detected, and determines the first prediction box and the first detection result from the image to be detected, thereby improving the feature recognition accuracy of the insulator to be detected in the image to be detected, thereby improving the accuracy and reliability of the target detection result of the insulator to be detected determined in step S105.
[0056] The target detection result of the insulator to be detected obtained in step S105 includes the defect position of the defect in the insulator to be detected and the defect type of the insulator to be detected.
[0057] When the first electronic device acquires the image to be detected, it can input the image to be detected into the target detection model.
[0058] Reference Figure 2 , shows a structural schematic diagram of a target detection model provided by an embodiment of the present invention, the target detection model includes a backbone network (Backbone), a neck network (Neck) and a head network (Head) connected in sequence; wherein the backbone network includes a first single-pass downsampler (SPD), a spatial pyramid pooling (ASPP) layer and a second single-pass downsampler connected in sequence; the head network includes an upsampling layer, a similarity-aware activation (SimAM) layer and a convolutional normalization layer connected in sequence.
[0059] The backbone network is used to obtain the first extracted feature map corresponding to the image to be detected in step S102, and input the first extracted feature map into the neck network; the neck network is used to perform feature fusion processing on the first extracted feature map input into the backbone network in step S103 to obtain a first fused feature map, and input the first fused feature map into the head network; the head network is used to determine the first prediction frame of the image to be detected and the first detection result corresponding to the first prediction frame based on the first fused feature map input into the neck network in step S104.
[0060] Understandably, Figure 2 The structure shown in is a necessary structure for implementing the embodiment of the present invention. In actual application scenarios, other components of the target detection model can be flexibly selected based on the detection needs of the image to be detected. In this way, each link of the target detection model can adopt the optimal configuration without compromising the performance of any link.
[0061] Specifically, in step S102, the first electronic device can use the first single-channel downsampler, the spatial pyramid pooling layer and the second single-channel downsampler to perform feature extraction processing on the image to be detected in sequence, and determine the feature map output by the second single-channel downsampler as the first extracted feature map corresponding to the image to be detected.
[0062] In an embodiment of the present invention, the neck network may include a pyramid pooling layer (SPP) and a convolutional batch normalization layer (CBL) connected in sequence; in step S103, when the first electronic device uses the neck network to perform feature fusion processing on the first extracted feature map, first, the pyramid pooling layer is used to perform pooling operations on the first extracted feature map at different scales to generate multiple feature maps of different scales, and these feature maps are concatenated to form a final first feature representation, and the first feature representation is transmitted to the convolutional batch normalization layer; then, the convolution kernel of the convolutional batch normalization layer is multiplied and summed with the sliding window of the first feature representation input by the pyramid pooling layer, and a bias term is added to generate a first fused feature map.
[0063] In step S104, the first electronic device can use the upsampling layer, the similarity perception activation layer and the convolution normalization layer to sequentially upsample, globally weight and convolutionally normalize the first fused feature map, obtain the prediction box output by the convolution normalization layer, and determine the prediction box output by the convolution normalization layer as the first prediction box; then match the first image in the first prediction box with the target image in the defect database, and determine the defect type corresponding to the target image matching the first image in the defect database as the first detection result corresponding to the first prediction box. The first prediction box is the location of the defect of the insulator to be detected in the image to be detected predicted by the target detection model based on the image to be detected.
[0064] In step S105, the first electronic device first obtains the first prediction box and the first detection result output by the head network of the target detection model; then determines the position indicated by the first prediction box as the defect position of the defect in the insulator to be detected, and determines the first detection result as the defect type of the insulator to be detected; finally, determines the defect position and defect type of the insulator to be detected as the target detection result of the insulator to be detected.
[0065] It should be noted that, when there is no defect in the insulator to be inspected, the first electronic device cannot determine the first prediction box and the first detection result through step S104; when the first electronic device cannot determine the first prediction box and the first detection result through step S104, it can directly execute step S105 to determine that the target detection result of the insulator to be inspected is that the quality of the insulator to be inspected is intact and without defects.
[0066] The insulator detection method provided by the embodiment of the present invention obtains an image to be detected including the insulator to be detected, and inputs the image to be detected into a target detection model. In the process of obtaining a first extracted feature map corresponding to the image to be detected by using the first single-channel downsampler, the spatial pyramid pooling layer and the second single-channel downsampler in the backbone network, the image to be detected can be downsampled without losing information of the image to be detected based on the first single-channel downsampler and the second single-channel downsampler, thereby improving the detection capability of small target insulators in the image to be detected. Multi-scale context information can be obtained from the image to be detected based on the spatial pyramid pooling layer, thereby improving the accuracy of the first extracted feature value obtained by the backbone network. In the process of obtaining a first extracted feature map corresponding to the image to be detected by using the upsampling layer, the spatial pyramid pooling layer and the second single-channel downsampler in the head network, the image to be detected can be downsampled without losing information of the image to be detected. In the process of determining the first prediction box and the first detection result based on the first fused feature map obtained by the neck network by the similarity-aware activation layer and the convolution normalization layer, the global context information of the first fused feature map is calculated by global attention weighting on the first fused feature map based on the similarity-aware activation layer, thereby improving the head network's ability to distinguish insulators in the image to be detected, and improving the accuracy of the first prediction box and the first detection result determined by the head network, thereby improving the accuracy and reliability of the target detection result of the insulator to be detected determined by the embodiment of the present invention according to the first prediction box and the first detection result. Moreover, the embodiment of the present invention can be automatically performed in a scenario without human participation, thereby reducing the labor cost in the insulator detection process and improving the efficiency of insulator detection.
[0067] Optionally, the step S102 uses the first single-channel downsampler, the spatial pyramid pooling layer, and the second single-channel downsampler in the backbone network to obtain a first extracted feature map corresponding to the image to be detected, including:
[0068] Step S1021: Use the first single-channel downsampler to perform a downsampling process on the image to be detected to obtain a first weighted feature map corresponding to the image to be detected.
[0069] Step S1022: Using the spatial pyramid pooling layer, perform convolution stacking processing on the first weighted feature map to obtain a first spatial feature map.
[0070] Step S1023: Use the second single-channel downsampler to perform secondary downsampling processing on the first spatial feature map to obtain a first extracted feature map corresponding to the image to be detected.
[0071] Specifically, in step S1021, the first electronic device uses a first single-channel downsampler to map the image to be detected X to a first weighted feature map X'. In the process of mapping the image to be detected X to the first weighted feature map X': first, the image to be detected X is sliced into at least scales by using the first single-channel downsampler. 2sub-feature maps f, where the parameter scale represents the slicing factor; then, each sub-feature map f is downsampled according to the given scaling factor scale; then, the first weight value corresponding to each sub-feature map f is calculated; finally, based on the first weight value corresponding to each sub-feature map f, all sub-feature maps f on the channel are connected to form a first weighted feature map X'.
[0072] As an example, the size of the image to be detected X is S×S×C1, and the sequence of sub-feature graphs f generated is:
[0073] f 0,0 =X[0:S:scale,0:S:scale],f 1,0 =X[1:S:scale,0:S:scale],...,f scale-1,0 =X[scale-1:S:scale,0:S:scale];
[0074] f 0,1 =X[0:S:scale,1:S:scale],f 1,1 ,...,f scale-1,1 =X[scale-1:S:scale,1:S:scale]; ......
[0076] f 0,scale-1 =X[0:S:scale,scale-1:S:scale],f 1, scale-1,...,f scale-1,scale-1 =X[scale-1:S:scale,scale-1:S:scale]; (1)
[0077] It can be understood that each sub-feature map f is a part of the image to be detected X, where the indexes i and j in X(i,j) are divisible by scale.
[0078] The first electronic device calculates the first weight value corresponding to each sub-feature graph f using the first single-channel downsampler, specifically:
[0079] First, the first electronic device uses the first single-channel downsampler to calculate the mean μ of each sub-feature graph f by the following formula 2, and calculates the variance σ of each sub-feature graph f by the following formula 3: 2 ;
[0080]
[0081] Where M represents the number of values on each channel; x i Represents the eigenvalues on the channel.
[0082] Finally, the first electronic device uses the first single-channel downsampler according to the variance σ of the sub-feature map f 2 The first weight value of the sub-feature graph f is calculated by the following formula:
[0083]
[0084] Among them, a represents the first weight value of the sub-feature map f; C represents the number of channels corresponding to the image to be detected X; β represents the learnable parameters in the process of model training for the target detection model; ε is a constant used to prevent numerical instability.
[0085] The first electronic device uses the first single-channel downsampler to connect all sub-feature graphs f on the channel based on the first weight value corresponding to each sub-feature graph f to form a first weighted feature graph X', specifically:
[0086] The first electronic device determines the first weighted feature map X' using the first single-channel downsampler based on the weight matrix A formed by the first weight value corresponding to each sub-feature map f and the image X to be detected formed by all sub-feature maps f, and the calculation formula is as follows:
[0087] X'=A⊙X (5)
[0088] In an embodiment of the present invention, when the first electronic device obtains the first weighted feature map corresponding to the image to be detected by using the first single-channel downsampler, the first weighted feature map is input into the spatial pyramid pooling layer, and step S1022 is executed to perform convolution stacking processing on the first weighted feature map by using the spatial pyramid pooling layer to obtain the first spatial feature map.
[0089] Among them, refer to Figure 3 , shows a schematic diagram of the structure of a spatial pyramid pooling layer provided by an embodiment of the present invention, wherein the spatial pyramid pooling layer includes at least two parallel branches with different expansion rates, each parallel branch performs a pooling operation on the input first weighted feature map to reduce the dimension of the first weighted feature map, applies a convolution kernel to extract features of different scales of the first weighted feature map, and the outputs of the parallel branches are spliced together and further processed to form a final first spatial feature map. The spatial pyramid pooling layer of an embodiment of the present invention uses convolution stacking to expand the expansion, and then stacks and cascades from the expansion coefficient from small to large, expands the receptive field area to extract the feature map connection containing different feature information, so that the obtained first spatial feature map has a larger range of receptive fields, which is conducive to improving the accuracy of the first extracted feature value obtained by the backbone network.
[0090] As an example, the first electronic device uses the spatial pyramid pooling layer to perform convolution stacking processing on the first weighted feature map to obtain a mathematical description of the first spatial feature map:
[0091] Y=H {3,6} (X')+H {3,12} (X')+H {3,18} (X')+H {3,18} (H {3,12} (H {3,6} (X'))) (6)
[0092] Wherein, Y represents the first spatial feature map.
[0093] In an embodiment of the present invention, when the first electronic device uses a spatial pyramid pooling layer to perform convolution stacking processing on the first weighted feature map to obtain a first spatial feature map, the first spatial feature map is input into a second single-channel downsampler, and step S1023 is executed to perform secondary downsampling processing on the first spatial feature map using the second single-channel downsampler to obtain a first extracted feature map corresponding to the image to be detected.
[0094] Among them, the specific processing process of the second single-channel downsampler performing secondary downsampling processing on the first spatial feature map to obtain the first extracted feature map corresponding to the image to be detected is the same as the processing process of the first single-channel downsampler performing one downsampling processing on the image to be detected to obtain the first weighted feature map corresponding to the image to be detected in step S1021. To avoid repetition, it will not be repeated here.
[0095] The insulator detection method provided by the embodiment of the present invention uses the first single-channel downsampler, the spatial pyramid pooling layer and the second single-channel downsampler in the backbone network to obtain the first extracted feature map corresponding to the image to be detected. Based on the first single-channel downsampler and the second single-channel downsampler, the image to be detected can be downsampled without losing the information of the image to be detected, thereby improving the detection capability of small target insulators in the image to be detected. Based on the spatial pyramid pooling layer, multi-scale context information can be obtained from the image to be detected, thereby improving the accuracy of the first extracted feature value obtained by the backbone network, thereby improving the accuracy and reliability of the target detection result of the insulator to be detected determined by the embodiment of the present invention.
[0096] Optionally, in step S104, using an upsampling layer, a similarity-aware activation layer, and a convolutional normalization layer in the head network to determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the first fused feature map includes:
[0097] Step S1041: Use the upsampling layer to upsample the first fused feature map to obtain a first upsampled feature map.
[0098] Step S1042: Use the similarity-aware activation layer to calculate the three-dimensional weight values of the first up-sampled feature map to obtain the energy value of each channel in the first up-sampled feature map.
[0099] Step S1043: using the similarity-aware activation layer to obtain a weighted enhanced feature map corresponding to the first up-sampled feature map according to the energy value.
[0100] Step S1044: Determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the weighted enhanced feature map using the convolution normalization layer.
[0101] In an embodiment of the present invention, the process in which the first electronic device uses the upsampling layer, the similarity-aware activation layer and the convolution normalization layer in the head network to determine the first prediction box of the image to be detected and the first detection result corresponding to the first prediction box based on the first fusion feature map can be implemented through steps S1041 to S1044.
[0102] Specifically, in step S1041, the first electronic device may utilize the upsampling layer to upsample the first fused feature map by an upsampling processing technique well known to those skilled in the art to obtain a first upsampled feature map.
[0103] When the first electronic device uses the upsampling layer to upsample the first fused feature map to obtain the first upsampled feature map, the first upsampled feature map is input into the similarity perception activation layer, and then executes steps S1042 to S1043 to obtain the weight enhanced feature map corresponding to the first upsampled feature map using the similarity perception activation layer.
[0104] Among them, refer to Figure 4 , which shows a schematic diagram of the structure of a similarity perception activation layer provided by an embodiment of the present invention, such as Figure 4 As shown, the first electronic device uses the similarity-aware activation layer to obtain the weight-enhanced feature map corresponding to the first upsampled feature map, specifically: using the similarity-aware activation layer to calculate the three-dimensional (3D) weight values corresponding to each channel in the first upsampled feature map based on the SimAM attention mechanism, and calculating the global context information of the first upsampled feature map through global attention weighting according to the three-dimensional weight values of each channel, to obtain the weight-enhanced feature map corresponding to the first upsampled feature map.
[0105] In the process of obtaining the weight-enhanced feature map corresponding to the first up-sampled feature map, the embodiment of the present invention uses a parameter-free similarity-aware activation layer to solve the problem of too many parameters, and improves the ability to distinguish the shape, size and texture of defects in the insulator to be detected by implementing global attention weighted calculation on the first up-sampled feature map. Generally speaking, channels far from the decision boundary are easy samples, and channels close to the decision boundary are difficult samples. The energy of each channel in the first up-sampled feature map is calculated by the 3D weight value. The smaller the energy, the more important the channel is, and the greater the weight of the channel is. The energy function definition of each channel conforms to the following formula:
[0106]
[0107] Where t represents the target neuron of the similarity-aware activation layer input to any channel in the first upsampled feature map; e t Represents the energy value of the channel input to the target neuron; x i represents other neurons other than the target neuron in the similarity-aware activation layer; w t represents the linear transformation weight; b t Represents the linear transformation bias; y represents the output value of the target neuron.
[0108] By linearly transforming the weights w in Formula 7 t and the linear transformation bias b t Taking partial derivatives, we get the expression of minimum energy as follows:
[0109]
[0110] in, represents the minimum energy value of the channel input to the target neuron; λ is a constant 1E-4; T represents the characteristic value of the channel input to the target neuron.
[0111] In e t for In the case of , the target neuron t has the greatest inhibitory effect on other surrounding neurons and is the most concerned. Therefore, we can calculate To obtain the importance of each neuron, the similarity-aware activation layer obtains the weighted enhanced feature map corresponding to the first up-sampled feature map according to the energy value. The mathematical description is as follows:
[0112]
[0113] in, represents the weight enhanced feature map; X'' represents the first upsampled feature map; ⊙ represents the dot product operation.
[0114] After the first electronic device uses the similarity perception activation layer to obtain the weight-enhanced feature map corresponding to the first up-sampled feature map through step S1043, the weight-enhanced feature map is input into the convolutional normalization layer, so that the first electronic device executes step S1044 to determine the first prediction frame of the image to be detected and the first detection result corresponding to the first prediction frame based on the weight-enhanced feature map using the convolutional normalization layer.
[0115] In the process of determining the first prediction box and the first detection result based on the first fusion feature map obtained by the neck network using the upsampling layer, the similarity-aware activation layer and the convolution normalization layer in the head network, the embodiment of the present invention calculates the global context information of the first upsampling feature map input by the upsampling layer through global attention weighting based on the similarity-aware activation layer, thereby improving the ability to distinguish the shape, size and texture of defects in the insulator to be detected, improving the accuracy of the first prediction box and the first detection result determined by the head network, and thereby improving the accuracy and reliability of the target detection result of the insulator to be detected determined by the embodiment of the present invention according to the first prediction box and the first detection result.
[0116] Reference Figure 5 , shows a flow chart of the steps of a model training method provided by an embodiment of the present invention, and the method may specifically include the following steps:
[0117] Step S201, obtaining an image data set corresponding to an insulator; the image data set includes training image samples, target frames corresponding to the training image samples, and label data corresponding to the target frames.
[0118] Step S202: In each round of training, the to-be-trained detection model is trained using the training image samples to obtain a second prediction box corresponding to the training image samples and a second detection result corresponding to the second prediction box.
[0119] Step S203: determine the loss value corresponding to the detection model to be trained according to the second prediction box and the second detection result, the target box and the label data.
[0120] Step S204: adjust the model parameters of the detection model to be trained according to the loss value, and perform the next round of training until the loss value meets the preset conditions to obtain the target detection model.
[0121] The model training method provided in the embodiment of the present invention can be applied to a second electronic device. The second electronic device includes but is not limited to a smart terminal, a computer, a personal digital assistant, a tablet computer, a vehicle-mounted device, etc. The second electronic device can be used alone to execute the model training method provided in the embodiment of the present invention.
[0122] The second electronic device and the aforementioned first electronic device may be the same device or different devices.
[0123] In step S201:
[0124] First, the second electronic device can use the image acquisition device set in the transmission line to obtain original images of different insulators in the transmission line and original images of the same insulator at different angles and different light intensities to obtain training image samples; the number of training image samples obtained by the second electronic device is greater than or equal to 500.
[0125] Then, the second electronic device labels the preset box in the training image sample according to the preset box and the label data corresponding to the preset box input by the user for the training image sample, obtains the target box corresponding to the training image sample, and obtains the label data corresponding to the target box by adding the label corresponding to the target box in the training image sample through the labelimg software; wherein the format of the label data is txt format, and the label data includes insulator, polymerdirty, pollution-flashover, Glassloss, and broken disc.
[0126] Finally, the second electronic device constructs an image data set corresponding to the insulator based on the training image samples, the target frames corresponding to the training image samples, and the label data corresponding to the target frames.
[0127] The detection model to be trained is obtained by improving the YOLOv7 model on the basis of the YOLOv7 model; the model to be trained includes a connected backbone network to be trained, a neck network to be trained and a head network to be trained; wherein the backbone network to be trained includes a first single-channel down-sampler to be trained, a spatial pyramid pooling layer to be trained and a second single-channel down-sampler to be trained which are connected in sequence; the head network to be trained includes an up-sampling layer to be trained, a similarity-aware activation layer to be trained and a convolutional normalization layer to be trained which are connected in sequence.
[0128] In step S202, in each round of training, the second electronic device inputs any training image sample in the image data set into the detection model to be trained for training, obtains the prediction box output by the detection model to be trained and the detection result corresponding to the prediction box, and determines the prediction box output by the detection model to be trained and the detection result corresponding to the prediction box as the second prediction box corresponding to the training image sample and the second detection result corresponding to the second prediction box.
[0129] In step S203, first, the second electronic device determines the first loss value corresponding to the detection model to be trained based on the second prediction box and the target box; then, the second electronic device determines the second loss value corresponding to the detection model to be trained based on the second detection result and the label data; finally, the second electronic device determines the sum or weighted sum of the first loss value and the second loss value as the loss value corresponding to the detection model to be trained.
[0130] In step S204, the second electronic device performs reverse gradient propagation according to the loss value determined in step S203, adjusts the model parameters of the detection model to be trained, and performs the next round of training until the loss value meets the preset conditions to obtain the target detection model as described in any of the above items.
[0131] As an example, the training process of the detection model to be trained includes 300 epochs and the batchsize is set to 32.
[0132] The model training method provided in an embodiment of the present invention improves the accuracy of the first prediction box and the first detection result determined by the first electronic device using the trained target detection model by training the detection model to be trained, thereby improving the accuracy and reliability of the target detection result of the insulator to be detected determined by the first electronic device based on the first prediction box and the first detection result.
[0133] As an optional implementation, the second electronic device obtains the image data set corresponding to the insulator through step S201, including a training set, a verification set and a test set, wherein the training set, the verification set and the test set each include a certain number of training image samples, target frames corresponding to the training image samples and label data corresponding to the target frames, and the training image samples in the training set, the verification set and the test set are different.
[0134] Specifically, the second electronic device uses the training image samples in the training set, the target boxes corresponding to the training image samples, and the label data corresponding to the target boxes to train the detection model to be trained through operations corresponding to steps S202 to S204 until the loss value of the model to be detected meets the preset conditions.
[0135] After the loss value of the model to be detected meets the preset conditions, the second electronic device can use the training image samples in the verification set, the target boxes corresponding to the training image samples, and the label data corresponding to the target boxes to evaluate the model to be trained whose loss value meets the preset conditions, so as to monitor the performance of the model to be trained, avoid overfitting, and determine the detection model to be trained with the best model parameters (such as learning rate, number of iterations, etc.) as the target detection model.
[0136] After evaluating the model to be trained whose loss value meets the preset conditions by using the training image samples in the validation set, the target frames corresponding to the training image samples, and the label data corresponding to the target frames to determine the target detection model, the second electronic device can use the training image samples in the test set, the target frames corresponding to the training image samples, and the label data corresponding to the target frames to evaluate the performance of the finally selected target detection model.
[0137] The embodiment of the present invention divides the image data set into a training set, a validation set, and a test set, which can improve the generalization ability of the final target detection model and avoid overfitting.
[0138] As an optional implementation, refer to Figure 6 , shows a logic block diagram of a model training provided by an embodiment of the present invention; specifically:
[0139] The second electronic device constructs an image data set corresponding to the insulator: first, the training image samples are determined based on self-collected data and / or public databases such as GitHub and Electricity, wherein the self-collected data is the original images of different insulators in the transmission line and the original images of the same insulator at different angles and different light intensities obtained by the second electronic device using an image acquisition device set in the transmission line; then, according to the preset frame for the training image sample and the label data corresponding to the preset frame input by the user, the preset frame is annotated in the training image sample to obtain the target frame corresponding to the training image sample; then, the label corresponding to the target frame is added to the training image sample by the labelimg software to obtain the label data corresponding to the target frame, and the format of the annotated training image sample and the label data is standardized; then, based on the training image samples, the target frames corresponding to the training image samples and the label data corresponding to the target frames, the image data set corresponding to the insulator is constructed; finally, the image data set corresponding to the insulator is divided into a training set, a validation set and a test set, wherein the training set, the validation set and the test set all include a certain number of training image samples, target frames corresponding to the training image samples and the label data corresponding to the target frames, and the training image samples in the training set, the validation set and the test set are different.
[0140] The second electronic device constructs a detection model to be trained: first, a first single-channel down-sampler to be trained, a spatial pyramid pooling layer to be trained, and a second single-channel down-sampler to be trained are set in sequence in the backbone network of the YOLOv7 model; then, an upsampling layer to be trained, a similarity-aware activation layer to be trained, and a convolutional normalization layer to be trained are set in sequence in the head network of the YOLOv7 model; finally, the loss function in the YOLOv7 model is replaced with a weighted intersection over Union loss function (Weighted Intersection over Union, WIOU), specifically, the complete intersection over Union loss function (Complete Intersection over Union, CIoU) in the YOLOv7 model is replaced with a weighted intersection over Union loss function, and the detection model to be trained constructed by an embodiment of the present invention is obtained.
[0141] Perform model training on the detection model to be trained: use the training set to train the detection model to be trained through the operations corresponding to the above steps S202 to S204 to obtain the model to be trained whose loss value meets the preset conditions; use the validation set to evaluate the model to be trained whose loss value meets the preset conditions and determine the detection model to be trained with the best model parameters as the target detection model; use the test set to evaluate the performance of the target detection model finally selected.
[0142] Optionally, the detection model to be trained includes a backbone network to be trained, a neck network to be trained, and a head network to be trained, which are connected in sequence; in each round of training described in step S202, the detection model to be trained is trained using the training image sample to obtain a second prediction frame corresponding to the training image sample and a second detection result corresponding to the second prediction frame, including:
[0143] Step S2021, using the first single-channel down-sampler to be trained, the spatial pyramid pooling layer to be trained and the second single-channel down-sampler to be trained in the backbone network to be trained, to obtain a second extracted feature map corresponding to the training image sample.
[0144] Step S2022: using the neck network to be trained, perform feature fusion processing based on the second extracted feature map to obtain a second fused feature map.
[0145] Step S2023: Determine the second prediction box of the training image sample and the second detection result corresponding to the second prediction box based on the second fused feature map by utilizing the upsampling layer to be trained, the similarity-aware activation layer to be trained and the convolution normalization layer to be trained in the head network to be trained.
[0146] Specifically, the implementation process of step S2021 to step S2023 can refer to the specific description of the aforementioned step S102 to step S104, and will not be repeated here to avoid repetition.
[0147] Optionally, the convolutional normalization layer to be trained includes a weighted intersection-over-union loss function layer; and the step S203 of determining the loss value corresponding to the detection model to be trained according to the second prediction box and the second detection result, the target box and the label data includes:
[0148] Step S2031: Use the weighted intersection-over-union loss function layer to obtain the first coordinate information of the second prediction box and the second coordinate information of the target box.
[0149] Step S2032: using the weighted intersection-over-union loss function layer, calculate the target distance between the second prediction box and the target box according to the first coordinate information and the second coordinate information.
[0150] Step S2033: using the weighted intersection-over-union loss function layer, determine a target weight value according to the target distance.
[0151] Step S2034: using the weighted intersection-over-union loss function layer, determine a first loss value corresponding to the detection model to be trained according to the target weight value, the first coordinate information, and the second coordinate information.
[0152] Step S2035: Determine a second loss value corresponding to the detection model to be trained according to the second detection result and the label data.
[0153] Step S2036: Determine the loss value corresponding to the detection model to be trained according to the first loss value and the second loss value.
[0154] In an embodiment of the present invention, the convolutional normalization layer to be trained in the head network to be trained includes a weighted intersection-over-union loss function layer. The weighted intersection-over-union loss function layer can more finely adjust the position and size of the prediction box through a dynamic non-monotonic focusing mechanism, thereby improving the detection accuracy of small targets, so that the detection model to be trained can focus on anchor frames of ordinary quality, thereby improving the accuracy and reliability of the first prediction box determined by the trained target detection model.
[0155] Specifically, in step S2031, the second electronic device may use a weighted intersection-over-union loss function layer to obtain first coordinate information of the second prediction frame and second coordinate information of the target frame. The first coordinate information includes the first center point coordinate (x p ,y p ) and the first vertex coordinates of the second prediction box; the second coordinate information includes the second center point coordinates of the target box (xt ,y t ) and the coordinates of the second vertex of the target box.
[0156] Reference Figure 7 , which shows a schematic diagram of the relationship between a second prediction frame and a target frame provided by an embodiment of the present invention, such as Figure 7 As shown, the coordinates of the first center point of the second prediction box are (x p ,y p ), the coordinates of the second center point of the target frame are (x t ,y t ). The first width of the second prediction box is w p , the first height is h p ; The second width of the target box is w t , the second height is h t The third width of the overlapping area between the second prediction box and the target box is w i , the third height is h i .
[0157] In step S2032, the second electronic device calculates the Euclidean distance between the center point of the second prediction box and the center point of the target box using the weighted intersection-over-union loss function layer according to the first center point coordinates and the second center point coordinates by the following formula, and determines the Euclidean distance as the target distance:
[0158]
[0159] Where d represents the target distance.
[0160] In step S2033, first, the second electronic device uses the weighted intersection-over-intersection loss function layer to calculate the diagonal length of the feature graph input to the weighted intersection-over-intersection loss function layer based on the Pythagorean theorem; then, the second electronic device uses the weighted intersection-over-intersection loss function layer to calculate the target weight according to the diagonal length and the target distance through the following formula:
[0161]
[0162] Among them, ω represents the target weight; α is a hyperparameter used to adjust the decrease speed of the target weight; D represents the length of the diagonal.
[0163] It can be understood that when the center point of the second prediction box coincides with the center point of the target box (ie, d≈0), the target weight ω≈1, and the weight of the overlapping area between the second prediction box and the target box is the largest; when the center point of the second prediction box and the center point of the target box are far apart (ie, d is close to D), the target weight ω≈1 will become very small, and the weight of the overlapping area between the second prediction box and the target box will also be small.
[0164] As an example, the size of the feature map input to the weighted intersection-over-union loss function layer is 640 pixels × 640 pixels, the first center point coordinates of the second prediction box are (350, 350), and the second center point coordinates of the target box are (320, 320); the target distance calculated by the second electronic device using the weighted intersection-over-union loss function layer through formula 10 is:
[0165]
[0166] The second electronic device uses the weighted intersection-over-union loss function layer to calculate the diagonal length of the feature map input to the weighted intersection-over-union loss function layer based on the Pythagorean theorem:
[0167]
[0168] When α is 1, the second electronic device uses the weighted intersection-over-union loss function layer to calculate the target weight according to the diagonal length and the target distance through formula 11:
[0169]
[0170] When ω≈0.95, it indicates that the offset between the center point of the second prediction box and the center point of the target box is small, and the weight of the overlapping area between the second prediction box and the target box is large (close to 1), indicating that the accuracy of the second prediction box output by the detection model to be trained is high.
[0171] When the second electronic device determines the target weight through step S2033, it can continue to use the weighted intersection-over-union loss function layer through step S2034 to determine the first loss value corresponding to the detection model to be trained according to the target weight value, the first coordinate information, and the second coordinate information.
[0172] Specifically, in step S2034, first, the second electronic device uses a weighted intersection-over-union loss function layer to calculate a first area of the intersection between the second prediction box and the target box, and a second area of the union between the second prediction box and the target box according to the first coordinate information and the second coordinate information; then, the second electronic device uses a weighted intersection-over-union loss function layer to calculate a first ratio of the first area to the second area; finally, the second electronic device uses a weighted intersection-over-union loss function layer to determine the product of the first ratio and the target weight value as a first loss value corresponding to the detection model to be trained; wherein the first loss value is used to reflect the degree of difference between the second prediction box and the target box output by the detection model to be trained.
[0173] In step S2025, the second electronic device can determine the second loss value corresponding to the detection model to be trained based on the second detection result and the label data; wherein the second loss value is used to reflect the degree of difference between the second detection result output by the detection model to be trained and the label data.
[0174] When determining the first loss value and the second loss value, the second electronic device may determine the sum or weighted sum of the first loss value and the second loss value as the loss value corresponding to the detection model to be trained.
[0175] In an embodiment of the present invention, when the target detection model is obtained through step S204, the mean average precision (mAP), precision (Precision), recall (Recall) and frames per second (FPS, Processing time per frame) can be selected to comprehensively evaluate the performance of the target detection model.
[0176] Among them, mAP is used to reflect the detection accuracy of the target detection model for the image to be detected, and the higher the mAP, the higher the detection accuracy of the target detection model; Precision is used to reflect the false detection rate of the target detection model, and the higher the Precision, the lower the false detection rate of the target detection model; Recall is used to reflect the missed detection rate of the target detection model, and the higher the Recall, the lower the missed detection rate of the target detection model; FPS is used to reflect the processing speed of the target detection model for the image to be detected.
[0177] Referring to Table 1, a performance evaluation table of a target detection model provided by an embodiment of the present invention is shown. As shown in Table 1, compared with the YOLOv7 model in the related art, the mAP of the target detection model provided by the embodiment of the present invention is improved by 5.12%, the Precision is improved by 7.52%, the Recall is improved by 7.31%, and the FPS is improved by 5.41%; it can be seen that by using the target detection model provided by the embodiment of the present invention to detect the image to be detected including the insulator to be detected, the false detection rate and the missed detection rate are reduced, the detection accuracy is improved, and the processing speed of the target detection model for the image to be detected is also improved, thereby improving the efficiency of detecting the insulator to be detected by using the target detection model.
[0178] Table 1
[0179] Detection Model mAP@0.5 Precision Recall FPS YOLOv7 56.70% 62.50% 60.20% 20.3% Object Detection Model 59.60% 67.20% 64.60% 21.4% Improvement in target detection model performance 5.12% 7.52% 7.31% 5.41%
[0180] It should be noted that the configuration environment parameters for evaluating the average accuracy mean, precision, recall rate and frames per second of the target detection model in the embodiment of the present invention are shown in Table 2, and the network initialization parameters are shown in Table 3.
[0181] Table 2
[0182] Configure the environment Version Python 3.8 Pytorch-CUDA 12.1 numpy 1.26.4 Tensorboard 2.29 PyQt5 5.15.10 Matplotlib 3.1.3 conda 22.9.0
[0183] Table 3
[0184] parameter Numeric batch_size 32 image_size 640*640 Epochs 300 number of classes 8
[0185] In summary, the insulator detection method provided by the embodiment of the present invention obtains an image to be detected including the insulator to be detected, and inputs the image to be detected into the target detection model. In the process of obtaining the first extracted feature map corresponding to the image to be detected by using the first single-channel downsampler, the spatial pyramid pooling layer and the second single-channel downsampler in the backbone network, the image to be detected can be downsampled without losing the information of the image to be detected based on the first single-channel downsampler and the second single-channel downsampler, thereby improving the detection capability of small target insulators in the image to be detected. Based on the spatial pyramid pooling layer, multi-scale context information can be obtained from the image to be detected, thereby improving the accuracy of the first extracted feature value obtained by the backbone network. In the process of using the upsampling layer, the spatial pyramid pooling layer and the second single-channel downsampler in the head network, the image to be detected can be downsampled without losing the information of the image to be detected. In the process of determining the first prediction box and the first detection result based on the first fused feature map obtained by the neck network by the similarity-aware activation layer and the convolution normalization layer, the global context information of the first fused feature map is calculated by global attention weighting on the first fused feature map based on the similarity-aware activation layer, thereby improving the head network's ability to distinguish insulators in the image to be detected, and improving the accuracy of the first prediction box and the first detection result determined by the head network, thereby improving the accuracy and reliability of the target detection result of the insulator to be detected determined by the embodiment of the present invention according to the first prediction box and the first detection result. Moreover, the embodiment of the present invention can be automatically performed in a scenario without human participation, thereby reducing the labor cost in the insulator detection process and improving the efficiency of insulator detection.
[0186] Device Embodiment
[0187] Reference Figure 8 , shows a logic block diagram of an insulator detection device provided by an embodiment of the present invention, and the device may include:
[0188] The first acquisition module 801 is used to acquire an image to be detected and input the image to be detected into a target detection model; the image to be detected includes an insulator to be detected; the target detection model includes a backbone network, a neck network and a head network connected in sequence;
[0189] The second acquisition module 802 is used to obtain a first extracted feature map corresponding to the image to be detected by using the first single-channel downsampler, the spatial pyramid pooling layer and the second single-channel downsampler in the backbone network;
[0190] An extraction module 803 is used to perform feature fusion processing on the first extracted feature map using the neck network to obtain a first fused feature map;
[0191] A first determination module 804 is used to determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the first fused feature map by using the upsampling layer, the similarity-aware activation layer and the convolution normalization layer in the head network;
[0192] The second determination module 805 is used to obtain the first prediction frame and the first detection result output by the target detection model, and determine the target detection result of the insulator to be detected according to the first prediction frame and the first detection result.
[0193] Optionally, the second acquisition module includes:
[0194] A first processing submodule, configured to perform a downsampling process on the image to be detected by using the first single-channel downsampler to obtain a first weighted feature map corresponding to the image to be detected;
[0195] A second processing submodule is used to perform convolution stacking processing on the first weighted feature map by using the spatial pyramid pooling layer to obtain a first spatial feature map;
[0196] The third processing submodule is used to use the second single-channel downsampler to perform secondary downsampling processing on the first spatial feature map to obtain a first extracted feature map corresponding to the image to be detected.
[0197] Optionally, the first determining module includes:
[0198] a fourth processing submodule, configured to perform upsampling processing on the first fused feature map by using the upsampling layer to obtain a first upsampling feature map;
[0199] A first calculation submodule, configured to calculate a three-dimensional weight value of the first up-sampled feature map by using the similarity-aware activation layer to obtain an energy value of each channel in the first up-sampled feature map;
[0200] A first acquisition submodule, configured to acquire a weighted enhanced feature map corresponding to the first up-sampled feature map according to the energy value by using the similarity-aware activation layer;
[0201] The first determination submodule is used to determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the weight enhanced feature map by using the convolution normalization layer.
[0202] Reference Fig. 9, shows a logic block diagram of a model training device provided by an embodiment of the present invention, and the device may include:
[0203] The third acquisition module 901 is used to acquire an image data set corresponding to the insulator; the image data set includes a training image sample, a target frame corresponding to the training image sample, and label data corresponding to the target frame;
[0204] A training module 902 is used to train the to-be-trained detection model using the training image samples in each round of training to obtain a second prediction box corresponding to the training image samples and a second detection result corresponding to the second prediction box;
[0205] A third determination module 903 is used to determine the loss value corresponding to the to-be-trained detection model according to the second prediction box and the second detection result, as well as the target box and the label data;
[0206] The training module 902 is also used to adjust the model parameters of the detection model to be trained according to the loss value, and perform the next round of training until the loss value meets the preset conditions to obtain the target detection model as described in any of the above items.
[0207] Optionally, the detection model to be trained includes a backbone network to be trained, a neck network to be trained and a head network to be trained which are connected in sequence; the training module includes:
[0208] A second acquisition submodule is used to acquire a second extracted feature map corresponding to the training image sample by using the first single-channel downsampler to be trained, the spatial pyramid pooling layer to be trained and the second single-channel downsampler to be trained in the backbone network to be trained;
[0209] A fifth processing submodule, configured to use the neck network to be trained to perform feature fusion processing based on the second extracted feature map to obtain a second fused feature map;
[0210] The second determination submodule is used to determine the second prediction box of the training image sample and the second detection result corresponding to the second prediction box based on the second fused feature map by using the upsampling layer to be trained, the similarity-aware activation layer to be trained and the convolution normalization layer to be trained in the head network to be trained.
[0211] Optionally, the convolution normalization layer to be trained includes a weighted intersection-over-union loss function layer; and the third determination module includes:
[0212] A third acquisition submodule is used to acquire the first coordinate information of the second prediction box and the second coordinate information of the target box by using the weighted intersection-over-union loss function layer;
[0213] A second calculation submodule, configured to calculate a target distance between the second prediction box and the target box according to the first coordinate information and the second coordinate information by using the weighted intersection-over-union loss function layer;
[0214] A third determination submodule is used to determine a target weight value according to the target distance by using the weighted intersection-over-union loss function layer;
[0215] A fourth determination submodule is used to determine a first loss value corresponding to the detection model to be trained according to the target weight value, the first coordinate information, and the second coordinate information by using the weighted intersection-over-union loss function layer;
[0216] A fifth determination submodule, used to determine a second loss value corresponding to the detection model to be trained according to the second detection result and the label data;
[0217] The sixth determination submodule is used to determine the loss value corresponding to the detection model to be trained according to the first loss value and the second loss value.
[0218] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0219] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0220] An embodiment of the present invention also provides an electronic device, which includes a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions enable the processor to execute the insulator detection method of the aforementioned embodiment, or execute the model training method of the aforementioned embodiment.
[0221] The embodiment of the present invention also provides a non-temporary computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device (server or terminal), the processor is able to execute Figure 1 Insulator inspection method shown, or perform Figure 5 The model training method shown.
[0222] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0223] It should be understood by those skilled in the art that the embodiments of the present invention can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0224] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0225] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0226] The insulator detection method, model training method, device, equipment and storage medium provided by the present invention are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. An insulator detection method, characterized in that: The method comprises: Acquire an image to be detected, and input the image to be detected into a target detection model; the image to be detected includes an insulator to be detected; the target detection model includes a backbone network, a neck network and a head network connected in sequence; Utilizing the first single-channel downsampler, the spatial pyramid pooling layer, and the second single-channel downsampler in the backbone network, obtaining a first extracted feature map corresponding to the image to be detected; Using the neck network, performing feature fusion processing on the first extracted feature map to obtain a first fused feature map; Determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the first fused feature map by using an upsampling layer, a similarity-aware activation layer, and a convolutional normalization layer in the head network; The first prediction frame and the first detection result output by the target detection model are obtained, and the target detection result of the insulator to be detected is determined according to the first prediction frame and the first detection result.
2. The method according to claim 1, characterized in that The method of using the first single-channel downsampler, the spatial pyramid pooling layer, and the second single-channel downsampler in the backbone network to obtain a first extracted feature map corresponding to the image to be detected includes: Using the first single-channel downsampler, downsampling the image to be detected is performed once to obtain a first weighted feature map corresponding to the image to be detected; Using the spatial pyramid pooling layer, performing convolution stacking processing on the first weighted feature map to obtain a first spatial feature map; The first spatial feature map is downsampled twice using the second single-channel downsampler to obtain a first extracted feature map corresponding to the image to be detected.
3. The method according to claim 1, characterized in that The method of using the upsampling layer, the similarity-aware activation layer, and the convolution normalization layer in the head network to determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the first fusion feature map includes: Performing upsampling processing on the first fused feature map by using the upsampling layer to obtain a first upsampling feature map; Calculating three-dimensional weight values of the first up-sampled feature map using the similarity-aware activation layer to obtain energy values of each channel in the first up-sampled feature map; Obtaining a weighted enhanced feature map corresponding to the first up-sampled feature map according to the energy value using the similarity-aware activation layer; The convolution normalization layer is used to determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the weight enhanced feature map.
4. A model training method, characterized in that: The method comprises: Acquire an image data set corresponding to the insulator; the image data set includes training image samples, target frames corresponding to the training image samples, and label data corresponding to the target frames; In each round of training, the training image sample is used to train the detection model to be trained, and a second prediction box corresponding to the training image sample and a second detection result corresponding to the second prediction box are obtained; Determine a loss value corresponding to the to-be-trained detection model according to the second prediction box and the second detection result, as well as the target box and the label data; The model parameters of the detection model to be trained are adjusted according to the loss value, and the next round of training is performed until the loss value meets the preset conditions, thereby obtaining the target detection model as described in any one of claims 1 to 3.
5. The method according to claim 4, characterized in that The detection model to be trained includes a backbone network to be trained, a neck network to be trained and a head network to be trained, which are connected in sequence; In each round of training, the training image sample is used to train the detection model to be trained to obtain a second prediction box corresponding to the training image sample and a second detection result corresponding to the second prediction box, including: Using the first single-channel down-sampler to be trained, the spatial pyramid pooling layer to be trained and the second single-channel down-sampler to be trained in the backbone network to be trained, obtaining a second extracted feature map corresponding to the training image sample; Using the neck network to be trained, performing feature fusion processing based on the second extracted feature map to obtain a second fused feature map; Utilizing the upsampling layer to be trained, the similarity-aware activation layer to be trained, and the convolutional normalization layer to be trained in the head network to be trained, based on the second fused feature map, determine the second prediction box of the training image sample and the second detection result corresponding to the second prediction box.
6. The method according to claim 5, characterized in that The convolution normalization layer to be trained includes a weighted intersection-over-union loss function layer; The determining, according to the second prediction box and the second detection result, the target box and the label data, a loss value corresponding to the detection model to be trained includes: Using the weighted intersection-over-union loss function layer, obtaining first coordinate information of the second prediction box and second coordinate information of the target box; Calculating a target distance between the second prediction box and the target box according to the first coordinate information and the second coordinate information by using the weighted intersection-over-union loss function layer; Determining a target weight value according to the target distance by using the weighted intersection-over-union loss function layer; Determine a first loss value corresponding to the detection model to be trained according to the target weight value, the first coordinate information, and the second coordinate information by using the weighted intersection-over-union loss function layer; Determine a second loss value corresponding to the to-be-trained detection model according to the second detection result and the label data; According to the first loss value and the second loss value, a loss value corresponding to the detection model to be trained is determined.
7. An insulator detection method and device, characterized in that: The device comprises: A first acquisition module is used to acquire an image to be detected and input the image to be detected into a target detection model; the image to be detected includes an insulator to be detected; the target detection model includes a backbone network, a neck network and a head network connected in sequence; A second acquisition module is used to acquire a first extracted feature map corresponding to the image to be detected by using the first single-channel downsampler, the spatial pyramid pooling layer and the second single-channel downsampler in the backbone network; An extraction module, used to perform feature fusion processing on the first extracted feature map using the neck network to obtain a first fused feature map; A first determination module, configured to determine a first prediction box of the image to be detected and a first detection result corresponding to the first prediction box based on the first fused feature map by using an upsampling layer, a similarity-aware activation layer, and a convolutional normalization layer in the head network; The second determination module is used to obtain the first prediction frame and the first detection result output by the target detection model, and determine the target detection result of the insulator to be detected based on the first prediction frame and the first detection result.
8. A model training device, characterized in that: The device comprises: A third acquisition module is used to acquire an image data set corresponding to the insulator; the image data set includes a training image sample, a target frame corresponding to the training image sample, and label data corresponding to the target frame; A training module, used to train the to-be-trained detection model using the training image samples in each round of training, to obtain a second prediction box corresponding to the training image samples and a second detection result corresponding to the second prediction box; A third determination module is used to determine the loss value corresponding to the to-be-trained detection model according to the second prediction box and the second detection result, the target box and the label data; The training module is also used to adjust the model parameters of the detection model to be trained according to the loss value, and perform the next round of training until the loss value meets the preset conditions, thereby obtaining the target detection model as described in any one of claims 1 to 3.
9. An electronic device, characterized in that: The electronic device comprises a processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store executable instructions, and the executable instructions enable the processor to execute the insulator detection method as described in any one of claims 1 to 3, or execute the model training method as described in any one of claims 4 to 6.
10. A readable storage medium, characterized in that: When the instructions in the readable storage medium are executed by a processor of an electronic device, the processor is enabled to execute the insulator detection method as described in any one of claims 1 to 3, or to execute the model training method as described in any one of claims 4 to 6.