Power grid operation site insulating glove wearing detection method and device, computer equipment, readable storage medium and program product

By using the neural network feature extraction module and channel space feature enhancement attention module at the grid operation site, and extracting insulated glove features and combining the multi-expansion aggregation module for feature fusion, the problem of low detection accuracy of insulated gloves in the prior art is solved, and higher detection accuracy and adaptability to complex scenes are achieved.

CN120183006AActive Publication Date: 2025-06-20ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510671877.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The object detection algorithm used in the prior art at the power grid operation site is difficult to accurately locate the specific area of ​​insulated gloves, resulting in low detection accuracy, especially when the background is complex, the human hand is blocked or the angle changes, it is difficult to effectively detect.

Method used

A method of wearing and detecting insulated gloves on the power grid operation site is adopted. Channel feature extraction is performed through the neural network feature extraction module, and the attention module is used to extract hand and glove features, and local and global features are fused through the multi-expansion aggregation module to improve detection accuracy.

Benefits of technology

The global context information of the feature map of the insulated gloves is enhanced, the feature extraction ability of the fuzzy target is improved, the target occlusion and angle changes are effectively dealt with, and the accuracy of insulated gloves is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183006A_ABST
    Figure CN120183006A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric power safety operation, and provides a power grid operation site insulating glove wearing detection method and device, computer equipment, a readable storage medium and a program product. The method comprises: according to a neural network feature extraction module of an insulating glove wearing detection model, extracting channel features of a power grid operation field personnel image to obtain a channel feature map; according to a channel space feature attention enhancement module of the insulating glove wearing detection model, extracting hand features of a person and features of the insulating glove in the channel feature map to obtain an insulating glove feature map; the channel space feature attention enhancement module is constructed according to global pooling operation in combination with a parameter-free attention mechanism; and according to a multi-expansion convergence module of the insulating glove wearing detection model, local and global features of the insulating glove feature map are extracted and then fusion processing is performed to obtain an insulating glove feature enhancement map so as to obtain an insulating glove wearing detection result. The method can improve the detection precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of electric power safety operations, and particularly to a detection method, device, computer equipment, computer-readable storage medium, and computer program product for wearing insulating gloves at the power grid operation site. Background Art

[0002] In the electric power industry, the safety management of the power grid operation site is of crucial importance. The power grid construction site involves the operation or maintenance of high-voltage electrical equipment. Especially when carrying out live operations, power grid construction personnel should wear insulating gloves that meet the requirements of the corresponding voltage level to ensure personal safety.

[0003] Currently, it is possible to detect the wearing of insulating gloves at the power grid operation site based on the object detection algorithm of deep learning. However, due to the complex and variable background of the power grid construction site, it is difficult to accurately locate the specific area where the gloves are worn when using the existing object detection algorithm for glove wearing detection, resulting in a low detection accuracy of insulating glove wearing; when the power grid construction personnel are far away from the detection device, the glove wearing features may appear as relatively small pixel areas in the image, and the existing object detection algorithms have difficulties in capturing these subtle features and are prone to ignoring key information, resulting in a low detection accuracy of insulating glove wearing; in the actual working environment, the hands of power grid construction workers may be partially blocked or the angle may be changed due to the operating posture and obstacles, and the existing object detection algorithms are difficult to effectively handle this kind of occlusion and angle change, resulting in a low detection accuracy of insulating glove wearing. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a detection method, device, computer equipment, computer-readable storage medium, and computer program product for wearing insulating gloves at the power grid operation site.

[0005] In a first aspect, this application provides a detection method for wearing insulating gloves at the power grid operation site, including:

[0006] Performing channel feature extraction on the image of personnel at the power grid operation site according to the neural network feature extraction module of the insulating glove wearing detection model to obtain a channel feature map;

[0007] Extracting the hand features of the personnel and the features of the insulating gloves in the channel feature map according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model to obtain an insulating glove feature map; the channel spatial feature enhancement attention module is constructed by combining global pooling operation with a parameter-free attention mechanism;

[0008] According to the multi-dilation convergence module of the insulating glove wearing detection model, the local and global features of the insulating glove feature map are extracted and then fused to obtain an enhanced insulating glove feature map;

[0009] According to the enhanced insulating glove feature map, it is determined whether the grid construction personnel in the grid operation site personnel image wear insulating gloves, whether the insulating gloves are worn correctly, and the target position in the grid operation site personnel image, so as to obtain the insulating glove wearing detection result of the grid operation site personnel image.

[0010] In one embodiment, the channel spatial feature enhanced attention module of the insulating glove wearing detection model extracts the hand features of the personnel and the features of the insulating gloves in the channel feature map to obtain the insulating glove feature map, including:

[0011] According to the channel spatial feature enhanced attention module of the insulating glove wearing detection model, a convolution operation is performed on the channel feature map to extract the hand features of the personnel and the features of the insulating gloves in the channel feature map to obtain a first intermediate feature map;

[0012] According to the channel spatial feature enhanced attention module of the insulating glove wearing detection model, global average pooling operation and global maximum pooling operation are respectively performed on the channel feature map to extract the hand features of the personnel and the features of the insulating gloves in the channel feature map to obtain a second intermediate feature map and a third intermediate feature map, so as to obtain channel feature weights;

[0013] According to the channel feature map and the channel feature weights, a fifth intermediate feature map is obtained;

[0014] According to the first intermediate feature map and the channel feature weights, a sixth intermediate feature map is obtained;

[0015] The channel feature map is input into a parameter-free attention mechanism to obtain a seventh intermediate feature map;

[0016] The fifth intermediate feature map, the sixth intermediate feature map and the seventh intermediate feature map are added element by element to obtain an eighth intermediate feature map, and a convolution operation is performed on the eighth intermediate feature map to obtain the insulating glove feature map.

[0017] In one embodiment, obtaining the channel feature weights includes:

[0018] The second intermediate feature map and the third intermediate feature map are added element by element to obtain a fourth intermediate feature map;

[0019] According to the fourth intermediate feature map and the activation function, the channel feature weights are obtained.

[0020] In one embodiment, the multi-dilation pooling module according to the insulating glove wearing detection model performs local and global feature extraction on the insulating glove feature map and then performs a fusion process to obtain an enhanced insulating glove feature map, including:

[0021] The multi-dilation pooling module according to the insulating glove wearing detection model performs a max pooling operation on the insulating glove feature map to obtain a first intermediate enhanced feature map;

[0022] The multi-dilation pooling module according to the insulating glove wearing detection model performs consecutive multiple dilated convolution operations on the insulating glove feature map to obtain a second intermediate enhanced feature map, a third intermediate enhanced feature map, and a fourth intermediate enhanced feature map;

[0023] The first intermediate enhanced feature map, the second intermediate enhanced feature map, the third intermediate enhanced feature map, and the fourth intermediate enhanced feature map are concatenated according to the channel dimension to obtain a fifth intermediate enhanced feature map;

[0024] A convolution operation is performed on the fifth intermediate enhanced feature map to obtain a sixth intermediate enhanced feature map;

[0025] According to the non-linear activation function, non-linear enhancement is performed on the sixth intermediate enhanced feature map to obtain an enhanced insulating glove feature map.

[0026] In one embodiment, the method further includes:

[0027] Obtain an insulating glove wearing detection training set and an insulating glove wearing detection validation set;

[0028] Input the insulating glove wearing detection training images in the insulating glove wearing detection training set into the model with parameters to be adjusted to obtain a first prediction result;

[0029] According to the first prediction result, the annotation data corresponding to the insulating glove wearing detection training image, and the adaptive intersection over union loss function, obtain a first training loss value; the adaptive intersection over union loss function is obtained according to the intersection over union loss, the center point distance loss, and the length and width loss;

[0030] Input the insulating glove wearing detection validation images in the insulating glove wearing detection validation set into the model with parameters to be adjusted to obtain a second prediction result;

[0031] According to the second prediction result, the annotation data corresponding to the insulating glove wearing detection validation image, and the adaptive intersection over union loss function, obtain a second training loss value;

[0032] Adjust the parameters of the model with parameters to be adjusted according to the first training loss value and the second training loss value until the second training loss value converges, so as to obtain an insulating glove wearing detection model.

[0033] In one embodiment, the obtaining of the insulating glove wearing detection training set and the insulating glove wearing detection verification set includes:

[0034] Collect real images of insulating gloves worn at the power grid operation site and during operation and maintenance.

[0035] Obtain a first simulated image of an insulating glove worn irregularly and a second simulated image of no insulating glove worn.

[0036] According to the real image, the first simulated image and the second simulated image, obtain an insulating glove wearing detection data set.

[0037] Perform wearing category annotation on the insulating glove wearing detection data set to obtain an insulating glove wearing detection annotated data set; the wearing categories include correct wearing of insulating gloves on the hand, no insulating glove worn on the hand, and irregular wearing of insulating gloves.

[0038] According to a set ratio, divide the insulating glove wearing detection annotated data set into an insulating glove wearing detection training set and an insulating glove wearing detection verification set.

[0039] In a second aspect, the present application also provides a detection device for wearing insulating gloves at a power grid operation site, including:

[0040] A channel feature map acquisition module, configured to perform channel feature extraction on an image of a person at a power grid operation site according to a neural network feature extraction module of an insulating glove wearing detection model to obtain a channel feature map.

[0041] An insulating glove feature map acquisition module, configured to extract the hand feature and the insulating glove feature of the person in the channel feature map according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model to obtain an insulating glove feature map; the channel spatial feature enhancement attention module is constructed by combining global pooling operation with a parameter-free attention mechanism.

[0042] A feature enhancement map acquisition module, configured to perform local and global feature extraction on the insulating glove feature map according to the multi-dilation aggregation module of the insulating glove wearing detection model and then perform fusion processing to obtain an insulating glove feature enhancement map.

[0043] A detection result acquisition module, configured to determine whether a power grid construction worker in the image of the person at the power grid operation site wears insulating gloves, whether the insulating gloves are worn correctly, and the target position in the image of the person at the power grid operation site according to the enhanced image of the insulating gloves, so as to obtain the detection result of insulating glove wearing in the image of the person at the power grid operation site.

[0044] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor executes the above method.

[0045] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and the computer program is executed by a processor to perform the above method.

[0046] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and the computer program is executed by a processor to perform the above method.

[0047] The above-mentioned detection method, device, computer equipment, computer-readable storage medium and computer program product for wearing insulating gloves at the power grid operation site extract channel features from the image of personnel at the power grid operation site according to the neural network feature extraction module of the insulating glove wearing detection model to obtain a channel feature map; according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model, extract the hand features of the personnel and the features of the insulating gloves in the channel feature map to obtain an insulating glove feature map; the channel spatial feature enhancement attention module is constructed by combining global pooling operation with a parameter-free attention mechanism; according to the multi-dilation aggregation module of the insulating glove wearing detection model, perform local and global feature extraction on the insulating glove feature map and then perform fusion processing to obtain an enhanced insulating glove feature map; according to the enhanced insulating glove feature map, determine whether the power grid construction personnel in the image of personnel at the power grid operation site wear insulating gloves, whether the insulating gloves are worn correctly, and the target position in the image of personnel at the power grid operation site, so as to obtain the detection result of wearing insulating gloves in the image of personnel at the power grid operation site. In this application, a channel spatial feature enhancement attention module of the insulating glove wearing detection model is constructed by combining global pooling operation with a parameter-free attention mechanism; according to this channel spatial feature enhancement attention module, extract the hand features of the personnel and the features of the insulating gloves in the channel feature map to obtain an insulating glove feature map; it can enhance the global context information of the insulating glove feature map and improve the feature extraction ability for fuzzy targets, thereby improving the detection accuracy of wearing insulating gloves; in the multi-dilation aggregation module of the insulating glove wearing detection model, by using dilated convolution operation and performing channel splicing and fusion on the insulating glove feature map, it can better capture the local and global relationships in the space of the insulating glove feature map, effectively handle the occlusion and angle change of the target, and further improve the detection accuracy of wearing insulating gloves. Description of the Drawings

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0049] Figure 1 It is an application environment diagram of the detection method for wearing insulating gloves at the power grid operation site in an embodiment;

[0050] Figure 2 It is a flowchart of the detection method for wearing insulating gloves at the power grid operation site in an embodiment;

[0051] Figure 3 It is a structural diagram of the channel spatial feature enhancement attention module in an embodiment;

[0052] Figure 4 It is a schematic structural diagram of a multi-expansion convergence module in an embodiment;

[0053] Figure 5 It is a schematic flowchart of a detection method for wearing insulating gloves at the power grid operation site in another embodiment;

[0054] Figure 6 It is a schematic flowchart of a rapid and effective detection algorithm for insulating gloves in an embodiment;

[0055] Figure 7 It is a structural block diagram of a detection device for wearing insulating gloves at the power grid operation site in an embodiment;

[0056] Figure 8 It is an internal structure diagram of a computer device in an embodiment. Specific embodiments

[0057] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] The detection method for wearing insulating gloves at the power grid operation site provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. The terminal 102 can input the image of the personnel at the power grid operation site into the insulating glove wearing detection model to obtain the detection result of wearing insulating gloves for the image of the personnel at the power grid operation site. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0059] In an exemplary embodiment, as Figure 2 shown, a detection method for wearing insulating gloves at the power grid operation site is provided. Taking the method applied to Figure 1 the terminal 102 in as an example, it includes the following steps S201 to step S204. Among them:

[0060] Step S201: According to the neural network feature extraction module of the insulating glove wearing detection model, extract the channel features of the image of the personnel at the power grid operation site to obtain a channel feature map.

[0061] The insulating glove wearing detection model can detect whether the power grid construction personnel in the image of the personnel at the power grid operation site wear insulating gloves. The insulating glove wearing detection model includes a neural network feature extraction module, a channel spatial feature enhancement attention module, and a multi-dilation aggregation module.

[0062] Images of power grid construction personnel during live work at the power grid operation site and during operation and maintenance can be collected as images of the personnel at the power grid operation site.

[0063] The neural network feature extraction module can be PConv_BN_ELU, which consists of three parts: a convolution layer (Convolution, Conv), a batch normalization layer (Batch Normalization, BN), and an activation function layer (Exponential Linear Unit, ELU).

[0064] The image of the personnel at the power grid operation site can be input into the insulating glove wearing detection model, so that the neural network feature extraction module of the insulating glove wearing detection model extracts the channel features of the image of the personnel at the power grid operation site to obtain a channel feature map.

[0065] Step S202: According to the channel spatial feature enhancement attention module of the insulating glove wearing detection model, extract the hand features of the personnel and the features of the insulating gloves in the channel feature map to obtain an insulating glove feature map; the channel spatial feature enhancement attention module is constructed based on the global pooling operation combined with a parameter-free attention mechanism.

[0066] It can be constructed based on the global pooling operation combined with a parameter-free attention mechanism (Similarity-based AttentionModule, SimAM) as Figure 3The shown Channel-Space Features Enhanced Attention Module (CSFE_AM). Among them, Global Average Pooling represents the global average pooling operation, Global Max Pooling represents the global maximum pooling operation, Conv represents the convolution operation, SimAM represents the parameter-free attention mechanism, and Sigmoid represents the S-shaped activation function. The Channel-Space Features Enhanced Attention Module effectively reduces the number of parameters in the insulating glove wearing detection model. At the same time, it can identify and strengthen the feature channels important for identifying the position of the insulating gloves, and suppress redundant or noisy channels. At the spatial feature level, the Channel-Space Features Enhanced Attention Module fuses the position details of the insulating gloves in the low-order feature map with the rich context semantic information in the high-order feature map, enabling the insulating glove wearing detection model to more accurately locate and identify whether the power grid construction personnel are wearing insulating gloves correctly.

[0067] Among them, the parameter-free attention mechanism is a lightweight and parameter-free attention mechanism. Based on the spatial inhibition theory in neuroscience, it infers the importance of each neuron in the feature map by optimizing an energy function, thereby generating three-dimensional attention weights. Different from traditional attention mechanisms, the parameter-free attention mechanism can effectively improve the representation ability and performance of Convolutional Neural Networks (CNNs) without adding any parameters to the original network. The implementation process of the parameter-free attention mechanism is simple and efficient, avoiding complex structural adjustments and heuristic methods, making its application in various visual tasks more flexible and convenient.

[0068] According to the Channel-Space Features Enhanced Attention Module of the insulating glove wearing detection model, it is possible to identify and strengthen the feature channels important for identifying the position of the insulating gloves, suppress redundant or noisy channels, so as to accurately extract the hand features of the personnel and the features of the insulating gloves in the channel feature map, and obtain the insulating glove feature map.

[0069] Step S203, according to the Multi-dilatation Aggregation Module of the insulating glove wearing detection model, perform local and global feature extraction on the insulating glove feature map and then perform fusion processing to obtain the insulating glove feature enhanced map.

[0070] The Multi-dilatation Aggregation Module (MDAM) of the insulating glove wearing detection model is as Figure 4As shown, where MaxPool represents the max pooling operation, Conv represents the convolution operation, and SiLU represents the Sigmoid linear unit activation function. The multi-dilation aggregation module uses consecutive dilated convolution operations to extract local and global features from the insulating glove feature map. By fusing multi-scale feature maps, local and global feature fusion processing is performed to obtain an enhanced insulating glove feature map, which can reduce the impact of reduced detection accuracy caused by occlusion and angle changes.

[0071] Among them, dilated convolution is a special convolution method that can obtain a larger receptive field without increasing the computational cost and reducing the resolution of the feature map.

[0072] Step S204: Based on the enhanced insulating glove feature map, determine whether the power grid construction personnel in the image of the personnel at the power grid operation site are wearing insulating gloves, whether the insulating gloves are worn correctly, and the target position in the image of the personnel at the power grid operation site, so as to obtain the detection result of insulating glove wearing in the image of the personnel at the power grid operation site.

[0073] The detection result of insulating glove wearing can be displayed on the image of the personnel at the power grid operation site, including whether the insulating gloves are worn correctly, the target position in the image of the personnel at the power grid operation site, and the detection confidence.

[0074] Based on the enhanced insulating glove feature map, it is possible to determine whether the power grid construction personnel in the image of the personnel at the power grid operation site are wearing insulating gloves, whether the insulating gloves are worn correctly, and the target position in the image of the personnel at the power grid operation site, so as to obtain the detection result of insulating glove wearing.

[0075] In the above method for detecting insulating glove wearing at the power grid operation site, a channel-spatial feature enhanced attention module of the insulating glove wearing detection model is constructed according to the global pooling operation combined with the parameter-free attention mechanism; according to this channel-spatial feature enhanced attention module, the hand features of the personnel and the features of the insulating gloves in the channel feature map are extracted to obtain the insulating glove feature map; it can enhance the global context information of the insulating glove feature map and improve the feature extraction ability for fuzzy targets, thereby improving the detection accuracy of insulating glove wearing; in the multi-dilation aggregation module of the insulating glove wearing detection model, by using dilated convolution operations and channel splicing fusion of the insulating glove feature map, the local and global relationships in the space of the insulating glove feature map can be better captured, effectively coping with target occlusion and angle changes, thereby further improving the detection accuracy of insulating glove wearing.

[0076] In one embodiment, according to the channel - spatial feature enhanced attention module of the insulating glove wearing detection model, the hand features of the person and the features of the insulating glove are extracted from the channel feature map to obtain the insulating glove feature map. The specific steps are as follows: According to the channel - spatial feature enhanced attention module of the insulating glove wearing detection model, a convolution operation is performed on the channel feature map to extract the hand features of the person and the features of the insulating glove in the channel feature map, obtaining a first intermediate feature map; According to the channel - spatial feature enhanced attention module of the insulating glove wearing detection model, global average pooling operation and global max - pooling operation are respectively performed on the channel feature map to extract the hand features of the person and the features of the insulating glove in the channel feature map, obtaining a second intermediate feature map and a third intermediate feature map to obtain the channel feature weight; According to the channel feature map and the channel feature weight, a fifth intermediate feature map is obtained; According to the first intermediate feature map and the channel feature weight, a sixth intermediate feature map is obtained; The channel feature map is input into a parameter - free attention mechanism to obtain a seventh intermediate feature map; The fifth intermediate feature map, the sixth intermediate feature map, and the seventh intermediate feature map are added element - by - element to obtain an eighth intermediate feature map, and a convolution operation is performed on the eighth intermediate feature map to obtain the insulating glove feature map.

[0077] According to the channel - spatial feature enhanced attention module of the insulating glove wearing detection model, a convolution operation with a convolution kernel size k of 3, a stride s of 1, and a padding value p of 1 can be performed on the channel feature map F0 to extract the hand features of the person and the features of the insulating glove in the channel feature map, obtaining a first intermediate feature map F1 with an unchanged spatial resolution size and a channel number of C. Where the size of the channel feature map F0 is C×H×W, C represents the number of channels, H represents the height, and W represents the width.

[0078] According to the channel - spatial feature enhanced attention module of the insulating glove wearing detection model, a global average pooling operation can be performed on the channel feature map F0 to extract the hand features of the person and the features of the insulating glove in the channel feature map, obtaining a second intermediate feature map F2 with a spatial resolution size of 1×1 and a channel number of C; According to the channel - spatial feature enhanced attention module of the insulating glove wearing detection model, a global max - pooling operation can be performed on the channel feature map F0 to extract the hand features of the person and the features of the insulating glove in the channel feature map, obtaining a third intermediate feature map F3 with a spatial resolution size of 1×1 and a channel number of C; The channel feature weight W1 can be obtained according to the second intermediate feature map F2 and the third intermediate feature map F3.

[0079] The channel feature map F0 and the channel feature weight W1 can be dot - multiplied to obtain the fifth intermediate feature map F5 with adaptively enhanced channel features; the first intermediate feature map F1 and the channel feature weight W1 can be element - by - element multiplied and dot - multiplied to obtain the sixth intermediate feature map F6 with adaptively enhanced channel features, where the spatial resolution of the sixth intermediate feature map F6 is H×W and the number of channels is C; the channel feature map F0 can be input into a parameter - free attention mechanism to obtain the seventh intermediate feature map F7 with adaptively enhanced spatial features, where the spatial resolution of the seventh intermediate feature map F7 is H×W and the number of channels is C.

[0080] The fifth intermediate feature map F5, the sixth intermediate feature map F6, and the seventh intermediate feature map F7 can be added element - by - element to obtain the eighth intermediate feature map F8 with a spatial resolution of H×W and a number of channels of C, and a convolution operation with a convolution kernel size k of 3, a stride s of 1, and a padding value p of 1 is performed on the eighth intermediate feature map F8 to obtain the insulating glove feature map F9 with a spatial resolution of H×W and a number of channels of C. Exemplarily, the spatial resolution can be 32×32 and the number of channels can be 256.

[0081] In this embodiment, according to the channel - space feature enhancement attention module of the insulating glove wearing detection model, the hand features of the person and the features of the insulating glove in the channel feature map are extracted to obtain the insulating glove feature map. Among them, the channel features are extracted through global average pooling and global maximum pooling, and the spatial features are strengthened by combining with a parameter - free attention mechanism, so that smaller - sized people in the distance in the power grid operation site image can be captured more, enabling the insulating glove wearing detection model to more accurately locate and identify whether the power grid construction personnel are wearing insulating gloves correctly.

[0082] In one of the embodiments, to obtain the channel feature weight, the specific steps are as follows: The second intermediate feature map and the third intermediate feature map are added element - by - element to obtain the fourth intermediate feature map; according to the fourth intermediate feature map and the activation function, the channel feature weight is obtained.

[0083] The second intermediate feature map F2 and the third intermediate feature map F3 can be added element - by - element to obtain the fourth intermediate feature map F4; the fourth intermediate feature map F9 can be calculated through a layer of Sigmoid activation function to obtain the channel feature weight W1.

[0084] In this embodiment, the second intermediate feature map and the third intermediate feature map are added element - by - element to obtain the fourth intermediate feature map to obtain the channel feature weight, so that the channel - space feature enhancement attention module can, according to the channel feature weight, strengthen the feature channels important for identifying the position of the insulating glove and suppress redundant or noisy channels.

[0085] In one embodiment, according to the multi-dilation pooling module of the insulating glove wearing detection model, local and global features of the insulating glove feature map are extracted and then fused to obtain an enhanced insulating glove feature map. The specific steps are as follows: According to the multi-dilation pooling module of the insulating glove wearing detection model, a max-pooling operation is performed on the insulating glove feature map to obtain a first intermediate enhanced feature map; According to the multi-dilation pooling module of the insulating glove wearing detection model, consecutive multiple dilated convolution operations are performed on the insulating glove feature map to obtain a second intermediate enhanced feature map, a third intermediate enhanced feature map, and a fourth intermediate enhanced feature map; The first intermediate enhanced feature map, the second intermediate enhanced feature map, the third intermediate enhanced feature map, and the fourth intermediate enhanced feature map are concatenated along the channel dimension to obtain a fifth intermediate enhanced feature map; A convolution operation is performed on the fifth intermediate enhanced feature map to obtain a sixth intermediate enhanced feature map; According to the non-linear activation function, non-linear enhancement is performed on the sixth intermediate enhanced feature map to obtain the enhanced insulating glove feature map.

[0086] According to the multi-dilation pooling module of the insulating glove wearing detection model, a max-pooling operation that does not change the spatial resolution size and the number of channels can be performed on the insulating glove feature map C0 to extract the local features of the insulating glove feature map C0, and a first intermediate enhanced feature map C1 with a spatial resolution size of H×W and a number of channels of C is obtained. Here, the insulating glove feature map F9 output by the channel spatial feature enhancement attention module is denoted as C0, and the size of the insulating glove feature map C0 is C×H×W, where C represents the number of channels, H represents the height, and W represents the width. Exemplarily, the spatial resolution size can be 128×128, and the number of channels can be 64.

[0087] According to the multi-dilation pooling module of the insulating glove wearing detection model, consecutive multiple dilated convolution operations can be performed on the insulating glove feature map C0 to extract the local and global features of the insulating glove feature map C0. Specifically, a dilated convolution operation with a convolution kernel size k of 3, a stride s of 1, a padding value p of 1, and a dilation value of 1 is performed on the insulating glove feature map C0 to obtain a second intermediate enhanced feature map C2 with an unchanged spatial resolution size and a number of channels of C; A dilated convolution operation with a convolution kernel size k of 3, a stride s of 1, a padding value p of 3, and a dilation value of 3 can be performed on the second intermediate enhanced feature map C2 to obtain a third intermediate enhanced feature map C3 with an unchanged spatial resolution size and a number of channels of C; A dilated convolution operation with a convolution kernel size k of 3, a stride s of 1, a padding value p of 1, and a dilation value of 1 can be performed on the third intermediate enhanced feature map C3 to obtain a fourth intermediate enhanced feature map C4 with an unchanged spatial resolution size and a number of channels of C.

[0088] The first intermediate feature enhancement map C1, the second intermediate feature enhancement map C2, the third intermediate feature enhancement map C3, and the fourth intermediate feature enhancement map C4 can be concatenated along the channel dimension to obtain a fifth intermediate feature enhancement map C5 with a spatial resolution of H×W and a channel number of 4C, so as to realize the fusion processing of the local features and global features of the insulating glove feature map C0; a convolution operation with a convolution kernel size k of 1, a stride s of 1, and a padding value p of 0 can be performed on the fifth intermediate feature enhancement map C5 to obtain a sixth intermediate feature enhancement map C6 with an unchanged spatial resolution and a channel number of C; according to the Sigmoid Linear Unit (SiLU) activation function, the sixth intermediate feature enhancement map C6 is non-linearly enhanced to obtain the insulating glove feature enhancement map C7.

[0089] In this embodiment, according to the multi-dilation pooling module of the insulating glove wearing detection model, a maximum pooling operation, multiple consecutive dilated convolution operations, and a convolution operation are performed on the insulating glove feature map. After extracting the local features and global features of the insulating glove feature map and performing fusion processing, the insulating glove feature enhancement map is obtained, which can filter out interference information, enhance small target feature information, improve the problem that target feature information is submerged during the convolution process, and solve the problem of target occlusion in the insulating glove wearing detection task.

[0090] In one of the embodiments, the method provided in this application further includes: obtaining an insulating glove wearing detection training set and an insulating glove wearing detection validation set; inputting the insulating glove wearing detection training images in the insulating glove wearing detection training set into the model with parameters to be adjusted to obtain a first prediction result; according to the first prediction result, the annotation data corresponding to the insulating glove wearing detection training images, and the adaptive intersection over union loss function, obtaining a first training loss value; the adaptive intersection over union loss function is obtained according to the intersection over union loss, the center point distance loss, and the length and width loss; inputting the insulating glove wearing detection validation images in the insulating glove wearing detection validation set into the model with parameters to be adjusted to obtain a second prediction result; according to the second prediction result, the annotation data corresponding to the insulating glove wearing detection validation images, and the adaptive intersection over union loss function, obtaining a second training loss value; according to the first training loss value and the second training loss value, adjusting the parameters of the model with parameters to be adjusted until the second training loss value converges, and obtaining the insulating glove wearing detection model.

[0091] The insulating glove wearing detection training set and the insulating glove wearing detection validation set can be obtained.

[0092] The channel spatial feature enhancement attention module of the model with parameters to be adjusted can be constructed according to the global pooling operation combined with the parameter-free attention mechanism. The multi-dilation pooling module of the model with parameters to be adjusted can be obtained according to the dilated convolution operation and the multi-scale feature map fusion operation.

[0093] Before model training, the parameters of the model to be adjusted can be initialized, and the hyperparameters related to the model to be adjusted can be set, such as the number of training epochs, the size of the batch, the choice of optimizer, the size of the learning rate, and the maximum value of the gradient clipping strategy.

[0094] The training images of insulating glove wearing detection in the insulating glove wearing detection training set can be input into the model with parameters to be adjusted to obtain a first prediction result; according to the Adaptive Intersection over Union (AD-IOU) loss function, the gap between the first prediction result and the labeled data corresponding to the insulating glove wearing detection training image can be measured to obtain a first training loss value.

[0095] The Adaptive Intersection over Union loss function is based on the Intersection over Union (IOU) loss function and adds a loss metric and , to solve the problem of inaccurate bounding box regression. The Adaptive Intersection over Union loss function is , and weighted sum, as shown in Equation (1).

[0096] (1)

[0097] (2)

[0098] (3)

[0099] (4)

[0100] where is the loss based on the Intersection over Union, is the loss based on the distance between the center points, is the loss based on the length and width.

[0101] In the loss of the Intersection over Union, represents 1 minus the Intersection over Union between the predicted box and the ground truth box, that is is the ratio of the overlapping part of the two boxes to their union. The smaller the value of , that is, the larger the

[0102] in , where , where is a hyperparameter, and are the center coordinates of the predicted bounding box and the ground truth bounding box respectively, represents the Euclidean distance between the center coordinate of the predicted bounding box and the center coordinate of the ground truth bounding box, and are the squares of the width and height of the ground truth bounding box respectively. The loss function is used to reduce the distance between the center coordinate of the predicted bounding box and the center coordinate of the ground truth bounding box, making the predicted bounding box closer to the ground truth bounding box.

[0103] In , and are the width and height of the predicted bounding box respectively. represents taking the absolute value. The loss function is used to reduce the difference between the length and width of the predicted bounding box and the length and width of the ground truth bounding box.

[0104] The detection verification images of insulating glove wearing in the insulating glove wearing detection verification set can be input into the model with parameters to be adjusted to obtain the second prediction result; according to the Adaptive Intersection over Union (AD-IOU) loss function, the gap between the second prediction result and the annotation data corresponding to the insulating glove wearing detection training images can be measured to obtain the second training loss value.

[0105] According to the first training loss value and the second training loss value, the parameters of the model with parameters to be adjusted can be adjusted until the second training loss value converges, and the insulating glove wearing detection model can be obtained.

[0106] In this embodiment, the insulating glove wearing detection training images in the insulating glove wearing detection training set are input into the model with parameters to be adjusted to obtain the first prediction result, so as to obtain the first training loss value; the insulating glove wearing detection verification images in the insulating glove wearing detection verification set are input into the model with parameters to be adjusted to obtain the second prediction result, so as to obtain the second training loss value; according to the first training loss value and the second training loss value, the parameters of the model with parameters to be adjusted are adjusted until the second training loss value converges, and a better-performing insulating glove wearing detection model can be obtained.

[0107] In one embodiment, an insulating glove wearing detection training set and an insulating glove wearing detection validation set are obtained. The specific steps are as follows: Collect real images of insulating gloves worn at the power grid operation site and during operation and maintenance; Obtain a first simulated image of an insulating glove worn in an irregular manner and a second simulated image of no insulating glove worn; Based on the real images, the first simulated image, and the second simulated image, obtain an insulating glove wearing detection data set; Perform wearing category annotation on the insulating glove wearing detection data set to obtain an insulating glove wearing detection annotation data set; The wearing categories include the hands correctly wearing insulating gloves, the hands not wearing insulating gloves, and the hands wearing insulating gloves in an irregular manner; According to a set ratio, divide the insulating glove wearing detection annotation data set into an insulating glove wearing detection training set and an insulating glove wearing detection validation set.

[0108] At the power grid operation site and during operation and maintenance, use a high-definition camera device to photograph power grid construction workers, and try to ensure that multiple power grid construction workers appear in the collected images to reflect various insulating glove wearing detection scenarios, so as to ensure the diversity and representativeness of the data.

[0109] To improve the accuracy of the model in identifying the wearing situation of insulating gloves, negative samples can be introduced. By simulating the situation of not wearing insulating gloves and the situation of wearing insulating gloves in an irregular manner, such as the glove slipping off and the fingers being exposed, a certain number of negative samples can be created to obtain a first simulated image of an insulating glove worn in an irregular manner and a second simulated image of no insulating glove worn.

[0110] Based on the real images, the first simulated image, and the second simulated image, obtain an insulating glove wearing detection data set. For each image in the insulating glove wearing detection data set, a professional annotation tool can be used to finely annotate the hand area of the power grid construction worker in the image to obtain an insulating glove wearing detection annotation data set. The annotation information includes the accurate position of the hand area and the wearing category of the insulating glove. A corresponding label file can be generated for each image, and the label file details the position information of the hand area and the wearing category of the insulating glove. The wearing categories include the hands correctly wearing insulating gloves, the hands not wearing insulating gloves, and the hands wearing insulating gloves in an irregular manner.

[0111] A series of preprocessing operations can be performed on the images in the obtained insulating glove wearing detection annotation data set, including adjusting the image size to meet the model input requirements and adding noise to improve the generalization ability of the model to ensure that the image quality meets the needs of model training.

[0112] According to the actual needs, determine the set ratio; According to the set ratio, divide the data set into an insulating glove wearing detection training set and an insulating glove wearing detection validation set according to the ratio, which are respectively used for model training and performance verification to ensure that the model can be fully evaluated and optimized at different stages.

[0113] In this embodiment, an insulation glove wearing detection data set is obtained based on a real image, a first simulated image, and a second simulated image, which can improve the diversity of the insulation glove wearing detection data set; according to a set ratio, the data set is divided into an insulation glove wearing detection training set and an insulation glove wearing detection validation set in proportion, so as to prepare data for the training and performance verification of the model.

[0114] To better understand the above method, the following elaborates in detail an application embodiment of the detection method for wearing insulation gloves at the power grid operation site of this application.

[0115] In the power industry, the safety management of the power grid operation site is of crucial importance. The power grid construction site involves the operation or maintenance of high-voltage electrical equipment. Especially when carrying out live work, power grid construction personnel should wear insulation gloves that meet the requirements of the corresponding voltage level to ensure personal safety. However, some power grid construction personnel may not wear insulation gloves for various reasons, which requires on-site management personnel to conduct patrol inspections, but it is also difficult to ensure that each construction worker will comply with safety regulations. With the rapid development of deep learning technology, especially the remarkable progress of object detection technology, a new solution has been provided for the safety management of the power grid operation site.

[0116] The object detection algorithm based on deep learning can efficiently and accurately identify specific objects, such as personnel, tools, and safety equipment, by automatically analyzing image or video data. In the task of detecting whether insulating gloves are worn at the power grid operation site, this type of object detection algorithm can capture the images of construction workers in real time and automatically detect whether they are wearing insulating gloves and whether the wearing is standard. Generally speaking, the basic idea of the method for detecting whether insulating gloves are worn at the power grid operation site based on deep learning is to use the construction images captured by the monitoring cameras at the power grid operation site as the input, and then preprocess the collected images to optimize the image quality. Then, by using the feature extraction ability of the deep convolutional neural network, the key features related to the wearing of insulating gloves are extracted from the preprocessed images. On the basis of feature extraction, advanced object detection models are adopted. Common models include the RCNN series, YOLO, SSD, etc. Among them, the full English name of R-CNN is Region-based Convolutional Neural Network, and the full Chinese name is Region-based Convolutional Neural Network. The full English name of YOLO is You Only Look Once, and the full English name of SSD is Single Shot MultiBox Detector, and the full Chinese name is Single Shot MultiBox Detector. These models can predict in real time the presence or absence of insulating gloves and whether the wearing status meets the safety standards. Finally, the detection results are presented in an intuitive way, such as marking the positions of construction workers who are not wearing or wearing non-standardly, and generating corresponding reports or alarms to facilitate timely corrective measures.

[0117] In summary, the method for detecting whether insulating gloves are worn at the power grid operation site based on deep learning, with its advantages of high efficiency, accuracy, and real-time performance, provides strong technical support for the safety management of the power grid operation site and effectively reduces the risk of safety accidents.

[0118] The current method for detecting whether insulating gloves are worn at the power grid operation site based on deep learning is as follows:

[0119] (1) A method for detecting the unsafe behavior of not wearing insulating gloves is proposed based on YOLOv5. First, the convolution in the feature extraction network is replaced with self-calibrated convolution, so that the network pays more attention to the information around the target to be detected when extracting target features. Then, a Selective Kernel (SK) attention mechanism is added at the end of the feature fusion network to make the network pay more attention to the small targets to be detected. Finally, the loss function in the original YOLOv5 network is modified to further accurately detect the position of the bounding box, and finally improve the detection accuracy of insulating gloves. Disadvantages of this method: For the cases of personnel overlap and hand occlusion, this algorithm may miss detections.

[0120] (2) Obtain the original image data in the live working scenario; the original image data includes the target object; preprocess the original image data to obtain the target image data; input the target image data into the trained insulation glove wearing detection model, and use the insulation glove wearing detection model to detect whether the target object in the original image data wears an insulation glove to obtain the target detection result; wherein, the insulation glove wearing detection model includes a feature extraction network, a feature fusion network and a prediction network. This method can effectively improve the accuracy of insulation glove wearing detection, and has higher detection efficiency, can reduce the human resource cost of supervision, and is beneficial to improving the safety of live working.

[0121] The current insulation glove wearing detection method based on deep learning in the power grid operation site mainly has the following problems:

[0122] (1) Insufficient model detection accuracy: The existing object detection algorithms have relatively low positioning accuracy when identifying the wearing state of insulation gloves. Due to the complex and variable background of the power grid construction site, when directly using the existing general models for detection, the single grid prediction box used by them is difficult to accurately locate the specific area where the gloves are worn, resulting in a large positioning error and making it difficult to accurately judge whether the staff wears insulation gloves.

[0123] (2) Difficulty in recognizing small target features: The wearing detection of insulation gloves involves the complete recognition of the human hand and the glove. When the power grid construction personnel are far away from the detection device, these features may appear as relatively small pixel areas in the image. The existing object detection algorithms have difficulties in capturing these subtle features and are prone to ignoring key information, resulting in misjudgment of the wearing state.

[0124] (3) Influence of occlusion and angle change: In the actual working environment, the worker's hand may be partially occluded or the angle may be changed due to factors such as the operating posture and occluders, which makes the wearing detection of insulation gloves more complex. The current technology often has difficulty in effectively coping with this occlusion and angle change, resulting in difficulty in accurately detecting the wearing situation of gloves in complex scenarios and increasing the risk of missed detection and false detection.

[0125] To solve the above problems, the technical solution provided in this embodiment offers an Effective and Fast Detection of Insulated Gloves (EFDIG) algorithm, which includes an Attention Module for Channel and Spatial Feature Enhancement (CSFE_AM). Channel features are extracted through global average pooling and global max pooling, and spatial features are strengthened by combining a parameter-free attention mechanism, enabling the network to capture more small-sized figures in the distance in the images of the power grid operation site, thus solving the problem of difficult small target detection in the task of insulated glove wearing detection. Meanwhile, the Multi-expansion aggregation module (MDAM) filters out interference information and enhances small target feature information through multiple dilated convolutions, improving the problem that target feature information is submerged during the convolution process and solving the problem of target occlusion in the task of insulated glove wearing detection. To address the problem of missed detection where individual insulated glove wearing situations are not detected, an Adaptive Intersection over Union (AD-IOU) loss function is introduced as the loss function of this algorithm. The above improvements effectively solve the problem of insufficient accuracy of the traditional insulated glove wearing detection model for small target detection and improve the detection accuracy of insulated glove wearing.

[0126] The specific steps of the technical solution provided in this embodiment include constructing a dataset for training insulated glove wearing detection, constructing an attention module for channel and spatial feature enhancement, obtaining a multi-expansion aggregation module, using a Feature Pyramid Network (FPN) module for feature fusion, obtaining an adaptive intersection over union loss function, model training and verification, and applying the model. The corresponding method flow chart is as Figure 5 shown.

[0127] The EFDIG algorithm structure of this embodiment is as Figure 6As shown. For the images of personnel at the power grid operation site, first, a two-layer neural network feature extraction module (PConv_BN_ELU) is used to obtain the channel feature map. Subsequently, the obtained channel feature map is input into the channel spatial feature enhancement attention module (CSFE_AM). The channel spatial feature enhancement attention module can filter out interference information and enhance the small target feature information, enhancing the feature extraction ability of the insulating glove wearing detection model for the entire image of personnel at the power grid operation site, especially the hand features and the features of insulating gloves of power grid construction personnel far from the video detection device. The multi-dilation aggregation module (MDAM) takes the insulating glove feature map output by the channel spatial feature enhancement attention module as input, first performs multi-scale feature extraction, and then fuses the feature maps of different scales to obtain an enhanced insulating glove feature map, enhancing the feature expression of the insulating gloves worn on the hands of power grid construction personnel in the image of personnel at the power grid operation site. Then, the enhanced insulating glove feature map is successively passed through a one-layer neural network feature extraction module, a channel spatial feature enhancement attention module, and a multi-dilation aggregation module, repeating the above operations twice. After that, the enhanced insulating glove feature maps extracted by each multi-domain adaptation module are respectively passed through the feature pyramid network (FPN), and then through the adaptive intersection over union loss function (AD-IOU) to improve the problem of false detection and missed detection of insulating gloves, and finally obtain the detection result of insulating glove wearing in the image of personnel at the power grid operation site.

[0128] S1. Construct an insulating glove wearing detection dataset for training:

[0129] In the actual scenarios of power grid construction operations and daily operation and maintenance, real images of insulating glove wearing and use can be collected. In addition, experimenters can simulate non-standard insulating glove wearing methods and the situation of not wearing insulating gloves to obtain the first simulated images of non-standard insulating glove wearing and the second simulated images of not wearing insulating gloves, so as to create a certain number of negative samples to balance the number of positive and negative samples.

[0130] Based on the real images, the first simulated images, and the second simulated images, an insulating glove wearing detection dataset is obtained. For each image in the insulating glove wearing detection dataset, a professional annotation tool can be used to finely annotate the hand area of the power grid construction personnel in the image to obtain an insulating glove wearing detection annotation dataset. The annotation information includes the accurate position of the hand area and the wearing category state of the insulating gloves. A corresponding label file can be generated for each image, and the label file details the position information of the hand area and the wearing category of the insulating gloves. The wearing categories include correctly wearing insulating gloves on the hands, not wearing insulating gloves on the hands, and non-standard wearing of insulating gloves.

[0131] Through such dataset construction and annotation work, staff can more comprehensively understand and analyze the wearing situation of insulating gloves of power grid construction workers during power grid construction operations, thereby providing strong data support for improving operation safety and standardizing operation behaviors.

[0132] Specific examples of constructing an insulating glove wearing detection dataset for training are as follows:

[0133] (1) Dataset collection: In the actual working environment, use high-definition camera equipment to photograph power grid construction workers, and try to ensure that multiple power grid construction workers appear in the collected images to reflect various insulating glove wearing detection scenarios. 5400 real images in different working scenarios can be collected to ensure the diversity and representativeness of the data.

[0134] (2) Negative sample creation: To improve the accuracy of the model in identifying the wearing situation of insulating gloves, negative samples are introduced. By simulating hand images without gloves and situations where gloves are worn irregularly, such as gloves slipping off and fingers being exposed, the first simulated images of irregularly worn insulating gloves and the second simulated images of no gloves worn are obtained to create a certain number of negative samples.

[0135] (3) Dataset annotation: Based on the real images, the first simulated images, and the second simulated images, an insulating glove wearing detection dataset is obtained. For each image in the insulating glove wearing detection dataset, use professional annotation tools to finely annotate the hand area of power grid construction workers. The annotation information includes the accurate position of the hand area and the wearing status of the insulating gloves.

[0136] (4) Label generation: Generate corresponding label files for each image in the insulating glove wearing detection dataset. The label files detail the position information of the hand area and the wearing status of the insulating gloves, which are divided into correctly worn, not worn, and irregular.

[0137] (5) Data preprocessing: Perform a series of preprocessing operations on the collected images, including resizing the images to meet the model input requirements, adding noise to improve the generalization ability of the model, and ensuring that the image quality meets the needs of model training.

[0138] (6) Dataset saving and partitioning: Organize and store the preprocessed images and corresponding label files in the insulating glove wearing detection annotation dataset. Subsequently, according to actual needs, determine the set ratio; according to the set ratio, divide the insulating glove wearing detection annotation dataset into a training set, a validation set, and a test set in proportion, which are used for model training, performance verification, and final testing respectively to ensure that the model can be fully evaluated and optimized at different stages.

[0139] S2. Construct a channel spatial feature enhanced attention module:

[0140] The channel-spatial feature enhanced attention module can effectively reduce the number of parameters in the insulating glove wearing detection model. At the same time, the channel-spatial feature enhanced attention module can identify and strengthen the feature channels important for the identification of the position of the insulating glove, and suppress redundant or noisy channels. At the spatial feature level, this module fuses the position details of the insulating glove in the low-order feature map with the rich context semantic information in the high-order feature map, enabling the insulating glove wearing detection model to more accurately locate and identify whether the construction worker is wearing the insulating glove correctly. The overall structure of the channel-spatial feature enhanced attention module is as Figure 3 shown.

[0141] The channel-spatial feature enhanced attention module takes the channel feature map output by the upper module neural network feature extraction module as the input feature map, and here the channel feature map is denoted as F0. The size of the channel feature map F0 is C×H×W, where C represents the number of channels, H represents the height, and W represents the width. The specific steps of the channel-spatial feature enhanced attention module are as follows:

[0142] S2.1: Perform a convolution operation on the channel feature map F0 with a convolution kernel size k of 3, a stride s of 1, and a padding value p of 1 to obtain the first intermediate feature map F1 with the same spatial resolution size and a channel number of C.

[0143] S2.2: Perform a global average pooling operation on the channel feature map F0 to extract the hand features of the person and the features of the insulating glove in the channel feature map, and obtain the second intermediate feature map F2 with a spatial resolution size of 1×1 and a channel number of C; perform a global maximum pooling operation on the channel feature map F0 to extract the hand features of the person and the features of the insulating glove in the channel feature map, and obtain the third intermediate feature map F3 with a spatial resolution size of 1×1 and a channel number of C; the second intermediate feature map F2 and the third intermediate feature map F3 can be added element-wise to obtain the fourth intermediate feature map F4; the fourth intermediate feature map F9 can be calculated through a sigmoid activation function layer to obtain the channel feature weight W1.

[0144] S2.3: Multiply F1 and W1 to obtain the intermediate feature map F_6 with enhanced channel feature adaptability, whose spatial resolution size is H×W and the channel number is C.

[0145] S2.4: Multiply the first intermediate feature map F1 and the channel feature weight W1 to obtain the sixth intermediate feature map F6 with enhanced channel feature adaptability, where the spatial resolution size of the sixth intermediate feature map F6 is H×W and the channel number is C.

[0146] S2.5: The channel feature map F0 can be input into a parameterless attention mechanism to obtain the seventh intermediate feature map F7 with spatially adaptively enhanced features. The spatial resolution of the seventh intermediate feature map F7 is H×W, and the number of channels is C.

[0147] S2.6: The fifth intermediate feature map F5, the sixth intermediate feature map F6, and the seventh intermediate feature map F7 can be added element-wise to obtain the eighth intermediate feature map F8 with a spatial resolution of H×W and a number of channels of C. Then, a convolution operation with a convolution kernel size k of 3, a stride s of 1, and a padding value p of 1 is performed on the eighth intermediate feature map F8 to obtain the insulating glove feature map F9 with a spatial resolution of H×W and a number of channels of C.

[0148] S3. Obtain the multi-dilation aggregation module:

[0149] Feature extraction of local and global features is performed on the input insulating glove feature map, and then local and global features are fused to reduce the impact of reduced detection accuracy caused by occlusion and angle changes. The structure of the multi-dilation aggregation module is as Figure 4 shown.

[0150] The insulating glove feature map output by the channel spatial feature enhancement attention module is used as the input feature map of the multi-dilation aggregation module. Here, the insulating glove feature map F9 output by the channel spatial feature enhancement attention module is denoted as C0. The size of the insulating glove feature map C0 is C×H×W, where C represents the number of channels, H represents the height, and W represents the width. The specific steps of the multi-dilation aggregation module are as follows:

[0151] S3.1: A max pooling operation that does not change the spatial resolution and the number of channels is performed on the insulating glove feature map C0 with a spatial resolution of H×W and a number of channels of C to obtain the first intermediate feature enhancement map C1 with a spatial resolution of H×W and a number of channels of C.

[0152] S3.2: A dilated convolution operation with a convolution kernel size k of 3, a stride s of 1, a padding value p of 1, and a dilation value of 1 is performed on the insulating glove feature map C0 to obtain the second intermediate feature enhancement map C2 with an unchanged spatial resolution and a number of channels of C.

[0153] S3.3: A dilated convolution operation with a convolution kernel size k of 3, a stride s of 1, a padding value p of 3, and a dilation value of 3 is performed on the second intermediate feature enhancement map C2 to obtain the third intermediate feature enhancement map C3 with an unchanged spatial resolution and a number of channels of C.

[0154] S3.4: Perform a dilated convolution operation with a convolution kernel size \(k = 3\), a stride \(s = 1\), a padding value \(p = 1\), and a dilation value of 1 on the third intermediate feature enhanced map \(C3\) to obtain a fourth intermediate feature enhanced map \(C4\) with the same spatial resolution size and a channel number of \(C\).

[0155] S3.5: Concatenate the first intermediate feature enhanced map \(C1\), the second intermediate feature enhanced map \(C2\), the third intermediate feature enhanced map \(C3\), and the fourth intermediate feature enhanced map \(C4\) along the channel dimension to obtain a fifth intermediate feature enhanced map \(C5\) with a spatial resolution size of \(H\times W\) and a channel number of \(4C\).

[0156] S3.6: Perform a convolution operation with a convolution kernel size \(k = 1\), a stride \(s = 1\), and a padding value \(p = 0\) on the fifth intermediate feature enhanced map \(C5\) to obtain a sixth intermediate feature enhanced map \(C6\) with the same spatial resolution size and a channel number of \(C\).

[0157] S3.7: Input the sixth intermediate feature enhanced map \(C6\) into the Sigmoid Linear Unit (SiLU) activation function for non - linear enhancement to obtain an insulating glove feature enhanced map \(C7\), which is used as the output of the multi - dilation pooling module.

[0158] S4. Use the Feature Pyramid Network module for feature fusion:

[0159] The Feature Pyramid Network (FPN) module is used to fuse insulating glove feature enhanced maps at different scales. The insulating glove feature enhanced maps 、 、 are the input feature maps of the FPN. The Feature Pyramid Network module performs feature fusion on 、 、 , that is, the insulating glove feature enhanced map fuses the upsampled insulating glove feature enhanced map of 、 . The fused insulating glove feature enhanced map is . The insulating glove feature enhanced map fuses the downsampled and the upsampled insulating glove feature enhanced map of . The fused insulating glove feature enhanced map is . The insulating glove feature enhanced map fuses the downsampled insulating glove feature enhanced maps of 、 . The fused insulating glove feature enhanced map is .

[0160] S5: Obtain the Adaptive Intersection over Union loss function:

[0161] The Adaptive Intersection over Union (AD-IOU) loss function is used to measure the gap between the prediction results of the insulating glove wearing detection model and the true label data.

[0162] The Adaptive Intersection over Union loss function is based on the Intersection over Union (IOU) loss function, with a loss metric added and , to address the problem of inaccurate Bounding Box regression. The Adaptive Intersection over Union loss function is , and 's weighted sum, as shown in Equation (1).

[0163] (1)

[0164] (2)

[0165] (3)

[0166] (4)

[0167] where is the loss based on Intersection over Union, is the loss based on the distance between the center points, is the loss based on the length and width.

[0168] In the loss of Intersection over Union, represents 1 minus the Intersection over Union between the predicted box and the true box, that is is the ratio of the overlapping part of the two boxes to their union. The smaller the value of , that is, the larger

[0169] is, the higher the overlapping degree of the predicted box and the true box, indicating the higher similarity between the predicted box and the true box. in , where is a hyperparameter, and are the center coordinates of the predicted box and the true box respectively, represents the Euclidean distance between the center coordinates of the predicted box and the true box, and are the squares of the width and height of the true box respectively. The loss function is used to reduce the distance between the center coordinates of the predicted bounding box and the center coordinates of the ground truth bounding box, making the predicted bounding box closer to the ground truth bounding box.

[0170] In , and are the width and height of the predicted bounding box respectively. denotes taking the absolute value. The loss function is used to reduce the difference between the length and width of the predicted bounding box and the length and width of the ground truth bounding box.

[0171] The loss calculation of IOU is as follows:

[0172] (5)

[0173] Wherein, and represent the ground truth bounding box and the predicted bounding box respectively, represents the intersection of the ground truth bounding box and the predicted bounding box, represents the union of the ground truth bounding box and the predicted bounding box.

[0174] Box regression refers to the process of adjusting the position and size of the bounding boxes predicted for each grid cell in the object detection task. For example, in the YOLO algorithm, the input image is divided into several grid cells, and each grid cell is responsible for predicting the objects whose center points fall within that grid. However, these initially predicted bounding boxes are often not accurate enough, so Box regression is needed to further adjust their positions and sizes to more accurately match the actual objects.

[0175] S6, Model Training and Validation:

[0176] Based on the channel spatial feature enhancement attention module and the multi-dilation aggregation module, a model with parameters to be adjusted is obtained. On the basis of the model with parameters to be adjusted, the parameters of each layer are trained and updated. Initialize all neural network parameters of the model with parameters to be adjusted, and set the hyperparameters related to the model with parameters to be adjusted, such as: the number of training epochs, the batch size, the selection of the optimizer, the learning rate, the maximum value of the gradient clipping strategy, etc.

[0177] After initializing the parameters, the insulation glove wearing detection dataset processed by S1 is divided into a training set, a validation set, and a test set. The training set and the validation set data are divided into multiple batches. Each time, a batch of training set data is input into the model with parameters to be adjusted for training, and the training loss value loss of this batch is obtained. When a round of training is completed for all batches of data in the entire training set (in the actual training process, there may be multiple rounds), the validation set is input into the model with parameters to be adjusted batch by batch, and the corresponding batch loss value batch loss is obtained. During training and validation, the model with parameters to be adjusted will automatically learn and adjust the parameters according to the situation of each loss and batch loss. When the training process has undergone one or more rounds until the batch loss value tends to converge, the training ends, and an insulation glove wearing detection model is obtained.

[0178] S7. Apply the model:

[0179] Apply the trained insulation glove wearing detection model to detect the images of personnel at the power grid operation site, and realize the automatic detection of the wearing of insulation gloves by power grid construction personnel at the power grid operation site.

[0180] In actual engineering applications, the collected images of personnel at the power grid operation site are input into the insulation glove wearing detection model. The insulation glove wearing detection model automatically detects the images of personnel at the power grid operation site, obtains the insulation glove wearing detection results of the images of personnel at the power grid operation site, and displays the detection results on the images of personnel at the power grid operation site, including whether the wearing is correct, the target position in the image, and the detection confidence.

[0181] Specific examples are as follows:

[0182] Use the insulation glove wearing detection dataset. The dataset contains 5400 images, including 16133 labeled objects and 3 categories, and is used to train an insulation glove wearing detection model.

[0183] (1) Dataset annotation: Annotate the wearing of insulation gloves on the hand areas of power grid construction personnel extracted from the insulation glove wearing detection dataset, and annotate the quality score for each insulation glove wearing situation.

[0184] (2) Dataset division: Randomly divide the dataset into a training set, a validation set, and a test set according to the ratio of 6:2:2, that is, there are 3240 training set pictures, and 1080 pictures for both the validation set and the test set.

[0185] (3) Model training: According to the data in the training set and the validation set, train an insulation glove wearing detection model. The entire training process is set to 450 epochs (rounds), and the parameters in the model are adjusted according to the loss value in each iteration process.

[0186] (4) Training results: After adjusting the relevant parameters, the accuracy of the final insulating glove wearing detection model tends to be stable, mainly fluctuating between 90% and 95%.

[0187] The technical solution provided in this embodiment has the following improvement points:

[0188] (1) The channel spatial feature enhanced attention module is constructed by using global pooling technology and combining a parameter-free attention mechanism at the same time, which enhances the global context information, improves the ability of the model backbone network to extract fuzzy targets, and has good detection ability for targets of different sizes.

[0189] (2) The multi-dilation convergence module can better capture the local and global relationships in the target feature space by using dilated convolution and channel splicing and fusion of feature maps.

[0190] (3) The adaptive intersection over union loss function comprehensively balances the loss calculations of position, shape, and overlap degree, and improves the positioning accuracy of insulating glove targets in complex environments;

[0191] The technical solution provided in this embodiment has the following beneficial effects:

[0192] (1) The channel spatial feature enhanced attention module extracts spatial features and spatial features at the same time, improving the model's ability to extract detailed features of the insulating glove wearing target on the hands of power grid construction personnel.

[0193] (2) By using multiple dilated convolutions and fusing multi-scale feature maps, the detection performance of the model for occluded insulating gloves in images of personnel at power grid operation sites is improved.

[0194] (3) The method of using the adaptive intersection over union loss function to improve the positioning accuracy of the insulating gloves worn on the hands of power grid construction personnel in complex power grid operation environments is achieved by calculating the Euclidean distance between the center points of the predicted box and the ground truth box and measuring the difference in the aspect ratio between the predicted box and the ground truth box.

[0195] In summary, the technical solution provided in this embodiment has higher detection accuracy, fewer parameters and calculations of the insulating glove wearing detection model, and can meet the high real-time requirements for the insulating glove wearing detection task in power grid operation scenarios.

[0196] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.

[0197] Based on the same inventive concept, an embodiment of the present application further provides a detection device for wearing insulating gloves at the grid operation site for implementing the detection method for wearing insulating gloves at the grid operation site involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the detection device for wearing insulating gloves at the grid operation site provided below can refer to the limitations on the detection method for wearing insulating gloves at the grid operation site in the above text, and will not be repeated here.

[0198] In an exemplary embodiment, as Figure 7 shown, a detection device for wearing insulating gloves at the grid operation site is provided, where:

[0199] The channel feature map acquisition module 701 is configured to perform channel feature extraction on the image of the personnel at the grid operation site according to the neural network feature extraction module of the insulating glove wearing detection model to obtain a channel feature map;

[0200] The insulating glove feature map acquisition module 702 is configured to extract the hand feature of the personnel and the feature of the insulating glove in the channel feature map according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model to obtain an insulating glove feature map; the channel spatial feature enhancement attention module is constructed by combining global pooling operation with a parameter-free attention mechanism;

[0201] The feature enhancement map acquisition module 703 is configured to perform local and global feature extraction on the insulating glove feature map according to the multi-dilation aggregation module of the insulating glove wearing detection model and then perform fusion processing to obtain an insulating glove feature enhancement map;

[0202] The detection result acquisition module 704 is configured to determine whether the grid construction personnel in the grid operation site personnel image wear insulating gloves, whether the insulating gloves are worn correctly, and the target position in the grid operation site personnel image according to the enhanced insulating glove feature map, so as to obtain the insulating glove wearing detection result of the grid operation site personnel image.

[0203] In one embodiment, the insulating glove feature map acquisition module 702 is further configured to: perform a convolution operation on the channel feature map according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model to extract the hand features of the personnel and the features of the insulating gloves in the channel feature map, so as to obtain a first intermediate feature map; perform global average pooling operation and global maximum pooling operation on the channel feature map respectively according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model to extract the hand features of the personnel and the features of the insulating gloves in the channel feature map, so as to obtain a second intermediate feature map and a third intermediate feature map, so as to obtain channel feature weights; obtain a fifth intermediate feature map according to the channel feature map and the channel feature weights; obtain a sixth intermediate feature map according to the first intermediate feature map and the channel feature weights; input the channel feature map into a parameter-free attention mechanism to obtain a seventh intermediate feature map; add the fifth intermediate feature map, the sixth intermediate feature map, and the seventh intermediate feature map element by element to obtain an eighth intermediate feature map, and perform a convolution operation on the eighth intermediate feature map to obtain an insulating glove feature map.

[0204] In one embodiment, the insulating glove feature map acquisition module 702 is further configured to: add the second intermediate feature map and the third intermediate feature map element by element to obtain a fourth intermediate feature map; obtain channel feature weights according to the fourth intermediate feature map and an activation function.

[0205] In one embodiment, the feature enhancement map acquisition module 703 is further configured to: perform a maximum pooling operation on the insulating glove feature map according to the multi-dilation aggregation module of the insulating glove wearing detection model to obtain a first intermediate feature enhancement map; perform continuous multiple dilation convolution operations on the insulating glove feature map according to the multi-dilation aggregation module of the insulating glove wearing detection model to obtain a second intermediate feature enhancement map, a third intermediate feature enhancement map, and a fourth intermediate feature enhancement map; splice the first intermediate feature enhancement map, the second intermediate feature enhancement map, the third intermediate feature enhancement map, and the fourth intermediate feature enhancement map in the channel dimension to obtain a fifth intermediate feature enhancement map; perform a convolution operation on the fifth intermediate feature enhancement map to obtain a sixth intermediate feature enhancement map; perform non-linear enhancement on the sixth intermediate feature enhancement map according to a non-linear activation function to obtain an insulating glove feature enhancement map.

[0206] In one embodiment, the device further includes a model training module, configured to: obtain an insulating glove wearing detection training set and an insulating glove wearing detection validation set; input the insulating glove wearing detection training images in the insulating glove wearing detection training set into a model with parameters to be adjusted, to obtain a first prediction result; obtain a first training loss value according to the first prediction result, the annotation data corresponding to the insulating glove wearing detection training images, and an adaptive intersection over union loss function; the adaptive intersection over union loss function is obtained according to an intersection over union loss, a center point distance loss, and a length-width loss; input the insulating glove wearing detection validation images in the insulating glove wearing detection validation set into the model with parameters to be adjusted, to obtain a second prediction result; obtain a second training loss value according to the second prediction result, the annotation data corresponding to the insulating glove wearing detection validation images, and the adaptive intersection over union loss function; adjust the parameters of the model with parameters to be adjusted according to the first training loss value and the second training loss value, until the second training loss value converges, to obtain an insulating glove wearing detection model.

[0207] In one embodiment, the model training module is further configured to: collect real images of insulating gloves worn at the power grid operation site and during operation and maintenance; obtain a first simulated image of an insulating glove worn in an irregular manner and a second simulated image of no insulating glove worn; obtain an insulating glove wearing detection data set according to the real images, the first simulated image, and the second simulated image; perform wearing category annotation on the insulating glove wearing detection data set, to obtain an insulating glove wearing detection annotation data set; the wearing categories include correct wearing of insulating gloves on the hand, no insulating glove worn on the hand, and irregular wearing of insulating gloves on the hand; divide the insulating glove wearing detection annotation data set into an insulating glove wearing detection training set and an insulating glove wearing detection validation set according to a set ratio.

[0208] Each module in the above-mentioned detection device for insulating glove wearing at the power grid operation site can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in the form of hardware or be independent of it, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0209] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data of embodiments of the detection method for wearing insulating gloves at the power grid operation site. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a detection method for wearing insulating gloves at the power grid operation site.

[0210] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0211] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0212] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0213] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0214] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0215] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0216] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.

[0217] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A detection method for wearing insulating gloves at the on-site operation of the power grid, characterized in that, The method includes: According to the neural network feature extraction module of the insulating glove wearing detection model, channel features of the image of personnel at the power grid operation site are extracted to obtain a channel feature map. According to the channel spatial feature enhancement attention module of the insulating glove wearing detection model, the hand features of the personnel and the features of the insulating gloves in the channel feature map are extracted to obtain an insulating glove feature map; the channel spatial feature enhancement attention module is constructed based on global pooling operation combined with a parameter-free attention mechanism. According to the multi-dilation aggregation module of the insulating glove wearing detection model, local and global features of the insulating glove feature map are extracted and then fused to obtain an enhanced insulating glove feature map. According to the enhanced insulating glove feature map, it is determined whether the power grid construction personnel in the image of the personnel at the power grid operation site wear insulating gloves, whether the insulating gloves are worn correctly, and the target position in the image of the personnel at the power grid operation site, so as to obtain the insulating glove wearing detection result of the image of the personnel at the power grid operation site.

2. The method according to claim 1, characterized in that, The step of extracting the hand features of the personnel and the features of the insulating gloves in the channel feature map according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model to obtain an insulating glove feature map includes: According to the channel spatial feature enhancement attention module of the insulating glove wearing detection model, a convolution operation is performed on the channel feature map to extract the hand features of the personnel and the features of the insulating gloves in the channel feature map to obtain a first intermediate feature map. According to the channel spatial feature enhancement attention module of the insulating glove wearing detection model, global average pooling operation and global max pooling operation are respectively performed on the channel feature map to extract the hand features of the personnel and the features of the insulating gloves in the channel feature map to obtain a second intermediate feature map and a third intermediate feature map, so as to obtain channel feature weights. According to the channel feature map and the channel feature weights, a fifth intermediate feature map is obtained. According to the first intermediate feature map and the channel feature weights, a sixth intermediate feature map is obtained. The channel feature map is input into a parameter-free attention mechanism to obtain a seventh intermediate feature map. The fifth intermediate feature map, the sixth intermediate feature map, and the seventh intermediate feature map are added element-wise to obtain an eighth intermediate feature map, and a convolution operation is performed on the eighth intermediate feature map to obtain an insulating glove feature map.

3. The method according to claim 2, characterized in that, Obtaining the channel feature weights includes: The second intermediate feature map and the third intermediate feature map are added element-wise to obtain a fourth intermediate feature map. According to the fourth intermediate feature map and an activation function, channel feature weights are obtained.

4. The method according to claim 1, characterized in that, The step of performing local and global feature extraction on the insulating glove feature map according to the multi-dilation aggregation module of the insulating glove wearing detection model and then performing a fusion process to obtain an enhanced insulating glove feature map includes: According to the multi-dilation aggregation module of the insulating glove wearing detection model, a max pooling operation is performed on the insulating glove feature map to obtain a first intermediate enhanced feature map. According to the multi-dilation pooling module of the insulating glove wearing detection model, perform continuous multiple dilation convolution operations on the insulating glove feature map to obtain a second intermediate feature enhancement map, a third intermediate feature enhancement map, and a fourth intermediate feature enhancement map; Concatenate the first intermediate feature enhancement map, the second intermediate feature enhancement map, the third intermediate feature enhancement map, and the fourth intermediate feature enhancement map along the channel dimension to obtain a fifth intermediate feature enhancement map; Perform a convolution operation on the fifth intermediate feature enhancement map to obtain a sixth intermediate feature enhancement map; According to the non-linear activation function, perform non-linear enhancement on the sixth intermediate feature enhancement map to obtain an insulating glove feature enhancement map.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: Obtain an insulating glove wearing detection training set and an insulating glove wearing detection validation set; Input the insulating glove wearing detection training images in the insulating glove wearing detection training set into the model with parameters to be adjusted to obtain a first prediction result; According to the first prediction result, the annotation data corresponding to the insulating glove wearing detection training image, and the adaptive intersection over union loss function, obtain a first training loss value; the adaptive intersection over union loss function is obtained according to the intersection over union loss, the center point distance loss, and the length and width loss; Input the insulating glove wearing detection validation images in the insulating glove wearing detection validation set into the model with parameters to be adjusted to obtain a second prediction result; According to the second prediction result, the annotation data corresponding to the insulating glove wearing detection validation image, and the adaptive intersection over union loss function, obtain a second training loss value; According to the first training loss value and the second training loss value, adjust the parameters of the model with parameters to be adjusted until the second training loss value converges, and obtain an insulating glove wearing detection model.

6. The method according to claim 5, characterized in that, The obtaining of the insulating glove wearing detection training set and the insulating glove wearing detection validation set includes: Collect real images of personnel wearing insulating gloves at the power grid operation site and during operation and maintenance; Obtain a first simulated image of wearing an insulating glove irregularly and a second simulated image of not wearing an insulating glove; According to the real images, the first simulated image, and the second simulated image, obtain an insulating glove wearing detection data set; Perform wearing category annotation on the insulating glove wearing detection data set to obtain an insulating glove wearing detection annotation data set; the wearing categories include correctly wearing an insulating glove on the hand, not wearing an insulating glove on the hand, and wearing an insulating glove irregularly; According to a set ratio, divide the insulating glove wearing detection annotation data set into an insulating glove wearing detection training set and an insulating glove wearing detection validation set.

7. A detection device for wearing insulating gloves at the on-site operation of a power grid, characterized in that, The device includes: A channel feature map acquisition module, configured to perform channel feature extraction on the power grid operation site personnel image according to the neural network feature extraction module of the insulating glove wearing detection model to obtain a channel feature map; The insulating glove feature map acquisition module is used to extract the hand features of the person and the features of the insulating glove in the channel feature map according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model, so as to obtain the insulating glove feature map; the channel spatial feature enhancement attention module is constructed based on the global pooling operation combined with the parameter-free attention mechanism; The feature enhancement map acquisition module is used to perform local and global feature extraction on the insulating glove feature map according to the multi-dilation convergence module of the insulating glove wearing detection model, and then perform fusion processing to obtain the insulating glove feature enhancement map; The detection result acquisition module is used to determine whether the grid construction personnel in the grid operation site personnel image wear insulating gloves, whether the insulating gloves are worn correctly, and the target position in the grid operation site personnel image according to the insulating glove feature enhancement map, so as to obtain the insulating glove wearing detection result of the grid operation site personnel image.

8. A computer device, including a memory and a processor, the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, including a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Safety helmet wearing detection method and device based on deep learning, equipment and medium

    CN114782986A

  • Edge detection system for power grid inspection and monitoring

    CN116846059A

  • Electric power safety operation target detection method and system based on deep learning

    CN118781511A

  • System and method for automatic detection and recognition of people wearing personal protective equipment using deep learning

    US20220058381A1