Method, device, computer equipment, readable storage medium and program product for detecting the wearing of insulating gloves at power grid operation sites

Through the neural network feature extraction and feature enhancement modules of the insulating gloves wearing detection model, the problem of low accuracy in insulating gloves wearing detection at power grid operation sites is solved, and higher detection accuracy is achieved.

CN120183006BActive Publication Date: 2025-09-30ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510671877.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-30
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Existing target detection algorithms have difficulty accurately locating the specific areas where insulating gloves are worn at power grid operation sites, especially in complex backgrounds and occlusions, resulting in low accuracy in insulating glove wearing detection.

Method used

An insulating glove wearing detection model is adopted. Through the neural network feature extraction module, the channel space feature enhanced attention module and the multi-expansion aggregation module, combined with the global pooling operation and the parameter-free attention mechanism, the hand and insulating glove features are extracted and fused to improve the detection accuracy.

Benefits of technology

The global context information of the insulating gloves feature map is enhanced, the feature extraction capability of blurred targets is improved, the occlusion and angle changes of the targets are effectively dealt with, and the accuracy of insulating gloves wearing detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183006B_ABST
    Figure CN120183006B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of safe power operations and provides a method, apparatus, computer equipment, readable storage medium, and program product for detecting the wearing of insulating gloves at power grid work sites. The method comprises: extracting channel features from images of personnel at power grid work sites using a neural network feature extraction module of an insulating glove wearing detection model to obtain a channel feature map; extracting the personnel's hand features and insulating glove features from the channel feature map using a channel spatial feature enhancement attention module of the insulating glove wearing detection model to obtain an insulating glove feature map; the channel spatial feature enhancement attention module is constructed using a global pooling operation combined with a parameter-free attention mechanism; and extracting local and global features from the insulating glove feature map using a multi-expansion convergence module of the insulating glove wearing detection model, performing fusion processing, and obtaining an insulating glove feature enhancement map to obtain an insulating glove wearing detection result. This method can improve detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of safe power operations, and in particular to a method, device, computer equipment, computer-readable storage medium, and computer program product for detecting the wearing of insulating gloves at power grid operation sites. Background Art

[0002] In the power industry, safety management of power grid operation sites is of vital importance. Power grid construction sites involve the operation or maintenance of high-voltage electrical equipment. Especially when performing live operations, power grid construction workers should wear insulating gloves that meet the requirements of the corresponding voltage level to ensure personal safety.

[0003] Currently, deep learning-based target detection algorithms can be used to detect the wearing of insulating gloves at power grid construction sites. However, due to the complex and ever-changing backgrounds at power grid construction sites, existing target detection algorithms struggle to precisely locate the specific areas where gloves are worn, resulting in low detection accuracy. When power grid workers are far from the detection equipment, glove-wearing features may appear as small pixel areas in the image. Existing target detection algorithms have difficulty capturing these subtle features and can easily overlook key information, resulting in low detection accuracy for insulating gloves. In actual work environments, power grid workers' hands may be partially obscured or shift in angle due to operating postures and obstructions. Existing target detection algorithms struggle to effectively address these occlusions and angle changes, resulting in low detection accuracy for insulating gloves. Summary of the Invention

[0004] Based on this, it is necessary to provide a detection method, device, computer equipment, computer-readable storage medium and computer program product for the wearing of insulating gloves at power grid operation sites to address the above technical problems.

[0005] In a first aspect, the present application provides a method for detecting the wearing of insulating gloves at a power grid operation site, comprising:

[0006] Based on the neural network feature extraction module of the insulating glove wearing detection model, channel feature extraction is performed on the images of personnel working on the power grid to obtain a channel feature map.

[0007] According to the channel space feature enhanced attention module of the insulating glove wearing detection model, the hand features of the person and the insulating glove features in the channel feature map are extracted to obtain the insulating glove feature map; the channel space feature enhanced attention module is constructed based on the global pooling operation combined with the parameter-free attention mechanism;

[0008] According to the multi-expansion convergence module of the insulating glove wearing detection model, local and global features are extracted from the insulating glove feature map and then fused to obtain an insulating glove feature enhancement map;

[0009] Based on the insulating gloves feature enhancement map, it is determined whether the power grid construction personnel in the power grid operation site personnel image are wearing insulating gloves, whether the insulating gloves are worn correctly, and the target position of the power grid operation site personnel image is determined, and an insulating gloves wearing detection result of the power grid operation site personnel image is obtained.

[0010] In one embodiment, the channel spatial feature enhancement attention module according to the insulating glove wearing detection model extracts the hand features of the person and the features of the insulating gloves in the channel feature map to obtain the insulating glove feature map, including:

[0011] According to the channel spatial feature enhancement attention module of the insulating glove wearing detection model, a convolution operation is performed on the channel feature map to extract the hand features of the person and the features of the insulating gloves in the channel feature map to obtain a first intermediate feature map;

[0012] According to the channel spatial feature enhancement attention module of the insulating glove wearing detection model, a global average pooling operation and a global maximum pooling operation are performed on the channel feature map, and the hand features of the person and the insulating glove features in the channel feature map are extracted to obtain a second intermediate feature map and a third intermediate feature map to obtain a channel feature weight;

[0013] Obtaining a fifth intermediate feature map according to the channel feature map and the channel feature weight;

[0014] Obtaining a sixth intermediate feature map according to the first intermediate feature map and the channel feature weights;

[0015] Inputting the channel feature map into the parameter-free attention mechanism to obtain a seventh intermediate feature map;

[0016] The fifth intermediate feature map, the sixth intermediate feature map, and the seventh intermediate feature map are element-wise added to obtain an eighth intermediate feature map, and a convolution operation is performed on the eighth intermediate feature map to obtain an insulating glove feature map.

[0017] In one embodiment, obtaining the channel feature weights includes:

[0018] Adding the second intermediate feature map and the third intermediate feature map element-wise to obtain a fourth intermediate feature map;

[0019] A channel feature weight is obtained according to the fourth intermediate feature map and the activation function.

[0020] In one embodiment, the multi-expansion convergence module according to the insulating glove wearing detection model extracts local and global features from the insulating glove feature map and then fuses them to obtain an insulating glove feature enhancement map, including:

[0021] performing a maximum pooling operation on the insulating glove feature map according to the multi-expansion convergence module of the insulating glove wearing detection model to obtain a first intermediate feature enhancement map;

[0022] According to the multi-expansion and convergence module of the insulating glove wearing detection model, the insulating glove feature map is subjected to multiple consecutive expansion and convolution operations to obtain a second intermediate feature enhancement map, a third intermediate feature enhancement map, and a fourth intermediate feature enhancement map;

[0023] splicing the first intermediate feature enhancement map, the second intermediate feature enhancement map, the third intermediate feature enhancement map, and the fourth intermediate feature enhancement map according to the channel dimension to obtain a fifth intermediate feature enhancement map;

[0024] performing a convolution operation on the fifth intermediate feature enhancement map to obtain a sixth intermediate feature enhancement map;

[0025] According to the nonlinear activation function, the sixth intermediate feature enhancement map is nonlinearly enhanced to obtain an insulating glove feature enhancement map.

[0026] In one embodiment, the method further comprises:

[0027] Obtain the insulating gloves wearing detection training set and the insulating gloves wearing detection verification set;

[0028] Inputting the insulating glove wearing detection training images in the insulating glove wearing detection training set into the model to be adjusted in parameters to obtain a first prediction result;

[0029] Obtaining a first training loss value based on the first prediction result, the labeled data corresponding to the insulating glove wearing detection training image, and an adaptive intersection-over-union loss function; the adaptive intersection-over-union loss function is obtained based on intersection-over-union loss, center point distance loss, and length-width loss;

[0030] Inputting the insulating glove wearing detection verification images in the insulating glove wearing detection verification set into the parameter-adjusted model to obtain a second prediction result;

[0031] Obtaining a second training loss value according to the second prediction result, the labeled data corresponding to the insulating glove wearing detection verification image, and the adaptive intersection-over-union loss function;

[0032] According to the first training loss value and the second training loss value, the parameters of the model to be adjusted are adjusted until the second training loss value tends to converge, thereby obtaining an insulating glove wearing detection model.

[0033] In one embodiment, obtaining an insulating glove wearing detection training set and an insulating glove wearing detection verification set includes:

[0034] Collect real images of people wearing insulating gloves at power grid operation sites and during operation and maintenance;

[0035] Acquire a first simulated image of an insulated glove not being worn properly and a second simulated image of an insulated glove not being worn;

[0036] Obtaining an insulating glove wearing detection dataset based on the real image, the first simulated image, and the second simulated image;

[0037] Performing wearing category labeling on the insulating gloves wearing detection dataset to obtain an insulating gloves wearing detection labeling dataset; the wearing categories include hands correctly wearing insulating gloves, hands not wearing insulating gloves, and hands wearing insulating gloves improperly;

[0038] According to a set ratio, the insulating gloves wearing detection annotated dataset is divided into an insulating gloves wearing detection training set and an insulating gloves wearing detection verification set.

[0039] In a second aspect, the present application also provides a detection device for wearing insulating gloves at a power grid operation site, comprising:

[0040] A channel feature map acquisition module is used to extract channel features from images of personnel working on the power grid operation site based on the neural network feature extraction module of the insulating glove wearing detection model to obtain a channel feature map;

[0041] An insulating glove feature map acquisition module is configured to enhance the attention module based on the channel spatial features of the insulating glove wearing detection model, extract the hand features of the person and the insulating glove features in the channel feature map, and obtain an insulating glove feature map; the channel spatial feature enhanced attention module is constructed based on a global pooling operation combined with a parameter-free attention mechanism;

[0042] a feature enhancement map acquisition module, configured to extract local and global features from the insulating glove feature map and then fuse them according to the multi-expansion convergence module of the insulating glove wearing detection model to obtain an insulating glove feature enhancement map;

[0043] The detection result acquisition module is used to determine whether the power grid construction personnel in the power grid operation site personnel image are wearing insulating gloves, whether the insulating gloves are worn correctly, and the target position of the power grid operation site personnel image based on the insulating gloves feature enhancement map, and obtain the insulating gloves wearing detection result of the power grid operation site personnel image.

[0044] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program and the processor executes the above method.

[0045] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is used by a processor to execute the above method.

[0046] In a fifth aspect, the present application further provides a computer program product, wherein the computer program product includes a computer program, and the computer program is executed by a processor to execute the above method.

[0047] The above-mentioned detection method, device, computer equipment, computer-readable storage medium and computer program product for the wearing of insulating gloves at the power grid operation site are as follows: according to the neural network feature extraction module of the insulating glove wearing detection model, channel features are extracted from the image of the power grid operation site personnel to obtain a channel feature map; according to the channel space feature enhancement attention module of the insulating glove wearing detection model, the hand features of the personnel and the insulating glove features in the channel feature map are extracted to obtain an insulating glove feature map; the channel space feature enhancement attention module is constructed based on the global pooling operation combined with the parameter-free attention mechanism; according to the multi-expansion convergence module of the insulating glove wearing detection model, local and global features are extracted from the insulating glove feature map and then fused to obtain an insulating glove feature enhancement map; according to the insulating glove feature enhancement map, whether the power grid construction personnel in the power grid operation site personnel image are wearing insulating gloves, whether the insulating gloves are worn correctly and the target position of the power grid operation site personnel image are determined to obtain the insulating glove wearing detection result of the power grid operation site personnel image. This application constructs a channel space feature enhancement attention module of the insulating glove wearing detection model based on the global pooling operation and the parameter-free attention mechanism; according to the channel space feature enhancement attention module, the hand features of the person and the features of the insulating gloves in the channel feature map are extracted to obtain the insulating glove feature map; the global context information of the insulating glove feature map can be enhanced, and the feature extraction capability of the blurred target can be improved, thereby improving the detection accuracy of the insulating glove wearing; in the multi-expansion convergence module of the insulating glove wearing detection model, the local and global relationships in the space of the insulating glove feature map can be better captured, and the occlusion and angle changes of the target can be effectively dealt with, thereby further improving the detection accuracy of the insulating glove wearing. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0049] Figure 1 A diagram illustrating an application environment of a method for detecting the wearing of insulating gloves at a power grid operation site according to an embodiment;

[0050] Figure 2 1 is a flow chart of a method for detecting the wearing of insulating gloves at a power grid operation site according to one embodiment;

[0051] Figure 3 Schematic diagram of the structure of a channel spatial feature enhanced attention module in one embodiment;

[0052] Figure 4 A schematic diagram of the structure of a multi-expansion and convergence module in one embodiment;

[0053] Figure 5 Schematic diagram of a flow chart of a method for detecting the wearing of insulating gloves at a power grid operation site according to another embodiment;

[0054] Figure 6 Schematic diagram of a flow chart of a fast and effective detection algorithm for insulating gloves in one embodiment;

[0055] Figure 7 This is a structural block diagram of a device for detecting the wearing of insulating gloves at a power grid operation site according to one embodiment;

[0056] Figure 8 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0058] The detection method for wearing insulating gloves at a power grid operation site provided by the embodiment of the present application can be applied to Figure 1 In the application environment shown. The terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can input the image of the personnel working on the power grid into the insulating glove wearing detection model to obtain the insulating glove wearing detection result of the image of the personnel working on the power grid. The terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0059] In an exemplary embodiment, Figure 2 As shown, a detection method for wearing insulating gloves at power grid operation sites is provided. Figure 1 The terminal 102 in the example is used as an example to illustrate the process, including the following steps S201 to S204.

[0060] Step S201: Based on the neural network feature extraction module of the insulating glove wearing detection model, channel feature extraction is performed on the image of the personnel at the power grid operation site to obtain a channel feature map.

[0061] The insulating glove wearing detection model can detect whether the power grid construction workers in the power grid operation site personnel image are wearing insulating gloves. The insulating glove wearing detection model includes a neural network feature extraction module, a channel space feature enhancement attention module and a multi-expansion convergence module.

[0062] Images of grid construction personnel performing live work at the grid operation site and during operation and maintenance can be collected as images of grid operation site personnel.

[0063] The neural network feature extraction module can be PConv_BN_ELU, which consists of three parts: convolution layer (Convolution, Conv), batch normalization layer (Batch Normalization, BN) and activation function layer (Exponential Linear Unit, ELU).

[0064] The images of personnel working on the power grid operation site can be input into the insulating glove wearing detection model, so that the neural network feature extraction module of the insulating glove wearing detection model can extract channel features from the images of personnel working on the power grid operation site to obtain a channel feature map.

[0065] Step S202: According to the channel space feature enhancement attention module of the insulating glove wearing detection model, the hand features of the person and the insulating glove features in the channel feature map are extracted to obtain the insulating glove feature map; the channel space feature enhancement attention module is constructed based on the global pooling operation combined with the parameter-free attention mechanism.

[0066] We can construct the following by combining the global pooling operation with the parameter-free attention mechanism (Similarity-based AttentionModule, SimAM): Figure 3The Channel-Space Features Enhanced Attention Module (CSFE_AM) shown in the figure is a CSFE-based model. Global Average Pooling represents the global average pooling operation, Global Max Pooling represents the global maximum pooling operation, Conv represents the convolution operation, SimAM represents the parameter-free attention mechanism, and Sigmoid represents the S-type activation function. This module effectively reduces the number of parameters in the insulating glove wearing detection model while identifying and strengthening feature channels that are important for insulating glove position recognition and suppressing redundant or noisy channels. At the spatial feature level, the module fuses the insulating glove position details in low-order feature maps with the rich contextual semantic information in high-order feature maps, enabling the insulating glove wearing detection model to more accurately locate and identify whether power grid construction workers are wearing insulating gloves correctly.

[0067] The parameter-free attention mechanism is a lightweight, parameter-free attention mechanism based on the theory of spatial inhibition in neuroscience. It infers the importance of each neuron in the feature map by optimizing an energy function, thereby generating three-dimensional attention weights. Unlike traditional attention mechanisms, the parameter-free attention mechanism effectively improves the representational power and performance of convolutional neural networks (CNNs) without adding any parameters to the original network. The parameter-free attention mechanism's implementation is simple and efficient, avoiding complex structural adjustments and heuristic methods, making its application in various visual tasks more flexible and convenient.

[0068] The attention module can be enhanced according to the channel spatial characteristics of the insulating gloves wearing detection model to identify and strengthen the feature channels that are important for insulating gloves position recognition, suppress redundant or noisy channels, and accurately extract the hand features of the person and the insulating gloves in the channel feature map to obtain the insulating gloves feature map.

[0069] In step S203, according to the multi-expansion convergence module of the insulating glove wearing detection model, local and global features are extracted from the insulating glove feature map and then fused to obtain an insulating glove feature enhancement map.

[0070] The Multi-dilatation Aggregation Module (MDAM) of the insulating gloves wearing detection model is as follows: Figure 4As shown in the figure, MaxPool represents the maximum pooling operation, Conv represents the convolution operation, and SiLU represents the Sigmoid linear unit activation function. The multi-expansion convergence module uses multiple consecutive dilated convolution operations to extract local and global features from the insulating glove feature map. By fusing the multi-scale feature maps, the local and global features are integrated to obtain the insulating glove feature enhancement map, which can reduce the impact of reduced detection accuracy caused by occlusion and angle changes.

[0071] Among them, dilated convolution is a special convolution method that can obtain a larger receptive field without increasing the amount of computation and reducing the resolution of the feature map.

[0072] Step S204: Determine whether the power grid construction worker in the power grid operation site worker image is wearing insulating gloves, whether the insulating gloves are worn correctly, and the target position of the power grid operation site worker image based on the insulating gloves feature enhancement map, and obtain the insulating gloves wearing detection result of the power grid operation site worker image.

[0073] The insulating gloves wearing detection results can be displayed on the image of the personnel at the power grid operation site, including whether the insulating gloves are worn correctly, the target position in the image of the personnel at the power grid operation site, and the detection confidence.

[0074] Based on the insulating gloves feature enhancement map, it can be determined whether the power grid construction workers in the power grid operation site personnel image are wearing insulating gloves, whether the insulating gloves are worn correctly, and the target position of the power grid operation site personnel image, thereby obtaining the insulating gloves wearing detection result.

[0075] In the above-mentioned detection method for insulating gloves wearing at power grid operation sites, a channel space feature enhancement attention module of the insulating gloves wearing detection model is constructed based on the global pooling operation and combined with the parameter-free attention mechanism; according to the channel space feature enhancement attention module, the hand features of the person and the features of the insulating gloves in the channel feature map are extracted to obtain the insulating gloves feature map; the global context information of the insulating gloves feature map can be enhanced, and the feature extraction capability of the blurred target can be improved, thereby improving the detection accuracy of the insulating gloves wearing; in the multi-expansion convergence module of the insulating gloves wearing detection model, the local and global relationships in the space of the insulating gloves feature map can be better captured, and the occlusion and angle changes of the target can be effectively dealt with, thereby further improving the detection accuracy of the insulating gloves wearing.

[0076] In one embodiment, according to the channel space feature enhanced attention module of the insulating glove wearing detection model, the hand features of the person and the features of the insulating gloves in the channel feature map are extracted to obtain the insulating glove feature map, and the specific steps are as follows: according to the channel space feature enhanced attention module of the insulating glove wearing detection model, a convolution operation is performed on the channel feature map to extract the hand features of the person and the features of the insulating gloves in the channel feature map to obtain a first intermediate feature map; according to the channel space feature enhanced attention module of the insulating glove wearing detection model, a global average pooling operation and a global maximum pooling operation are performed on the first intermediate feature map respectively to extract the hand features of the person and the features of the insulating gloves in the channel feature map to obtain a second intermediate feature map and a third intermediate feature map to obtain a channel feature weight; according to the channel feature map and the channel feature weight, a fifth intermediate feature map is obtained; according to the first intermediate feature map and the channel feature weight, a sixth intermediate feature map is obtained; the channel feature map is input into the parameter-free attention mechanism to obtain a seventh intermediate feature map; the fifth intermediate feature map, the sixth intermediate feature map and the seventh intermediate feature map are element-wise added to obtain an eighth intermediate feature map, and a convolution operation is performed on the eighth intermediate feature map to obtain an insulating glove feature map.

[0077] The attention module can be enhanced based on the channel spatial features of the insulating glove wearing detection model. A convolution operation with a kernel size k of 3, a step size s of 1, and a padding value p of 1 is performed on the channel feature map F0. The hand features of the person and the insulating glove features in the channel feature map are extracted to obtain the first intermediate feature map F1 with a constant spatial resolution and a number of channels C. The size of the channel feature map F0 is C×H×W, where C represents the number of channels, H represents the height, and W represents the width.

[0078] The attention module can be enhanced according to the channel spatial features of the insulating gloves wearing detection model, and a global average pooling operation can be performed on the first intermediate feature map F1 to extract the hand features of the person and the features of the insulating gloves in the channel feature map, thereby obtaining a second intermediate feature map F2 with a spatial resolution of 1×1 and a number of channels of C; the attention module can be enhanced according to the channel spatial features of the insulating gloves wearing detection model to perform a global maximum pooling operation on the first intermediate feature map F1 to extract the hand features of the person and the features of the insulating gloves in the channel feature map, thereby obtaining a third intermediate feature map F3 with a spatial resolution of 1×1 and a number of channels of C; the channel feature weight W1 can be obtained based on the second intermediate feature map F2 and the third intermediate feature map F3.

[0079] The channel feature map F0 and the channel feature weight W1 can be dot-multiplied to obtain the fifth intermediate feature map F5 with adaptive channel feature enhancement; the first intermediate feature map F1 and the channel feature weight W1 can be element-wise multiplied, and the first intermediate feature map F1 and the channel feature weight W1 can be dot-multiplied to obtain the sixth intermediate feature map F6 with adaptive channel feature enhancement, where the spatial resolution of the sixth intermediate feature map F6 is H×W and the number of channels is C; the channel feature map F0 can be input into the parameter-free attention mechanism to obtain the seventh intermediate feature map F7 with adaptive spatial feature enhancement, where the spatial resolution of the seventh intermediate feature map F7 is H×W and the number of channels is C.

[0080] The fifth intermediate feature map F5, the sixth intermediate feature map F6, and the seventh intermediate feature map F7 can be element-wise added to obtain an eighth intermediate feature map F8 having a spatial resolution of H×W and a number of channels C. A convolution operation is then performed on the eighth intermediate feature map F8 with a convolution kernel size k of 3, a stride s of 1, and a padding value p of 1 to obtain an insulating glove feature map F9 having a spatial resolution of H×W and a number of channels C. By way of example, the spatial resolution can be 32×32 and the number of channels can be 256.

[0081] In this embodiment, an attention module is enhanced based on the channel spatial features of the insulating glove wearing detection model to extract the hand features of the person and the features of the insulating gloves in the channel feature map to obtain an insulating glove feature map. Specifically, channel features are extracted through global average pooling and global maximum pooling, and the spatial features are enhanced by combining a parameter-free attention mechanism. This can capture more distant small-sized people in the power grid operation site image, so that the insulating glove wearing detection model can more accurately locate and identify whether the power grid construction personnel are wearing insulating gloves correctly.

[0082] In one embodiment, the channel feature weight is obtained by the following specific steps: adding the second intermediate feature map and the third intermediate feature map element by element to obtain a fourth intermediate feature map; and obtaining the channel feature weight according to the fourth intermediate feature map and the activation function.

[0083] The second intermediate feature map F2 and the third intermediate feature map F3 can be element-wise added to obtain a fourth intermediate feature map F4; the fourth intermediate feature map F9 can be subjected to a layer of sigmoid activation function to obtain a channel feature weight W1.

[0084] In this embodiment, the second intermediate feature map and the third intermediate feature map are element-wise added to obtain a fourth intermediate feature map to obtain channel feature weights, so that the channel spatial feature enhanced attention module can strengthen feature channels important for insulating glove position identification according to the channel feature weights and suppress redundant or noisy channels.

[0085] In one embodiment, according to the multi-expansion convergence module of the insulating glove wearing detection model, local and global features are extracted from the insulating glove feature map and then fused to obtain an insulating glove feature enhancement map. The specific steps are as follows: according to the multi-expansion convergence module of the insulating glove wearing detection model, a maximum pooling operation is performed on the insulating glove feature map to obtain a first intermediate feature enhancement map; according to the multi-expansion convergence module of the insulating glove wearing detection model, a plurality of continuous expansion convolution operations are performed on the insulating glove feature map to obtain a second intermediate feature enhancement map, a third intermediate feature enhancement map and a fourth intermediate feature enhancement map; the first intermediate feature enhancement map, the second intermediate feature enhancement map, the third intermediate feature enhancement map and the fourth intermediate feature enhancement map are spliced ​​according to the channel dimension to obtain a fifth intermediate feature enhancement map; a convolution operation is performed on the fifth intermediate feature enhancement map to obtain a sixth intermediate feature enhancement map; and according to a nonlinear activation function, the sixth intermediate feature enhancement map is nonlinearly enhanced to obtain an insulating glove feature enhancement map.

[0086] Based on the multi-expansion convergence module of the insulating glove wearing detection model, a maximum pooling operation can be performed on the insulating glove feature map C0 without changing the spatial resolution and number of channels. Local features of the insulating glove feature map C0 can be extracted to obtain a first intermediate feature enhancement map C1 with a spatial resolution of H×W and a number of channels of C. Here, the insulating glove feature map F9 output by the channel spatial feature enhancement attention module is denoted as C0. The size of the insulating glove feature map C0 is C×H×W, where C represents the number of channels, H represents the height, and W represents the width. For example, the spatial resolution can be 128×128 and the number of channels can be 64.

[0087] According to the multi-expansion convergence module of the insulating glove wearing detection model, the insulating glove feature map C0 can be subjected to multiple continuous expansion convolution operations to extract local features and global features of the insulating glove feature map C0. Specifically, the insulating glove feature map C0 is subjected to an expansion convolution operation with a convolution kernel size k of 3, a step size s of 1, a padding value p of 1, and an expansion value of 1 to obtain a second intermediate feature enhancement map C2 with an unchanged spatial resolution and a number of channels of C; the second intermediate feature enhancement map C2 can be subjected to an expansion convolution operation with a convolution kernel size k of 3, a step size s of 1, a padding value p of 3, and an expansion value of 3 to obtain a third intermediate feature enhancement map C3 with an unchanged spatial resolution and a number of channels of C; the third intermediate feature enhancement map C3 can be subjected to an expansion convolution operation with a convolution kernel size k of 3, a step size s of 1, a padding value p of 1, and an expansion value of 1 to obtain a fourth intermediate feature enhancement map C4 with an unchanged spatial resolution and a number of channels of C.

[0088] The first intermediate feature enhancement map C1, the second intermediate feature enhancement map C2, the third intermediate feature enhancement map C3 and the fourth intermediate feature enhancement map C4 can be spliced ​​according to the channel dimension to obtain a fifth intermediate feature enhancement map C5 with a spatial resolution of H×W and a channel number of 4C, so as to realize the fusion processing of the local features and global features of the insulating gloves feature map C0; the fifth intermediate feature enhancement map C5 can be subjected to a convolution operation with a convolution kernel size k of 1, a step size s of 1, and a padding value p of 0 to obtain a sixth intermediate feature enhancement map C6 with an unchanged spatial resolution and a channel number of C; the sixth intermediate feature enhancement map C6 is nonlinearly enhanced according to the Sigmoid Linear Unit (SiLU) activation function to obtain the insulating gloves feature enhancement map C7.

[0089] In this embodiment, according to the multi-expansion convergence module of the insulating glove wearing detection model, a maximum pooling operation, multiple consecutive expansion convolution operations and convolution operations are performed on the insulating glove feature map, and the local features and global features of the insulating glove feature map are extracted and fused to obtain an insulating glove feature enhancement map. This can filter out interference information, enhance small target feature information, and improve the problem of target feature information being submerged during the convolution process, thereby solving the problem of target occlusion in the insulating glove wearing detection task.

[0090] In one embodiment, the method provided by the present application also includes: obtaining an insulating glove wearing detection training set and an insulating glove wearing detection verification set; inputting the insulating glove wearing detection training image in the insulating glove wearing detection training set into the parameter to be adjusted model to obtain a first prediction result; obtaining a first training loss value based on the first prediction result, the labeled data corresponding to the insulating glove wearing detection training image, and the adaptive intersection-over-union loss function; the adaptive intersection-over-union loss function is obtained based on the intersection-over-union loss, the center point distance loss, and the length-width loss; inputting the insulating glove wearing detection verification image in the insulating glove wearing detection verification set into the parameter to be adjusted model to obtain a second prediction result; obtaining a second training loss value based on the second prediction result, the labeled data corresponding to the insulating glove wearing detection verification image, and the adaptive intersection-over-union loss function; adjusting the parameters of the parameter to be adjusted model based on the first training loss value and the second training loss value until the second training loss value converges, thereby obtaining an insulating glove wearing detection model.

[0091] You can obtain the insulating gloves wearing detection training set and the insulating gloves wearing detection verification set.

[0092] By combining the global pooling operation with the parameter-free attention mechanism, a channel-space feature enhancement attention module for the model whose parameters are to be adjusted can be constructed. By combining the dilated convolution operation with the multi-scale feature map fusion operation, a multi-dilated convergence module for the model whose parameters are to be adjusted can be constructed.

[0093] Before model training, you can initialize the parameters of the model to be adjusted and set the hyperparameters related to the model to be adjusted, such as the number of training rounds, batch size, optimizer selection, learning rate, and maximum value of the gradient clipping strategy.

[0094] The insulating glove wearing detection training images in the insulating glove wearing detection training set can be input into the parameter-adjusted model to obtain a first prediction result. The first training loss value can be obtained by measuring the difference between the first prediction result and the labeled data corresponding to the insulating glove wearing detection training images using an adaptive intersection over union (AD-IOU) loss function.

[0095] The adaptive intersection over union loss function adds a loss metric based on the intersection over union (IOU) loss function. and , to solve the problem of inaccurate bounding box regression. The adaptive intersection-over-union loss function is 、 and The weighted sum of is shown in formula (1).

[0096] (1)

[0097] (2)

[0098] (3)

[0099] (4)

[0100] in It is the loss based on intersection-over-union. is the loss based on the center point distance, It is based on the loss of length and width.

[0101] In the intersection-over-union loss, It represents 1 minus the intersection-over-union ratio between the predicted box and the true box, that is, is the ratio of the overlap of two boxes to their union. The smaller the value of The larger it is, the higher the overlap between the predicted box and the true box, which means the higher the similarity between the predicted box and the true box.

[0102] exist in ,in is a hyperparameter, and are the center coordinates of the predicted box and the real box respectively, Represents the Euclidean distance between the center coordinates of the predicted box and the center coordinates of the real box, and are the squares of the width and height of the ground-truth box, respectively. The loss function is used to reduce the distance between the center coordinates of the predicted box and the center coordinates of the real box, making the predicted box closer to the real box.

[0103] exist middle, and are the width and height of the prediction box respectively. Indicates taking the absolute value. The loss function is used to reduce the difference between the length and width of the predicted box and the length and width of the real box.

[0104] The insulating glove wearing detection verification images in the insulating glove wearing detection verification set can be input into the parameter-adjusted model to obtain a second prediction result. The second training loss value can be obtained by measuring the gap between the second prediction result and the labeled data corresponding to the insulating glove wearing detection training images based on the adaptive intersection over union (AD-IOU) loss function.

[0105] The parameters of the model to be adjusted can be adjusted according to the first training loss value and the second training loss value until the second training loss value converges, thereby obtaining an insulating glove wearing detection model.

[0106] In this embodiment, the insulating glove wearing detection training images in the insulating glove wearing detection training set are input into the parameter-to-be-adjusted model to obtain a first prediction result, thereby obtaining a first training loss value; the insulating glove wearing detection verification images in the insulating glove wearing detection verification set are input into the parameter-to-be-adjusted model to obtain a second prediction result, thereby obtaining a second training loss value; and the parameters of the parameter-to-be-adjusted model are adjusted according to the first training loss value and the second training loss value until the second training loss value converges, thereby obtaining an insulating glove wearing detection model with better performance.

[0107] In one embodiment, an insulating glove wearing detection training set and an insulating glove wearing detection verification set are obtained, and the specific steps are as follows: collecting real images of insulating gloves worn at power grid operation sites and during operation and maintenance; obtaining a first simulated image of insulating gloves worn improperly and a second simulated image of not wearing insulating gloves; obtaining an insulating glove wearing detection dataset based on the real image, the first simulated image, and the second simulated image; annotating the insulating glove wearing detection dataset by wearing categories to obtain an insulating glove wearing detection labeled dataset; the wearing categories include hands wearing insulating gloves correctly, hands not wearing insulating gloves, and hands wearing insulating gloves improperly; and dividing the insulating glove wearing detection labeled dataset into an insulating glove wearing detection training set and an insulating glove wearing detection verification set according to a set ratio.

[0108] At power grid operation sites and during operation and maintenance, high-definition cameras are used to film power grid construction workers, ensuring that multiple power grid construction workers appear in the captured images as much as possible, reflecting various insulating glove wearing detection scenarios to ensure data diversity and representativeness.

[0109] To improve the model's accuracy in identifying the wearing conditions of insulating gloves, negative samples can be introduced. A certain number of negative samples can be created by simulating situations where the insulating gloves are not worn and situations where the gloves are improperly worn, such as gloves slipping off or fingers being exposed. This results in a first simulated image of the person wearing the insulating gloves improperly and a second simulated image of the person not wearing the insulating gloves.

[0110] Based on the real image, the first simulated image, and the second simulated image, an insulating glove wearing detection dataset is obtained. For each image in the insulating glove wearing detection dataset, professional annotation tools can be used to finely annotate the hand area of ​​the power grid construction worker in the image to obtain an insulating glove wearing detection annotation dataset. The annotation information includes the exact location of the hand area and the wearing category of the insulating gloves. A corresponding label file can be generated for each image. The label file records in detail the location information of the hand area and the wearing category of the insulating gloves. The wearing categories include hands wearing insulating gloves correctly, hands not wearing insulating gloves, and hands wearing insulating gloves improperly.

[0111] A series of preprocessing operations can be performed on the images in the insulating gloves wearing detection annotation dataset, including adjusting the image size to adapt to the model input requirements, adding noise to improve the generalization ability of the model, and ensuring that the image quality meets the requirements of model training.

[0112] The set ratio is determined according to actual needs; according to the set ratio, the dataset is divided into an insulating glove wearing detection training set and an insulating glove wearing detection verification set, which are used for model training and performance verification respectively to ensure that the model can be fully evaluated and optimized at different stages.

[0113] In this embodiment, an insulating glove wearing detection dataset is obtained based on the real image, the first simulated image, and the second simulated image, which can improve the diversity of the insulating glove wearing detection dataset; according to the set ratio, the dataset is proportionally divided into an insulating glove wearing detection training set and an insulating glove wearing detection verification set, thereby preparing data for model training and performance verification.

[0114] In order to better understand the above method, an application example of the detection method for wearing insulating gloves at a power grid operation site of the present application is described in detail below.

[0115] Safety management at power grid operation sites is crucial in the power industry. Grid construction sites involve the operation or maintenance of high-voltage electrical equipment. Especially when performing live work, grid construction workers must wear insulating gloves that meet the corresponding voltage level requirements to ensure personal safety. However, some grid construction workers may not wear insulating gloves for various reasons. This necessitates on-site management personnel to conduct patrol inspections, but it is difficult to ensure that every worker complies with safety regulations. The rapid development of deep learning technology, particularly significant advances in object detection technology, has provided new solutions for safety management at power grid operation sites.

[0116] Deep learning-based object detection algorithms automatically analyze image or video data to efficiently and accurately identify specific targets, such as people, tools, and safety equipment. For the task of detecting the wearing of insulating gloves at power grid work sites, this type of object detection algorithm can capture images of construction workers in real time and automatically detect whether they are wearing insulating gloves and whether they are wearing them correctly. Generally speaking, the basic idea behind deep learning-based methods for detecting the wearing of insulating gloves at power grid work sites is to use construction images captured by surveillance cameras at the work site as input. The captured images are then preprocessed to optimize image quality. Next, the feature extraction capabilities of deep convolutional neural networks are leveraged to extract key features related to the wearing of insulating gloves from the preprocessed images. Based on this feature extraction, advanced object detection models are employed. Common models include the R-CNN series, YOLO, and SSD. R-CNN stands for Region-based Convolutional Neural Network, YOLO stands for You Only Look Once, and SSD stands for Single Shot MultiBox Detector. These models can predict in real time whether insulating gloves are present and whether they are worn in compliance with safety standards. Finally, the detection results are presented intuitively, such as by marking the locations of construction workers who are not wearing gloves or are wearing them improperly, and corresponding reports or alerts are generated so that corrective measures can be taken promptly.

[0117] In summary, the deep learning-based insulating glove wearing detection method for power grid operation sites, with its advantages of high efficiency, accuracy and real-time performance, provides strong technical support for safety management of power grid operation sites and effectively reduces the risk of safety accidents.

[0118] The current deep learning-based method for detecting the wearing of insulating gloves at power grid work sites is as follows:

[0119] (1) Based on YOLOv5, a method for detecting unsafe behaviors such as not wearing insulating gloves is proposed. First, the convolution in the feature extraction network is replaced with a self-calibration convolution, so that the network pays more attention to the information around the target to be detected when extracting target features. Then, a selective kernel (SK) attention mechanism is added to the end of the feature fusion network to make the network pay more attention to small targets to be detected. Finally, the loss function in the original YOLOv5 network is modified to further refine the detection box position, ultimately improving the detection accuracy of insulating gloves. The disadvantage of this method is that the algorithm may miss detections in the case of overlapping people and hand occlusion.

[0120] (2) Obtaining raw image data from live working scenarios; the raw image data includes the target object; preprocessing the raw image data to obtain target image data; inputting the target image data into a trained insulating glove wearing detection model, and using the insulating glove wearing detection model to detect whether the target object in the raw image data is wearing insulating gloves, obtaining a target detection result; wherein the insulating glove wearing detection model includes a feature extraction network, a feature fusion network, and a prediction network. This method can effectively improve the accuracy of insulating glove wearing detection, and has higher detection efficiency, can reduce the human resource cost of supervision, and is conducive to improving the safety of live working.

[0121] The current deep learning-based detection method for insulating gloves at power grid work sites has the following main problems:

[0122] (1) Insufficient model detection accuracy: The existing target detection algorithm has relatively low positioning accuracy when identifying the wearing status of insulating gloves. Due to the complex and changeable background of power grid construction sites, when directly using existing general models for detection, the single grid prediction frame used is difficult to accurately locate the specific area where the gloves are worn, resulting in large positioning errors and difficulty in accurately determining whether workers are wearing insulating gloves.

[0123] (2) Difficulty in identifying small target features: Detecting the wearing of insulating gloves requires the complete recognition of both the human hand and the glove. When power grid workers are far away from the detection equipment, these features may appear as small pixel areas in the image. Existing target detection algorithms have difficulty capturing these subtle features and are prone to overlooking key information, leading to misjudgment of the wearing status.

[0124] (3) Impact of occlusion and angle changes: In actual working environments, workers’ hands may be partially obscured or their angles may change due to factors such as operating posture and obstructions, which makes the detection of wearing insulating gloves more complicated. Current technologies often have difficulty effectively dealing with such occlusion and angle changes, making it difficult to accurately detect the wearing of gloves in complex scenarios, increasing the risk of missed detection and false detection.

[0125] To address the aforementioned issues, the technical solution provided in this embodiment provides an effective and fast detection of insulated gloves (EFDIG) algorithm. This algorithm includes an Attention Module for Channel and Spatial Feature Enhancement (CSFE_AM). This algorithm extracts channel features through global average pooling and global max pooling, and enhances spatial features with a parameter-free attention mechanism. This allows the network to better capture small, distant figures in images of power grid operations, addressing the difficulty of detecting small targets in the insulating glove wear detection task. Furthermore, a Multi-Expansion Aggregation Module (MDAM) uses multiple dilated convolutions to filter out interference information and enhance the features of small targets, mitigating the problem of target feature information being overwhelmed during the convolution process. This addresses the issue of target occlusion in the insulating glove wear detection task. To address the issue of missed detection of individual insulating glove wear conditions, the algorithm introduces an Adaptive Intersection over Union (AD-IOU) loss function. The above improvements effectively solve the problem of insufficient accuracy of small target detection in the traditional insulating glove wearing detection model, and improve the detection accuracy of insulating gloves wearing.

[0126] The technical solution provided in this embodiment includes the following specific steps: constructing an insulating glove wearing detection dataset for training, constructing a channel space feature enhancement attention module, obtaining a multi-expansion convergence module, using a feature pyramid network (FPN) module for feature fusion, obtaining an adaptive intersection-over-union loss function, model training and verification, and applying the model. The corresponding method flow chart is as follows: Figure 5 shown.

[0127] The EFDIG algorithm structure of this embodiment is as follows Figure 6As shown. For the images of personnel on the power grid operation site, the channel feature map is first obtained by a two-layer neural network feature extraction module (PConv_BN_ELU); then the obtained channel feature map is input into the channel space feature enhancement attention module (CSFE_AM). The channel space feature enhancement attention module can filter out interference information, enhance the feature information of small targets, and enhance the insulating gloves wearing detection model's ability to extract the features of the hand features and insulating gloves of the entire power grid operation site personnel image, especially those of the power grid construction personnel who are far away from the video detection equipment; the multi-expansion aggregation module (MDAM) takes the insulating gloves feature map output by the channel space feature enhancement attention module as input, and first performs multi-scale feature extraction. The feature maps of different scales are extracted, and then the feature maps of different scales are fused to obtain the insulating gloves feature enhancement map, which enhances the feature expression of the insulating gloves worn by the power grid construction workers in the power grid operation site personnel image; the insulating gloves feature enhancement map is then passed through a layer of neural network feature extraction module, a channel space feature enhancement attention module and a multi-expansion aggregation module in sequence, and the above operation is repeated twice; then the insulating gloves feature enhancement map extracted by each multi-domain adaptation module is passed through the feature pyramid network (FPN) and the adaptive intersection-over-union loss function (AD-IOU) to improve the false detection and missed detection problems of insulating gloves, and finally the insulating gloves wearing detection result of the power grid operation site personnel image is obtained.

[0128] S1. Construct an insulating glove wearing detection dataset for training:

[0129] Real images of insulating gloves being worn can be collected during actual grid construction operations and routine maintenance. Furthermore, researchers can simulate improperly worn insulating gloves and situations where no gloves are worn at all, generating a first simulated image of improperly worn insulating gloves and a second simulated image of no gloves. This creates a certain number of negative samples, aiming to balance the number of positive and negative samples.

[0130] Based on the real image, the first simulated image, and the second simulated image, an insulating glove wearing detection dataset is obtained. For each image in the insulating glove wearing detection dataset, a professional annotation tool can be used to finely annotate the hand area of ​​the power grid construction worker in the image to obtain an insulating glove wearing detection annotation dataset. The annotation information includes the exact location of the hand area and the wearing category of the insulating gloves. A corresponding label file can be generated for each image. The label file records in detail the location information of the hand area and the wearing category of the insulating gloves. The wearing categories include hands wearing insulating gloves correctly, hands not wearing insulating gloves, and hands wearing insulating gloves improperly.

[0131] Through such data set construction and annotation work, staff can more comprehensively understand and analyze the wearing of insulating gloves by power grid construction workers during power grid construction operations, thereby providing strong data support for improving work safety and standardizing operating behaviors.

[0132] The specific example of constructing an insulating glove wearing detection dataset for training is as follows:

[0133] (1) Dataset Collection: In a real working environment, high-definition cameras were used to film power grid construction workers. Try to ensure that multiple power grid construction workers appear in the captured images, reflecting various insulating glove wearing detection scenarios. 5,400 real images of different working scenarios were collected to ensure data diversity and representativeness.

[0134] (2) Negative sample creation: To improve the accuracy of the model in identifying the wearing conditions of insulating gloves, negative samples were introduced. By simulating images of hands without gloves and improper wearing of gloves, such as gloves slipping off and fingers exposed, a first simulated image of improperly wearing insulating gloves and a second simulated image of not wearing insulating gloves were obtained to create a certain number of negative samples.

[0135] (3) Dataset annotation: Based on the real image, the first simulated image, and the second simulated image, an insulating glove wearing detection dataset is obtained. For each image in the insulating glove wearing detection dataset, a professional annotation tool is used to finely annotate the hand area of ​​the power grid construction worker. The annotation information includes the exact location of the hand area and the wearing status of the insulating gloves.

[0136] (4) Label generation: Generate a corresponding label file for each image in the insulating gloves wearing detection dataset. The label file records in detail the position information of the hand area and the wearing status of the insulating gloves, which are divided into correct wearing, not wearing, and irregular.

[0137] (5) Data preprocessing: A series of preprocessing operations are performed on the collected images, including adjusting the image size to adapt to the model input requirements, adding noise to improve the generalization ability of the model, and ensuring that the image quality meets the requirements of model training.

[0138] (6) Dataset storage and division: The pre-processed images and corresponding label files are organized and stored in the insulating gloves wearing detection annotated dataset. Subsequently, the set ratio is determined according to actual needs; according to the set ratio, the insulating gloves wearing detection annotated dataset is proportionally divided into training set, validation set, and test set, which are used for model training, performance verification, and final testing, respectively, to ensure that the model can be fully evaluated and optimized at different stages.

[0139] S2. Construct channel space feature enhanced attention module:

[0140] The channel space feature enhanced attention module can effectively reduce the number of parameters in the insulating gloves wearing detection model. At the same time, the channel space feature enhanced attention module can identify and strengthen the feature channels that are important for insulating gloves position recognition and suppress redundant or noisy channels. At the spatial feature level, the module fuses the insulating gloves position details in the low-order feature map with the rich contextual semantic information in the high-order feature map, so that the insulating gloves wearing detection model can more accurately locate and identify whether the construction workers are wearing insulating gloves correctly. The overall structure of the channel space feature enhanced attention module is as follows: Figure 3 shown.

[0141] The channel-space feature enhanced attention module uses the channel feature map output by the upper-layer neural network feature extraction module as the input feature map. Here, the channel feature map is denoted as F0. The size of the channel feature map F0 is C×H×W, where C represents the number of channels, H represents the height, and W represents the width. The specific steps of the channel-space feature enhanced attention module are as follows:

[0142] S2.1: Perform a convolution operation on the channel feature map F0 with a convolution kernel size k of 3, a step size s of 1, and a padding value p of 1 to obtain the first intermediate feature map F1 with the same spatial resolution and the number of channels C.

[0143] S2.2: Perform a global average pooling operation on the first intermediate feature map F1 to extract the hand features of the person and the features of the insulating gloves in the channel feature map, and obtain a second intermediate feature map F2 with a spatial resolution of 1×1 and a number of channels C; perform a global maximum pooling operation on the first intermediate feature map F1 to extract the hand features of the person and the features of the insulating gloves in the channel feature map, and obtain a third intermediate feature map F3 with a spatial resolution of 1×1 and a number of channels C; the second intermediate feature map F2 and the third intermediate feature map F3 can be element-wise added to obtain a fourth intermediate feature map F4; the fourth intermediate feature map F9 can be calculated through a layer of S-type (Sigmoid) activation function to obtain the channel feature weight W1.

[0144] S2.3: Perform a dot product of F0 and W1 to obtain the intermediate feature map F5 with adaptive enhancement of the channel feature, whose spatial resolution is H×W and the number of channels is C.

[0145] S2.4: Perform a dot product of the first intermediate feature map F1 and the channel feature weight W1 to obtain a sixth intermediate feature map F6 with adaptively enhanced channel features, where the spatial resolution of the sixth intermediate feature map F6 is H×W and the number of channels is C.

[0146] S2.5: The channel feature map F0 can be input into the parameter-free attention mechanism to obtain the seventh intermediate feature map F7 with adaptively enhanced spatial features, where the spatial resolution of the seventh intermediate feature map F7 is H×W and the number of channels is C.

[0147] S2.6: The fifth intermediate feature map F5, the sixth intermediate feature map F6 and the seventh intermediate feature map F7 can be added element by element to obtain the eighth intermediate feature map F8 with a spatial resolution of H×W and a number of channels of C, and a convolution operation with a convolution kernel size k of 3, a step size s of 1, and a padding value p of 1 is performed on the eighth intermediate feature map F8 to obtain an insulating glove feature map F9 with a spatial resolution of H×W and a number of channels of C.

[0148] S3. Obtain multiple expansion and convergence modules:

[0149] The input insulating gloves feature map is extracted from local and global features, and then the local and global features are fused to reduce the impact of reduced detection accuracy caused by occlusion and angle changes. The structure of the multi-expansion convergence module is as follows: Figure 4 shown.

[0150] The insulating glove feature map output by the channel-space feature enhancement attention module is used as the input feature map of the multi-expansion convergence module. Here, the insulating glove feature map F9 output by the channel-space feature enhancement attention module is denoted as C0. The size of the insulating glove feature map C0 is C×H×W, where C represents the number of channels, H represents the height, and W represents the width. The specific steps of the multi-expansion convergence module are as follows:

[0151] S3.1: Perform a maximum pooling operation on the insulating glove feature map C0 with a spatial resolution of H×W and a number of channels C without changing the spatial resolution and the number of channels, to obtain a first intermediate feature enhancement map C1 with a spatial resolution of H×W and a number of channels C.

[0152] S3.2: Perform a dilated convolution operation on the insulating glove feature map C0 with a convolution kernel size k of 3, a step size s of 1, a padding value p of 1, and a dilation value of 1 to obtain a second intermediate feature enhancement map C2 with the same spatial resolution and C number of channels.

[0153] S3.3: Perform an expansion convolution operation on the second intermediate feature enhancement image C2 with a convolution kernel size k of 3, a step size s of 1, a padding value p of 3, and a dilation value of 3 to obtain a third intermediate feature enhancement image C3 with the same spatial resolution and C number of channels.

[0154] S3.4: Perform an expansion convolution operation on the third intermediate feature enhancement image C3 with a convolution kernel size k of 3, a step size s of 1, a padding value p of 1, and a dilation value of 1 to obtain a fourth intermediate feature enhancement image C4 with an unchanged spatial resolution and a number of channels C.

[0155] S3.5: Concatenate the first intermediate feature enhancement map C1, the second intermediate feature enhancement map C2, the third intermediate feature enhancement map C3, and the fourth intermediate feature enhancement map C4 according to the channel dimension to obtain the fifth intermediate feature enhancement map C5 with a spatial resolution of H×W and a channel number of 4C.

[0156] S3.6: Perform a convolution operation on the fifth intermediate feature enhancement image C5 with a convolution kernel size k of 1, a step size s of 1, and a padding value p of 0 to obtain a sixth intermediate feature enhancement image C6 with an unchanged spatial resolution and a number of channels C.

[0157] S3.7: Input the sixth intermediate feature enhancement map C6 into the Sigmoid Linear Unit (SiLU) activation function for nonlinear enhancement to obtain the insulating glove feature enhancement map C7 as the output of the multi-expansion convergence module.

[0158] S4. Use the feature pyramid network module for feature fusion:

[0159] The Feature Pyramid Network (FPN) module is used to fuse the enhanced feature maps of insulating gloves at different scales. 、 、 is the input feature map of FPN, the feature pyramid network module is 、 、 Perform feature fusion, i.e., insulating gloves feature enhancement map Fusion 、 The enhanced feature map of the insulating gloves after upsampling and the enhanced feature map of the insulating gloves after fusion are as follows: ,Insulating gloves feature enhancement diagram Fusion Downsampling and The enhanced feature map of the insulating gloves after upsampling and the enhanced feature map of the insulating gloves after fusion are as follows: ,Insulating gloves feature enhancement diagram Fusion 、 The enhanced feature map of the insulating gloves after downsampling and the enhanced feature map of the insulating gloves after fusion are as follows: .

[0160] S5: Get the adaptive intersection-over-union loss function:

[0161] The adaptive intersection over union (AD-IOU) loss function is used to measure the gap between the prediction results of the insulating glove wearing detection model and the true label data.

[0162] The adaptive intersection over union loss function adds a loss metric based on the intersection over union (IOU) loss function. and , to solve the problem of inaccurate bounding box regression. The adaptive intersection-over-union loss function is 、 and The weighted sum of is shown in formula (1).

[0163] (1)

[0164] (2)

[0165] (3)

[0166] (4)

[0167] in It is the loss based on intersection-over-union. is the loss based on the center point distance, It is based on the loss of length and width.

[0168] In the intersection-over-union loss, It represents 1 minus the intersection-over-union ratio between the predicted box and the true box, that is, is the ratio of the overlap of two boxes to their union. The smaller the value of The larger it is, the higher the overlap between the predicted box and the true box, which means the higher the similarity between the predicted box and the true box.

[0169] exist in ,in is a hyperparameter, and are the center coordinates of the predicted box and the real box respectively, Represents the Euclidean distance between the center coordinates of the predicted box and the center coordinates of the real box, and are the squares of the width and height of the ground-truth box, respectively. The loss function is used to reduce the distance between the center coordinates of the predicted box and the center coordinates of the real box, making the predicted box closer to the real box.

[0170] exist middle, and are the width and height of the prediction box respectively. Indicates taking the absolute value. The loss function is used to reduce the difference between the length and width of the predicted box and the length and width of the real box.

[0171] The IOU loss is calculated as follows:

[0172] (5)

[0173] in, and Represent the true box and the predicted box respectively, represents the intersection of the real box and the predicted box, Represents the union of the true box and the predicted box.

[0174] Box regression refers to the process of adjusting the position and size of the bounding box predicted for each grid cell in object detection tasks. For example, in the YOLO algorithm, the input image is divided into several grid cells, and each grid cell is responsible for predicting an object whose center point falls within that grid cell. However, these initial predicted bounding boxes are often not accurate enough, so box regression is needed to further adjust their position and size to more accurately match the actual object.

[0175] S6. Model training and validation:

[0176] Based on the channel-space features, the attention module and the multi-expansion convergence module are enhanced to obtain the parameter-adjusted model. Based on the parameter-adjusted model, the parameters of each layer are trained and updated. All neural network parameters of the parameter-adjusted model are initialized, and hyperparameters related to the parameter-adjusted model are set, such as the number of training rounds, batch size, optimizer selection, learning rate, and maximum value of the gradient clipping strategy.

[0177] After initializing the parameters, the insulating glove wearing detection dataset processed by S1 is divided into a training set, a validation set, and a test set. The training and validation set data are divided into multiple batches. Each batch of training set data is input into the parameter adjustment model for training, and the training loss value 1oss for that batch is obtained. After completing one round of training for all batches of data in the entire training set (in actual training, there will likely be multiple rounds), the validation set is input into the parameter adjustment model according to the batch, and the corresponding batch loss value batch loss is obtained. During training and validation, the parameter adjustment model will automatically learn and adjust its parameters based on each loss and batch loss. When the training process has been carried out for one or more rounds until the batch loss value converges, the training ends and the insulating glove wearing detection model is obtained.

[0178] S7. Apply the model:

[0179] The trained insulating glove wearing detection model is used to detect images of personnel at power grid operation sites, realizing automatic detection of insulating gloves worn by power grid construction personnel at power grid operation sites.

[0180] In actual engineering applications, the collected images of grid operation site personnel are input into the insulating glove wearing detection model. The insulating glove wearing detection model automatically detects the images of grid operation site personnel, obtains the insulating glove wearing detection results of the grid operation site personnel images, and displays the detection results on the grid operation site personnel images, including whether the wearing is correct, the target position in the image, and the detection confidence.

[0181] Specific examples are as follows:

[0182] The insulating gloves wearing detection dataset is used. The dataset contains 5,400 images, including 16,133 labeled objects and 3 categories, which is used to train the insulating gloves wearing detection model.

[0183] (1) Dataset annotation: The hand areas of power grid construction workers extracted from the insulating gloves wearing detection dataset are annotated to show the wearing of insulating gloves, and the quality score is annotated for each wearing condition of insulating gloves.

[0184] (2) Dataset division: The dataset is randomly divided into training set, validation set, and test set in a ratio of 6:2:2, that is, the training set contains 3240 images, and the validation set and test set each contain 1080 images.

[0185] (3) Model training: Based on the data in the training set and the validation set, the insulating glove wearing detection model is trained. The entire training process is set to 450 epochs (rounds), and the parameters in the model are adjusted according to the loss value in each iteration process.

[0186] (4) Training results: After adjusting the relevant parameters, the accuracy of the insulating gloves wearing detection model tends to be stable, mainly fluctuating between 90% and 95%.

[0187] The technical solution provided in this embodiment has the following improvements:

[0188] (1) The channel spatial feature enhanced attention module is constructed by using global pooling technology and combining it with a parameter-free attention mechanism. It enhances global context information and improves the ability of the model backbone network to extract blurred targets. It has good detection capabilities for targets of different sizes.

[0189] (2) The multi-expansion convergence module can better capture the local and global relationships in the target feature space by using dilated convolution and channel splicing and fusion of feature maps.

[0190] (3) Adaptive intersection-over-union loss function, which comprehensively balances the loss calculation of position, shape and overlap, improves the positioning accuracy of insulating gloves in complex environments;

[0191] The technical solution provided in this embodiment has the following beneficial effects:

[0192] (1) In the channel spatial feature enhancement attention module, spatial features and spatial features are extracted simultaneously to improve the model's ability to extract the detailed features of the insulating gloves worn by power grid construction workers.

[0193] (2) By using multiple dilated convolutions and fusing multi-scale feature maps, the model’s detection performance of occluded insulating gloves in images of personnel working on power grid sites is improved.

[0194] (3) An adaptive intersection-over-union loss function is used to improve the model's positioning accuracy for grid construction workers wearing insulating gloves in complex grid operation environments. This is achieved by calculating the Euclidean distance between the center point of the predicted frame and the center point of the true frame, and measuring the difference in the aspect ratio between the predicted frame and the true frame.

[0195] In summary, the technical solution provided by this embodiment has higher detection accuracy, and the insulating glove wearing detection model has fewer parameters and calculations, which can meet the high real-time requirements of the insulating glove wearing detection task in power grid operation scenarios.

[0196] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0197] Based on the same inventive concept, embodiments of the present application also provide a device for detecting the wearing of insulating gloves at power grid work sites, which is used to implement the aforementioned method for detecting the wearing of insulating gloves at power grid work sites. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the device for detecting the wearing of insulating gloves at power grid work sites provided below can be found in the aforementioned method for detecting the wearing of insulating gloves at power grid work sites, and will not be further elaborated here.

[0198] In an exemplary embodiment, Figure 7 As shown, a detection device for wearing insulating gloves at a power grid operation site is provided, wherein:

[0199] The channel feature map acquisition module 701 is used to extract channel features from the images of the personnel working on the power grid operation site according to the neural network feature extraction module of the insulating glove wearing detection model to obtain a channel feature map;

[0200] The insulating glove feature map acquisition module 702 is configured to extract the hand features of the person and the insulating glove features in the channel feature map based on the channel spatial feature enhancement attention module of the insulating glove wearing detection model, thereby obtaining an insulating glove feature map. The channel spatial feature enhancement attention module is constructed based on a global pooling operation combined with a parameter-free attention mechanism.

[0201] A feature enhancement map acquisition module 703 is configured to extract local and global features from the insulating glove feature map and then fuse them according to the multi-expansion convergence module of the insulating glove wearing detection model to obtain an insulating glove feature enhancement map;

[0202] The detection result acquisition module 704 is used to determine whether the power grid construction personnel in the power grid operation site personnel image are wearing insulating gloves, whether the insulating gloves are worn correctly, and the target position of the power grid operation site personnel image based on the insulating gloves feature enhancement map, and obtain the insulating gloves wearing detection result of the power grid operation site personnel image.

[0203] In one embodiment, the insulating glove feature map acquisition module 702 is further used to: perform a convolution operation on the channel feature map according to the channel space feature enhancement attention module of the insulating glove wearing detection model, extract the hand features of the person and the features of the insulating gloves in the channel feature map, and obtain a first intermediate feature map; perform a global average pooling operation and a global maximum pooling operation on the channel feature map according to the channel space feature enhancement attention module of the insulating glove wearing detection model, extract the hand features of the person and the features of the insulating gloves in the channel feature map, and obtain a second intermediate feature map and a third intermediate feature map to obtain a channel feature weight; obtain a fifth intermediate feature map according to the channel feature map and the channel feature weight; obtain a sixth intermediate feature map according to the first intermediate feature map and the channel feature weight; input the channel feature map into the parameter-free attention mechanism to obtain a seventh intermediate feature map; add the fifth intermediate feature map, the sixth intermediate feature map, and the seventh intermediate feature map element-wise to obtain an eighth intermediate feature map, and perform a convolution operation on the eighth intermediate feature map to obtain an insulating glove feature map.

[0204] In one embodiment, the insulating glove feature map acquisition module 702 is further used to: add the second intermediate feature map and the third intermediate feature map element-wise to obtain a fourth intermediate feature map; and obtain a channel feature weight based on the fourth intermediate feature map and the activation function.

[0205] In one embodiment, the feature enhancement map acquisition module 703 is further used to: perform a maximum pooling operation on the insulating glove feature map according to the multi-expansion convergence module of the insulating glove wearing detection model to obtain a first intermediate feature enhancement map; perform multiple consecutive expansion convolution operations on the insulating glove feature map according to the multi-expansion convergence module of the insulating glove wearing detection model to obtain a second intermediate feature enhancement map, a third intermediate feature enhancement map and a fourth intermediate feature enhancement map; splice the first intermediate feature enhancement map, the second intermediate feature enhancement map, the third intermediate feature enhancement map and the fourth intermediate feature enhancement map according to the channel dimension to obtain a fifth intermediate feature enhancement map; perform a convolution operation on the fifth intermediate feature enhancement map to obtain a sixth intermediate feature enhancement map; and perform nonlinear enhancement on the sixth intermediate feature enhancement map according to a nonlinear activation function to obtain an insulating glove feature enhancement map.

[0206] In one embodiment, the device also includes a model training module, which is used to: obtain an insulating glove wearing detection training set and an insulating glove wearing detection verification set; input the insulating glove wearing detection training image in the insulating glove wearing detection training set into the parameter to be adjusted model to obtain a first prediction result; obtain a first training loss value based on the first prediction result, the labeled data corresponding to the insulating glove wearing detection training image and the adaptive intersection-over-union loss function; the adaptive intersection-over-union loss function is obtained based on the intersection-over-union loss, the center point distance loss and the length-width loss; input the insulating glove wearing detection verification image in the insulating glove wearing detection verification set into the parameter to be adjusted model to obtain a second prediction result; obtain a second training loss value based on the second prediction result, the labeled data corresponding to the insulating glove wearing detection verification image and the adaptive intersection-over-union loss function; adjust the parameters of the parameter to be adjusted model based on the first training loss value and the second training loss value until the second training loss value tends to converge, thereby obtaining an insulating glove wearing detection model.

[0207] In one embodiment, the model training module is further used to: collect real images of insulating gloves being worn at power grid operation sites and during operation and maintenance; obtain a first simulated image of insulating gloves being worn improperly and a second simulated image of not wearing insulating gloves; obtain an insulating glove wearing detection dataset based on the real image, the first simulated image, and the second simulated image; label the insulating glove wearing detection dataset with wearing categories to obtain an insulating glove wearing detection labeled dataset; the wearing categories include hands wearing insulating gloves correctly, hands not wearing insulating gloves, and hands wearing insulating gloves improperly; and divide the insulating glove wearing detection labeled dataset into an insulating glove wearing detection training set and an insulating glove wearing detection verification set according to a set ratio.

[0208] Each module in the aforementioned device for detecting the wearing of insulating gloves at power grid work sites can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0209] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data of an embodiment of a method for detecting the wearing of insulating gloves at a power grid operation site. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for detecting the wearing of insulating gloves at a power grid operation site is implemented.

[0210] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0211] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0212] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0213] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0214] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0215] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0216] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0217] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for detecting the wearing of insulating gloves at a power grid operation site, characterized in that: The method comprises: Based on the neural network feature extraction module of the insulating glove wearing detection model, channel feature extraction is performed on the images of personnel working on the power grid to obtain a channel feature map. According to the channel spatial feature enhancement attention module of the insulating glove wearing detection model, a convolution operation is performed on the channel feature map to extract the hand features of the person and the features of the insulating gloves in the channel feature map to obtain a first intermediate feature map; Performing a global average pooling operation on the first intermediate feature map according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model to obtain a second intermediate feature map; Performing a global maximum pooling operation on the first intermediate feature map according to the channel spatial feature enhancement attention module of the insulating glove wearing detection model to obtain a third intermediate feature map; Adding the second intermediate feature map and the third intermediate feature map element-wise to obtain a fourth intermediate feature map; Calculate the fourth intermediate feature map through a sigmoid activation function to obtain a channel feature weight; Obtaining a fifth intermediate feature map according to the channel feature map and the channel feature weight; Obtaining a sixth intermediate feature map according to the first intermediate feature map and the channel feature weights; Inputting the channel feature map into the parameter-free attention mechanism to obtain a seventh intermediate feature map; adding the fifth intermediate feature map, the sixth intermediate feature map, and the seventh intermediate feature map element-wise to obtain an eighth intermediate feature map, and performing a convolution operation on the eighth intermediate feature map to obtain an insulating glove feature map; performing a maximum pooling operation on the insulating glove feature map according to the multi-expansion convergence module of the insulating glove wearing detection model to obtain a first intermediate feature enhancement map; According to the multi-expansion and convergence module of the insulating glove wearing detection model, the insulating glove feature map is subjected to multiple consecutive expansion and convolution operations to obtain a second intermediate feature enhancement map, a third intermediate feature enhancement map, and a fourth intermediate feature enhancement map; splicing the first intermediate feature enhancement map, the second intermediate feature enhancement map, the third intermediate feature enhancement map, and the fourth intermediate feature enhancement map according to the channel dimension to obtain a fifth intermediate feature enhancement map; performing a convolution operation on the fifth intermediate feature enhancement map to obtain a sixth intermediate feature enhancement map; performing nonlinear enhancement on the sixth intermediate feature enhancement map according to a nonlinear activation function to obtain an insulating glove feature enhancement map; Based on the insulating gloves feature enhancement map, it is determined whether the power grid construction personnel in the power grid operation site personnel image are wearing insulating gloves, whether the insulating gloves are worn correctly, and the target position of the power grid operation site personnel image is determined, and an insulating gloves wearing detection result of the power grid operation site personnel image is obtained.

2. The method according to claim 1, characterized in that The method further comprises: Obtain the insulating gloves wearing detection training set and the insulating gloves wearing detection verification set; Inputting the insulating glove wearing detection training images in the insulating glove wearing detection training set into the model to be adjusted in parameters to obtain a first prediction result; Obtaining a first training loss value based on the first prediction result, the labeled data corresponding to the insulating glove wearing detection training image, and an adaptive intersection-over-union loss function; the adaptive intersection-over-union loss function is obtained based on intersection-over-union loss, center point distance loss, and length-width loss; Inputting the insulating glove wearing detection verification images in the insulating glove wearing detection verification set into the parameter-adjusted model to obtain a second prediction result; Obtaining a second training loss value according to the second prediction result, the labeled data corresponding to the insulating glove wearing detection verification image, and the adaptive intersection-over-union loss function; According to the first training loss value and the second training loss value, the parameters of the model to be adjusted are adjusted until the second training loss value tends to converge, thereby obtaining an insulating glove wearing detection model.

3. The method according to claim 2, characterized in that Obtaining a first training loss value according to the first prediction result, the labeled data corresponding to the insulating glove wearing detection training image, and an adaptive intersection-over-union loss function includes: According to the adaptive intersection-over-union loss function, the gap between the first prediction result and the labeled data corresponding to the insulating glove wearing detection training image is measured to obtain a first training loss value.

4. The method according to claim 2, characterized in that The obtaining of the insulating gloves wearing detection training set and the insulating gloves wearing detection verification set includes: Collect real images of people wearing insulating gloves at power grid operation sites and during operation and maintenance; Acquire a first simulated image of an insulated glove not being worn properly and a second simulated image of an insulated glove not being worn; Obtaining an insulating glove wearing detection dataset based on the real image, the first simulated image, and the second simulated image; Performing wearing category labeling on the insulating gloves wearing detection dataset to obtain an insulating gloves wearing detection labeling dataset; the wearing categories include hands correctly wearing insulating gloves, hands not wearing insulating gloves, and hands wearing insulating gloves improperly; According to a set ratio, the insulating gloves wearing detection annotated dataset is divided into an insulating gloves wearing detection training set and an insulating gloves wearing detection verification set.

5. A detection device for wearing insulating gloves at a power grid operation site, characterized in that: The device comprises: A channel feature map acquisition module is used to extract channel features from images of personnel working on the power grid operation site based on the neural network feature extraction module of the insulating glove wearing detection model to obtain a channel feature map; The insulating gloves feature map acquisition module is used to perform a convolution operation on the channel feature map according to the channel space feature enhancement attention module of the insulating gloves wearing detection model, extract the hand features of the person and the features of the insulating gloves in the channel feature map, and obtain a first intermediate feature map; according to the channel space feature enhancement attention module of the insulating gloves wearing detection model, perform a global average pooling operation on the first intermediate feature map to obtain a second intermediate feature map; according to the channel space feature enhancement attention module of the insulating gloves wearing detection model, perform a global maximum pooling operation on the first intermediate feature map to obtain a third intermediate feature map; the second intermediate feature map is obtained by adding the second intermediate feature map to the channel space feature enhancement attention module. The feature map and the third intermediate feature map are element-wise added to obtain a fourth intermediate feature map; the fourth intermediate feature map is calculated by a sigmoid activation function to obtain a channel feature weight; a fifth intermediate feature map is obtained based on the channel feature map and the channel feature weight; a sixth intermediate feature map is obtained based on the first intermediate feature map and the channel feature weight; the channel feature map is input into a parameter-free attention mechanism to obtain a seventh intermediate feature map; the fifth intermediate feature map, the sixth intermediate feature map, and the seventh intermediate feature map are element-wise added to obtain an eighth intermediate feature map, and a convolution operation is performed on the eighth intermediate feature map to obtain an insulating glove feature map; A feature enhancement map acquisition module is configured to perform a maximum pooling operation on the insulating glove feature map according to the multi-expansion convergence module of the insulating glove wearing detection model to obtain a first intermediate feature enhancement map; perform multiple consecutive expansion convolution operations on the insulating glove feature map according to the multi-expansion convergence module of the insulating glove wearing detection model to obtain a second intermediate feature enhancement map, a third intermediate feature enhancement map, and a fourth intermediate feature enhancement map; concatenate the first intermediate feature enhancement map, the second intermediate feature enhancement map, the third intermediate feature enhancement map, and the fourth intermediate feature enhancement map according to the channel dimension to obtain a fifth intermediate feature enhancement map; perform a convolution operation on the fifth intermediate feature enhancement map to obtain a sixth intermediate feature enhancement map; and perform nonlinear enhancement on the sixth intermediate feature enhancement map according to a nonlinear activation function to obtain an insulating glove feature enhancement map. The detection result acquisition module is used to determine whether the power grid construction personnel in the power grid operation site personnel image are wearing insulating gloves, whether the insulating gloves are worn correctly, and the target position of the power grid operation site personnel image based on the insulating gloves feature enhancement map, and obtain the insulating gloves wearing detection result of the power grid operation site personnel image.

6. The device according to claim 5, characterized in that The device also includes a model training module for obtaining an insulating glove wearing detection training set and an insulating glove wearing detection verification set; inputting the insulating glove wearing detection training images in the insulating glove wearing detection training set into the parameter-to-be-adjusted model to obtain a first prediction result; obtaining a first training loss value based on the first prediction result, the labeled data corresponding to the insulating glove wearing detection training images, and an adaptive intersection-over-union loss function; the adaptive intersection-over-union loss function is obtained based on the intersection-over-union loss, the center point distance loss, and the length-width loss; inputting the insulating glove wearing detection verification images in the insulating glove wearing detection verification set into the parameter-to-be-adjusted model to obtain a second prediction result; obtaining a second training loss value based on the second prediction result, the labeled data corresponding to the insulating glove wearing detection verification images, and the adaptive intersection-over-union loss function; adjusting the parameters of the parameter-to-be-adjusted model based on the first training loss value and the second training loss value until the second training loss value converges, thereby obtaining an insulating glove wearing detection model.

7. The device according to claim 6, characterized in that The model training module is also used to: collect real images of people wearing insulating gloves at power grid operation sites and during operation and maintenance; obtain a first simulated image of people wearing insulating gloves improperly and a second simulated image of people not wearing insulating gloves; An insulating glove wearing detection dataset is obtained based on the real image, the first simulated image, and the second simulated image; the insulating glove wearing detection dataset is labeled with wearing categories to obtain an insulating glove wearing detection labeled dataset; the wearing categories include hands wearing insulating gloves correctly, hands not wearing insulating gloves, and hands wearing insulating gloves improperly; and according to a set ratio, the insulating glove wearing detection labeled dataset is divided into an insulating glove wearing detection training set and an insulating glove wearing detection verification set.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Safety helmet wearing detection method and device based on deep learning, equipment and medium

    CN114782986A

  • Edge detection system for power grid inspection and monitoring

    CN116846059A