Power worker goggles wearing detection method based on instance segmentation
By adopting an instance segmentation method in the wear detection of goggles for electric workers, combining multi-scale feature enhancement and noise perturbation mechanisms, the problems of complex background processing and boundary refinement in the prior art are solved, and higher detection accuracy and safety are achieved.
Patent Information
- Application Number
- CN202510381179.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing detection methods for goggles wear for power workers have shortcomings in complex background processing, border refinement, multi-instance distinction ability and target state perception ability, which makes it difficult to meet high requirements for detection accuracy and safety.
Using an instance segmentation-based detection method, the EGSNet network is designed, combining multi-scale feature enhancement, goggles prompt embedded features and noise perturbation mechanism to improve the robustness and accuracy of detection.
It significantly improves the accuracy and safety of goggles wear detection, reduces the rates of missed and false detection, and performs excellently in complex backgrounds and multiple instance scenarios.
Smart Images

Figure CN120147756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electric power, and particularly to a method for detecting the wearing of safety goggles for electric power workers based on instance segmentation. Background Technique
[0002] In electric power operations, safety goggles are a key part of workers' protective equipment, which can effectively protect the eyes from potential harm during operations. However, in the actual operation environment, due to reasons such as the heavy weight of the goggle equipment, uncomfortable wearing, and personal usage habits, electric power workers may sometimes not wear the safety goggles correctly or even not wear them at all, thus resulting in operation safety risks. Therefore, real-time monitoring of the wearing situation of safety goggles for electric power workers has become one of the important measures to ensure operation safety.
[0003] Currently, common detection methods are mostly based on semantic segmentation or image detection techniques. Although these methods can judge the wearing situation of safety goggles to a certain extent, they have obvious deficiencies in practical applications. Among them, the semantic segmentation method identifies the target by dividing the image into different category regions. However, in the complex scenarios of electric power operations, especially in the case of complex backgrounds (such as environments with colors similar to those of the safety goggles), target occlusion, or changing lighting conditions, its ability to perform fine-grained processing on the boundaries of safety goggles is insufficient, easily resulting in blurred boundaries or missed targets; the image detection method mainly identifies the presence or absence of the target, but cannot accurately judge the wearing state of the safety goggles (such as position, angle, and tightness), so it is difficult to meet the high requirements for detection accuracy in actual operations.
[0004] Based on the above, the current detection of the wearing of safety goggles for electric power workers mainly uses deep learning methods, which have the following disadvantages:
[0005] 1. Existing methods pay more attention to high-level semantic information and ignore the rich local information in shallow features, resulting in insufficient robustness of the model to small targets;
[0006] 2. Most existing methods directly generate masks based on feature maps, and the feature maps cannot accurately capture the fine-grained boundary information of instances, and cannot generate fine safety goggle boundaries, resulting in overly blurred safety goggle boundaries when generating masks; the above problems will be particularly obvious when the target shape is complex or there is occlusion;
[0007] 3. Existing methods often have missed detections or false detections when dealing with complex backgrounds, especially when the texture or color between the background and the safety goggles is similar (such as the safety goggles are similar in color to the surrounding environment), and when the safety goggles are worn irregularly (such as the safety goggles do not completely cover the wearing area or are worn loosely), the misjudgment rate is relatively high.
[0008] Therefore, an instance segmentation-based method for detecting whether a power worker wears safety goggles is invented to solve the problems of insufficient processing of complex backgrounds, boundary refinement, multi-instance discrimination ability, and target state perception ability in the prior art. By combining multi-scale feature enhancement, safety goggles prompt embedding features, and a noise perturbation mechanism, the present invention significantly improves the robustness and accuracy of safety goggles wearing detection, provides an efficient and reliable solution for the safety monitoring of power operations, will greatly enhance the safety of power operations, reduce the risk of work-related injuries, and bring more reliable operation guarantee for power enterprises. Summary of the Invention
[0009] To solve the above technical problems, according to one aspect of the present invention, the following technical solutions are provided:
[0010] An instance segmentation-based method for detecting whether a power worker wears safety goggles, which includes the following specific steps:
[0011] S1: Construct a dataset D of original images of power workers wearing safety goggles during operations 1 ;
[0012] S2: Preprocess D 1 to obtain a new dataset D of images of power workers wearing safety goggles during operations 2 ;
[0013] S3: Design an EGSNet network and use the network to process any image of a power worker during operation in D 2 to obtain the detection result of whether the power worker wears safety goggles;
[0014] S4: Load the training set and validation set in D 2 into the EGSNet network for training to update the network parameters, and then use the test set in D 2 as the input to verify the network effect;
[0015] S5: Apply the trained EGSNet network to the safety goggles detection task and integrate it into the power construction site monitoring system;
[0016] The specific steps of S3 are as follows:
[0017] S31: Construct a multi-scale feature map generation module to generate multi-scale feature maps with enhanced features;
[0018] S32: Use the multi-scale feature map fusion module to fuse different-scale P 1 , P 2 , P 3 and P 4 to output a mask feature map F mask ;
[0019] S33: Construct the Q2P-Cross module, and use the cross-attention mechanism and the multi-layer perceptron MLP to process P 1 、P 2 、P 3 and P 4 for feature fusion and enhancement processing to generate a goggle prompt embedding feature matrix;
[0020] S34: Construct the B2F-Diffu module to generate a noise filter NF related to each goggle instance;
[0021] S35: Use the BMD module to construct a power worker goggle instance mask feature map to obtain the wearing detection result of the power worker goggles.
[0022] As a preferred solution of a power worker goggle wearing detection method based on instance segmentation according to the present invention, wherein: the specific steps of S1 are as follows:
[0023] S11: Install, including but not limited to, a cloth control ball and a camera shooting device at the power grid construction operation site, and collect power worker operation images under different lighting, weather, shooting angles, and different postures of power workers;
[0024] S12: Among all the collected power worker operation images, use the images with correctly worn goggles as positive samples;
[0025] S13: Use the images without wearing, wearing incorrectly, and wearing other glasses as negative samples, and the number of positive and negative samples is relatively balanced, and finally form an original dataset D of power worker goggle wearing operation images 1 .
[0026] As a preferred solution of a power worker goggle wearing detection method based on instance segmentation according to the present invention, wherein: the specific steps of S2 are as follows:
[0027] S21: Perform data cleaning operations on the collected D 1 to remove fuzzy, low-quality, and duplicate samples;
[0028] S22: Construct a ground truth label GT sample, and the GT sample contains the bounding box, instance mask, and instance identification annotation data of each instance in the image. Specifically, use the annotation tool to annotate D 1 and annotate the bounding box of the goggle wearing area of the power worker in D 1 . The bounding box attributes include its position information and category information, where the category information includes two types: goggles worn, goggles not worn; in addition, use the annotation tool to annotate D 1Mask for the drawing example of the operation image of a power worker wearing safety goggles, D 1 For the operation image of a power worker not wearing safety goggles, no instance mask needs to be drawn;
[0029] S23: To enhance the data diversity and improve the model robustness, data augmentation techniques will be adopted to perform random rotation, cropping, flipping, and lighting adjustment operations on the images in D 1 to obtain the dataset after data augmentation;
[0030] S24: Segment the dataset after data augmentation and divide it into a training set, a validation set, and a test set according to the ratio of 8:1:1 to obtain the new dataset D 2 of the operation images of power workers wearing safety goggles.
[0031] As a preferred solution of a method for detecting whether a power worker wears safety goggles based on instance segmentation according to the present invention, wherein: the specific steps of S31 are as follows:
[0032] S311: Perform a pooling operation on the image Y 0 ∈D 2 with a spatial resolution of H×W and a channel number of C to obtain the first shallow feature map Y 1 of safety goggles with a spatial resolution of H / 2×W / 2 and a channel number of C, and then perform two Conv1×1 convolution operations to obtain the second shallow feature map Y 2 of safety goggles with a spatial resolution of H / 2×W / 2 and a channel number of 2C;
[0033] S312: Input Y 2 into the feature extraction sub-module for safety goggle wearing detection to extract multi-scale feature maps, including the first layer feature map F 1 , the second layer feature map F 2 , the third layer feature map F 3 and the fourth layer feature map F 4 ;
[0034] S313: Design an F-Agg network and adopt the propagation method from deep semantic features to shallow features to perform feature enhancement on F 1 , F 2 , F 3 and F 4 to obtain the multi-scale feature maps with enhanced safety goggle features, including the first layer enhanced feature map P 1 , the second layer enhanced feature map P 2 , the third layer enhanced feature map P 3 and the fourth layer enhanced feature map P 4 .
[0035] As a preferred solution of a method for detecting the wearing of safety goggles for electric power workers based on instance segmentation according to the present invention, wherein: the specific steps of S312 are as follows:
[0036] S3121: Input Y 2 to the first layer Conv_Layer1 of the safety goggles wearing detection feature extraction sub-module. Perform a Conv5×5 convolution operation on Y 2 using a convolution kernel of size 5×5 to capture local features. Perform a batch normalization BN operation to reduce the problem of gradient disappearance. Perform an operation using the ReLU activation function to introduce non-linearity and enhance the model's expression ability, thereby achieving preliminary feature extraction and obtaining the first-layer feature map F of the safety goggles wearing detection 1 ;
[0037] S3122: Input F 1 to the second layer Conv_Layer2 of the safety goggles wearing detection feature extraction sub-module. Perform a Conv1×1 convolution operation on F 1 using a convolution kernel of size 1×1 to adjust the number of channels and change the depth of the feature map. Perform BN operation and ReLU function operation to reduce gradient disappearance and introduce non-linearity. Then perform a Conv3×3 convolution operation using a convolution kernel of size 3×3 to capture local information in the image. Perform a Conv1×1 convolution operation using a convolution kernel of size 1×1 to restore the number of channels. Perform BN operation and ReLU function operation on the feature map after restoring the channels, and obtain the second-layer feature map F of the safety goggles wearing detection 2 ;
[0038] S3123: After performing residual addition on F 1 and F 2 , input it to the third layer Conv_Layer3 of the safety goggles wearing detection feature extraction sub-module. Perform a Conv1×1 convolution operation using a convolution kernel of size 1×1 to adjust the number of channels and change the depth of the feature map. Perform BN operation and ReLU function operation to reduce gradient disappearance and introduce non-linearity. Then perform a Conv3×3 convolution operation using a convolution kernel of size 3×3 to capture local information in the image. Perform a Conv1×1 convolution operation using a convolution kernel of size 1×1 to restore the number of channels. Perform BN operation and ReLU function operation on the feature map after restoring the channels, and obtain the third-layer feature map F of the safety goggles wearing detection 3 ;
[0039] S3124: Input F 2 and F 3After performing residual addition, it is input to the fourth layer Conv_Layer4 of the goggle-wearing detection feature extraction sub-module. A 1×1 convolutional kernel is used for Conv1×1 convolution operation to adjust the number of channels and change the depth of the feature map. BN operation and ReLU function operation are used to reduce gradient disappearance and introduce non-linearity. Then, a 3×3 convolutional kernel is used for Conv3×3 convolution operation to capture local information in the image. A 1×1 convolutional kernel is used for Conv1×1 convolution operation to restore the number of channels. BN operation and ReLU function operation are performed on the feature map after restoring the channels to obtain the fourth layer feature map F of goggle-wearing detection. 4 ;
[0040] The specific steps of S313 are as follows:
[0041] S3131: Taking F 4 as the input, in order to unify the number of channels, a 1×1 convolutional kernel is used to perform Conv1×1 convolution operation on F 4 to obtain the feature map F' 4 . For F' 4 , a 3×3 convolutional kernel is used to perform Conv3×3 convolution operation to enhance spatial information, and ReLU activation function operation is used to introduce non-linearity to obtain the fourth layer enhanced feature map P of goggle-wearing detection 4 ;
[0042] S3132: A 1×1 convolutional kernel is used to perform Conv1×1 convolution operation on F 3 to obtain the feature map F' 3 . The feature maps F' 4 , F' 3 and P 4 are added together to obtain M 3 . A 3×3 convolutional kernel is used to perform Conv3×3 convolution operation on M 3 to obtain the third layer enhanced feature map P of goggle-wearing detection 3 ;
[0043] S3133: A 1×1 convolutional kernel is used to perform Conv1×1 convolution operation on F 2 to obtain the feature map F' 2 . The feature maps M 3 , F' 2 and P 3 are added together to obtain M 2 . A 3×3 convolutional kernel is used to perform Conv3×3 convolution operation on M 2 to obtain the second layer enhanced feature map P of goggle-wearing detection 2 ;
[0044] S3134: Take F 1 as the input, and obtain the feature map F' after performing a 1×1 convolution 1 ; add the feature maps M 2 , F' 1 and P 2 to get M 1 ; perform a 3×3 convolution operation on M 1 to obtain the first enhanced feature map P for goggle wearing detection 1 ;
[0045] S3135: Through convolution operations and the fusion of feature maps, the multi-scale feature maps required for goggle wearing detection are enhanced, and finally the multi-scale feature maps P 1 , P 2 , P 3 and P 4 for goggle feature enhancement are obtained;
[0046] In the F-Agg network, the definitions for enhancing the goggle features of F 1 , F 2 , F 3 and F 4 are as follows:
[0047] F′ i = Φ Conv1×1 (F i ), 1 ≤ i ≤ 4
[0048] P j = Φ Conv3×3 (M j ), 1 ≤ j ≤ 4
[0049] M k = M k+1 + P k+1 + F′ k , 1 ≤ k ≤ 3
[0050] where F i is the input multi-scale feature map; Φ Conv1×1 represents the 1×1 convolution operation, which is used to adjust the number of channels through convolution; F' i represents the feature map after adjusting the channels, where F' 1 is the feature map with adjusted channels obtained after performing a 1×1 convolution operation on F 1 , F' 2 is the feature map with adjusted channels obtained after performing a 1×1 convolution operation on F 2 , F' 3 is the feature map with adjusted channels obtained after performing a 1×1 convolution operation on F 3 , F'4 is the feature map with adjusted channels obtained after performing a Conv1×1 convolution operation on F 4 ; Φ Conv3×3 represents a Conv3×3 convolution operation for enhancing spatial information, M j represents the feature map after the fusion of goggles features, where M 4 is the deepest feature map, i.e., F' 4 , M 3 is obtained by fusing and adding the features of M 4 , P 4 and F' 3 ; M 2 is obtained by fusing and adding the features of M 3 , P 3 and F' 2 ; M 1 is obtained by fusing and adding the features of M 2 , P 2 and F' 1 ; P j is the multi-scale feature map for enhancing goggles features, P 1 is the first-layer enhanced feature map obtained after performing a Conv3×3 convolution operation on M 1 ; P 2 is the second-layer enhanced feature map obtained after performing a Conv3×3 convolution operation on M 2 ; P 3 is the third-layer enhanced feature map obtained after performing a Conv3×3 convolution operation on M 3 ; P 4 is the fourth-layer enhanced feature map obtained after performing a Conv3×3 convolution operation on M 4 .
[0051] As a preferred solution of the method for detecting the wearing of electrician goggles based on instance segmentation according to the present invention, where: the specific steps of S32 are as follows:
[0052] S321: Perform a Fusion operation on P 1 , P 2 , P 3 and P 4 to obtain global semantic information and fine local detail information, and obtain the fused fifth-layer enhanced feature map P 5 ;
[0053] S322: After performing a Conv1×1 convolution operation on P 5 using a convolution kernel of size 1×1, then perform a batch normalization BN operation and a ReLU function operation to obtain the mask feature map F mask .
[0054] As a preferred solution of a method for detecting the wearing of safety goggles for electric power workers based on instance segmentation according to the present invention, wherein: the specific steps of S33 are as follows:
[0055] S331: Use the cross-attention mechanism to perform attention weight operations on P 1 、P 2 、P 3 and P 4 to obtain a safety goggle query matrix, including the first safety goggle query matrix Q 1 、the second safety goggle query matrix Q 2 、the third safety goggle query matrix Q 3 、the fourth safety goggle query matrix Q 4 ;
[0056] S332: Use the multi-layer perceptron MLP to perform feature enhancement on Q 1 、Q 2 、Q 3 and Q 4 to generate a safety goggle prompt embedding matrix, including the first safety goggle prompt embedding matrix E 1 、the second safety goggle prompt embedding matrix E 2 、the third safety goggle prompt embedding matrix E 3 and the fourth safety goggle prompt embedding matrix E 4 , which is used to guide the subsequent mask generation process. The formula is:
[0057] E i =Φ mlp (Q i ), 1≤i≤4
[0058] where, Q i represents the safety goggle query matrix; Φ mlp represents the multi-layer perceptron MLP, and MLP consists of multiple fully connected layers and ReLU activation functions; E i represents the safety goggle prompt embedding matrix, where E 1 is the first safety goggle prompt embedding matrix obtained by performing a multi-layer perceptron operation on Q 1 , E 2 is the second safety goggle prompt embedding matrix obtained by performing a multi-layer perceptron operation on Q 2 , E 3 is the third safety goggle prompt embedding matrix obtained by performing a multi-layer perceptron operation on Q 3 , E 4 is the fourth safety goggle prompt embedding matrix obtained by performing a multi-layer perceptron operation on Q 4 ;
[0059] S333: Combine E 1 、E2 , E 3 and E 4 generate the goggle prompt embedding feature matrix through the sine transform sin, including the first goggle prompt embedding feature matrix T 1 , the second goggle prompt embedding feature matrix T 2 , the third goggle prompt embedding feature matrix T 3 , the fourth goggle prompt embedding feature matrix T 4 . The above goggle prompt embedding feature matrices are used to provide the position information and shape feature prompts of the goggles. The calculation formula is:
[0060] T i = E i + sin(E i ), 1 ≤ i ≤ 4
[0061] where E i represents the goggle prompt embedding matrix; sin(E i ) represents the sine transform of E i ; T i represents the goggle prompt embedding feature matrix, where T 1 is the first goggle prompt embedding feature matrix obtained by adding E 1 and sin(E 1 ), T 2 is the second goggle prompt embedding feature matrix obtained by adding E 2 and sin(E 2 ), T 3 is the third goggle prompt embedding feature matrix obtained by adding E 3 and sin(E 3 ), T 4 is the fourth goggle prompt embedding feature matrix obtained by adding E 4 and sin(E 4 );
[0062] The specific steps of the S331 are as follows:
[0063] S3311: Preset Q 0 as the goggle query matrix that has been initialized to zero. The dimension of Q 0 is (N q , d), where N q represents the number of query features, and d represents the query feature dimension. Calculate the attention weights between P 1 and Q 0 using the cross-attention mechanism, and obtain the attention weight matrix A 1 by Softmax normalization. Multiply A 1 with Q 0Add them to generate the first goggle query matrix Q 1 ;
[0064] S3312: Calculate the attention weights between P 2 and Q 1 using the cross-attention mechanism, and obtain the attention weight matrix A 2 by Softmax normalization. Add A 2 and Q 1 to generate the second goggle query matrix Q 2 ;
[0065] S3313: Calculate the attention weights between P 3 and Q 2 using the cross-attention mechanism, and obtain the attention weight matrix A 3 by Softmax normalization. Add A 3 and Q 2 to generate the third goggle query matrix Q 3 ;
[0066] S3314: Calculate the attention weights between P 4 and Q 3 using the cross-attention mechanism, and obtain the attention weight matrix A 4 by Softmax normalization. Add A 4 and Q 3 to generate the fourth goggle query matrix Q 4 , and the formula is:
[0067]
[0068] Q i =A i +Q i-1 , 1 ≤ i ≤ 4
[0069] where A i represents the attention weight matrix calculated through the cross-attention mechanism; Q i represents the goggle query matrix; the dimensions of A i and Q i are both (N q , d), N q represents the number of query features; d represents the query feature dimension; P i is the feature map with the dimension of (H × W, d); Φ attention (P i , Q i-1 ) represents the cross-attention function, which is used to calculate the information interaction between the goggle query matrix Q i-1 and the feature map P i ; WQ , W K and W V are linear transformation matrices that project Q i-1 and P i onto the same eigen-dimension d; is an N q -row and H×W-column matrix representing attention weights; Softmax represents an activation function that ensures the weights are between 0 and 1 and sum to 1, normalizing the attention weight distribution; by adding A i and Q i-1 to obtain Q i , which can accumulate the information obtained from the feature map P i and gradually learn the key features of the goggles.
[0070] As a preferred solution of the method for detecting the wearing of electrician goggles based on instance segmentation according to the present invention, wherein: the specific steps of S34 are as follows;
[0071] S341: Perturb GT and enter the forward Forword stage to gradually add Gaussian noise to the GT samples to generate NBB. The specific formula for the noise perturbation process is as follows:
[0072]
[0073] where q(x t |x t-1 ) represents the noise perturbation process, x 0 represents the real image without noise, t represents the t-th addition of Gaussian noise, T represents the total number of times of adding noise, represents the Gaussian distribution, β t represents the variance of the Gaussian distribution, which is used to control the magnitude of the added noise, I represents the covariance matrix of the noise, which is an identity matrix, and its size is the same as the dimension of the input data x t or x t-1 ;
[0074] S342: Sample a Gaussian independent noise variable ∈~N(0, I) from x 0 to obtain a sample of x t . The specific formula is as follows:
[0075]
[0076] where, is the weight coefficient, representing the cumulative value of the variance at each time step; ∈ represents the Gaussian independent noise variable, which is the random noise introduced in each diffusion process. As t increases, x t gets closer and closer to pure noise. When T→∞, x Tis pure Gaussian noise, and the noise bounding box of the goggles NBB ∈ {x 1 , x 2 , …, x T};
[0077] S343: Input NBB together with T 1 , T 2 , T 3 , T 4 into the reverse phase for noise reduction, which is used to reverse the noise addition process and sample from p(x t-1 |x t ). The formula for the reverse process is as follows:
[0078]
[0079] During the noise reduction process, represents the Gaussian distribution; μ(x t , t) represents the mean, and ∑(x t , t) is the covariance. The mean and covariance are optimized and learned by minimizing the loss function and are parameterized by the neural network; p(x t-1 |x t ) represents inferring the state x t at the previous moment based on the current state x t-1 . The model restores the image features before noise addition by learning the parameters, and the bounding box features BBF after removing the noise are obtained in this process;
[0080] S344: Send BBF into the fully connected layer FC. After linear transformation and ReLU activation function operations, a noise filter NF related to the goggles instance is generated. The process of generating NF is as follows:
[0081] NF = η(f(x t , t))
[0082] where η represents the fully connected layer, which is used to learn the mapping relationship between the features after denoising and the filter, and f(x t , t) are the bounding box features after removing the noise.
[0083] As a preferred solution of the method for detecting whether a power worker wears goggles based on instance segmentation according to the present invention, wherein: the specific steps of S35 are as follows:
[0084] S351: Use NF to perform a convolution operation on the F mask convolution, extract the shape, boundary, and other key features of a specific goggles instance, and obtain a feature map S 1 with bounding box information, S 1Contains the position and size information of the bounding box. Then, a 1×1 Conv1×1 convolution operation is used to perform convolution on S 1 to adjust the number of channels of the feature map, thereby obtaining S 2 ;
[0085] S352: Apply the Sigmoid activation function operation to S 2 to constrain the output value of S 2 between [0, 1]. Specifically, set a pixel determination threshold H 1 , and binarize the output value of S 2 . If the output value is less than H 1 , it indicates that the pixel belongs to the background, and the corresponding output value is 0; if the output value is greater than H 1 , it indicates that the pixel belongs to the goggles instance, and the corresponding output value is 1. Finally, a power worker goggles instance mask feature map BM composed of 0s and 1s is obtained. In BM, 0 indicates that the pixel point belongs to the background area, and 1 indicates that the pixel point belongs to the goggles area;
[0086] S353: Set a goggles wearing detection threshold H 2 . In the bounding box area of the goggles instance, if the goggles area output in BM is greater than H 2 , it is judged that the goggles are worn. If the goggles area output in BM is less than H 2 , it is judged that the goggles are not worn, and the wearing detection result of the power worker goggles can be obtained.
[0087] As a preferred solution of a power worker goggles wearing detection method based on instance segmentation according to the present invention, wherein: The specific steps of S4 are as follows:
[0088] S41: According to D 2 , train and verify the EGSNet network constructed for S3. First, initialize all neural network parameters and related hyperparameters;
[0089] S42: After initializing the parameters, divide the training set in D 2 into multiple batches according to the batch size, input them batch by batch for training, and obtain the training loss value loss of each batch. After all batches of the training set are trained, divide the validation set in D 2 according to the batch size and input it for verification to obtain the corresponding batch loss value batch_loss;
[0090] S43: During training and validation, the algorithm learns according to the corresponding loss value situation and adjusts the parameters. The training process conducts multiple rounds of training according to the preset number of training epochs. When the loss value of the algorithm tends to converge, the training of the instance segmentation algorithm for detecting the wearing of electrician safety glasses ends;
[0091] S44: After training and validation are completed, apply D 2 to the preprocessed test set, input the test data into the instance algorithm for detecting the wearing of electrician safety glasses with the optimal network parameters after training, conduct the detection test on the wearing situation of electrician safety glasses, and use accuracy and recall rate as the verification indicators of the algorithm to verify the network effect;
[0092] The specific steps of S5 are as follows:
[0093] S51: Use the trained EGSNet network as one of the modules of the electric power construction site monitoring system. The system is connected to the cameras at the electric power construction site to monitor the wearing situation of electrician safety glasses in real time;
[0094] S52: If there is a situation of not wearing or wearing improperly, the system will give an alarm, and at the same time, the relevant images will be recorded in the system for relevant personnel to take corresponding handling measures.
[0095] Compared with the prior art:
[0096] 1. Through the noise perturbation and reverse denoising mechanism, the present invention dynamically generates noise filters related to each safety glasses instance, effectively captures the boundary features of the safety glasses, and realizes more accurate boundary processing; the mask feature map provides global context information, further enhancing the robustness of segmentation in complex backgrounds, significantly reducing the missed detection rate and false detection rate, especially performing excellently when the safety glasses are similar in texture or color to the background;
[0097] 2. The present invention combines the safety glasses prompt embedding feature matrix to dynamically adapt to the positioning and shape changes of the safety glasses, generates features for each instance, and can accurately distinguish multiple target instances to avoid mask overlap; at the same time, the multi-scale feature fusion of the safety glasses retains the shallow detail information, improves the perception and segmentation ability of small targets, and adapts to the segmentation requirements of diverse scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Figure 1 is a schematic flow chart of the present invention;
[0099] Figure 2 is a schematic diagram of the EGSNet network structure of the present invention;
[0100] Figure 3 is a schematic diagram of the multi-scale feature map generation module of the present invention;
[0101] Figure 4 Schematic diagram of the goggle wearing detection feature extraction sub-module of the present invention;
[0102] Figure 5 Schematic diagram of the F-Agg network structure of the present invention;
[0103] Figure 6 Schematic diagram of the multi-scale feature fusion module of the present invention;
[0104] Figure 7 Schematic diagram of the Q2P-Cross module of the present invention;
[0105] Figure 8 Schematic diagram of the B2F-Diffu module of the present invention;
[0106] Figure 9 Schematic diagram of the BMD module of the present invention. Specific implementation manners
[0107] To make the objectives, technical solutions and advantages of the present invention clearer, the following will further describe the implementation manners of the present invention in detail with reference to the accompanying drawings.
[0108] The present invention provides a method for detecting the wearing of goggles by power workers based on instance segmentation. Please refer to Figures 1-9 , and the specific steps are as follows:
[0109] S1: Construct a dataset D of original images of power workers wearing goggles for operation 1 ;
[0110] The specific steps of S1 are as follows:
[0111] S11: Install devices including but not limited to cloth control balls and camera shooting devices at the power grid construction operation sites, and collect images of power workers operating under different lighting conditions, weather conditions, shooting angles, and different postures of power workers;
[0112] S12: Among all the collected images of power workers operating, use the images of power workers wearing goggles correctly as positive samples;
[0113] S13: Use the images of not wearing goggles, wearing goggles incorrectly, and wearing other glasses as negative samples, and keep the number of positive and negative samples relatively balanced, and finally form a dataset D of original images of power workers wearing goggles for operation 1 ;
[0114] S2: Preprocess D 1 to obtain a new dataset D of images of power workers wearing goggles for operation 2 ;
[0115] The specific steps of S2 are as follows:
[0116] S21: For the collected D 1 perform data cleaning operations to eliminate fuzzy, low-quality, and duplicate samples;
[0117] S22: Construct ground truth label GT (Ground Truth) samples. The GT samples contain bounding boxes, instance masks, and instance identification annotation data for each instance in the image. Specifically, use an annotation tool to annotate D 1 and annotate the bounding box for the area where the power workers wear safety goggles in D 1 . The attributes of the bounding box include its position information and category information. The category information includes two types: wearing safety goggles and not wearing safety goggles. In addition, use an annotation tool to draw an instance mask for the operation images of power workers wearing safety goggles in D 1 . There is no need to draw an instance mask for the operation images of power workers not wearing safety goggles in D 1 ;
[0118] S23: To enhance data diversity and improve model robustness, data augmentation techniques will be adopted to perform random rotation, cropping, flipping, and lighting adjustment operations on the images in D 1 to obtain an augmented dataset;
[0119] S24: Split the augmented dataset and divide it into a training set, a validation set, and a test set according to the ratio of 8:1:1 to obtain a new dataset D 2 of operation images of power workers wearing safety goggles;
[0120] Among them: In the experiments of the present invention, the existing LabelMe software is used as the annotation tool, and this software is not implemented in this patent; Data cleaning: Manually eliminate fuzzy, low-quality, and duplicate image samples to ensure the quality of the dataset; Instance mask: It is a pixel-level binary annotation matrix (0 / 1) used to accurately identify the contour area of each safety goggle wearing instance in the image; Random rotation: With the center of the image as the origin, randomly generate a rotation angle from -15° to +15°, and use bilinear interpolation to fill the blank area; Cropping: Randomly generate a cropping ratio within the range of 80%-95% of the image, and retain the area where the power workers wear safety goggles in the image; Flipping: Horizontally flip the image; Lighting adjustment: Adjust the RGB three-channel values through a linear transformation of R±20, G±15, B±10, and finally generate an enhanced image;
[0121] S3: Design an EGSNet network and use this network to process any operation image of a power worker in D 2 to obtain the wearing detection result of the power worker's safety goggles;
[0122] Design an Electric Goggles Segmentation Network (EGSNet) to detect whether electric workers are wearing goggles from electric worker operation images. This network includes the following modules: a multi-scale feature map generation module, a multi-scale feature map fusion module, a Q2P-Cross module, a B2F-Diffu module, and a BMD module. The structure of the EGSNet network is as Figure 2 shown;
[0123] The specific steps of S3 are as follows:
[0124] S31: Construct a multi-scale feature map generation module to generate multi-scale feature maps with enhanced features. The structure of the multi-scale feature map generation module is as Figure 3 shown;
[0125] The specific steps of S31 are as follows:
[0126] S311: Perform a pooling operation on an image Y with a spatial resolution of H×W and a channel number of C 0 ∈D 2 to obtain the first goggles shallow feature map Y with a spatial resolution of H / 2×W / 2 and a channel number of C 1 , and then through two Conv1×1 convolution operations, obtain the second goggles shallow feature map Y with a spatial resolution of H / 2×W / 2 and a channel number of 2C 2 ;
[0127] S312: Input Y 2 into the goggles wearing detection feature extraction sub-module for feature extraction to obtain multi-scale feature maps, including the first layer feature map F 1 , the second layer feature map F 2 , the third layer feature map F 3 and the fourth layer feature map F 4 ; The specific structure of the goggles wearing detection feature extraction sub-module is as Figure 4 shown;
[0128] The specific steps of S312 are as follows:
[0129] S3121: Input Y 2 into the first Conv_Layer1 of the goggles wearing detection feature extraction sub-module, and for Y 2Perform Conv5×5 convolution operation using a 5×5 convolutional kernel to capture local features, use batch normalization (BN) operation to reduce the vanishing gradient problem, use ReLU activation function for operation to introduce non-linearity, enhance the model's expressive ability, and achieve preliminary feature extraction to obtain the first-layer feature map F of goggle wearing detection 1 ;
[0130] S3122: Input F 1 to the second layer Conv_Layer2 of the goggle wearing detection feature extraction sub-module, and perform Conv1×1 convolution operation on F 1 using a 1×1 convolutional kernel to adjust the number of channels and change the depth of the feature map, use BN operation and ReLU function operation to reduce the vanishing gradient and introduce non-linearity, then perform Conv3×3 convolution operation using a 3×3 convolutional kernel to capture local information in the image, perform Conv1×1 convolution operation using a 1×1 convolutional kernel to restore the number of channels, and use BN operation and ReLU function operation on the feature map after restoring the channels to obtain the second-layer feature map F of goggle wearing detection 2 ;
[0131] S3123: After performing residual addition on F 1 and F 2 , input it to the third layer Conv_Layer3 of the goggle wearing detection feature extraction sub-module, perform Conv1×1 convolution operation using a 1×1 convolutional kernel to adjust the number of channels and change the depth of the feature map, use BN operation and ReLU function operation to reduce the vanishing gradient and introduce non-linearity, then perform Conv3×3 convolution operation using a 3×3 convolutional kernel to capture local information in the image, perform Conv1×1 convolution operation using a 1×1 convolutional kernel to restore the number of channels, and use BN operation and ReLU function operation on the feature map after restoring the channels to obtain the third-layer feature map F of goggle wearing detection 3 ;
[0132] S3124: Input F 2 and F 3After performing residual addition, it is input to the fourth layer Conv_Layer4 of the goggle wearing detection feature extraction sub-module. A 1×1 convolutional kernel is used for Conv1×1 convolution operation to adjust the number of channels and change the depth of the feature map. BN operation and ReLU function operation are adopted to reduce gradient vanishing and introduce non-linearity. Then, a 3×3 convolutional kernel is used for Conv3×3 convolution operation to capture local information in the image, and a 1×1 convolutional kernel is used for Conv1×1 convolution operation to restore the number of channels. BN operation and ReLU function operation are performed on the feature map after restoring the channels to obtain the fourth layer feature map F of goggle wearing detection. 4 ;
[0133] The advantage of the goggle wearing detection feature extraction sub-module is that as the number of layers progresses, the number of channels and the depth also increase, and it can extract more and more complex and abstract features. Therefore, the finally obtained feature maps F 1 、F 2 、F 3 and F 4 have the same scale complementarity, thus forming a series of multi-scale feature maps;
[0134] S313: Design an F-Agg network, and use the propagation method from deep semantic features to shallow features to enhance the features of F 1 、F 2 、F 3 and F 4 to obtain multi-scale feature maps with enhanced goggle features, including the first layer enhanced feature map P 1 、the second layer enhanced feature map P 2 、the third layer enhanced feature map P 3 and the fourth layer enhanced feature map P 4 ;
[0135] The specific steps of the above S313 are as follows:
[0136] S3131: Take F 4 as the input. To unify the number of channels, use a 1×1 convolutional kernel to perform Conv1×1 convolution operation on F 4 to obtain the feature map F' 4 . Perform Conv3×3 convolution operation on F' 4 using a 3×3 convolutional kernel to enhance the spatial information, and use the ReLU activation function operation to introduce non-linearity to obtain the fourth layer enhanced feature map P of goggle wearing detection 4 ;
[0137] S3132: Perform Conv1×1 convolution operation on F 3 using a 1×1 convolutional kernel to obtain the feature map F'3 , add the feature map F' 4 , F' 3 and P 4 to obtain M 3 ; perform a Conv3×3 convolution operation on M 3 using a 3×3 convolutional kernel to obtain the third enhanced feature map P for goggle wearing detection 3 ;
[0138] S3133: Perform a Conv1×1 convolution operation on F 2 using a 1×1 convolutional kernel to obtain the feature map F' 2 , add the feature map M 3 , F' 2 and P 3 to obtain M 2 ; perform a Conv3×3 convolution operation on M 2 using a 3×3 convolutional kernel to obtain the second enhanced feature map P for goggle wearing detection 2 ;
[0139] S3134: Use F 1 as the input, and after Conv1×1 convolution, obtain the feature map F' 1 , add the feature map M 2 , F' 1 and P 2 to obtain M 1 ; perform a convolution using the Conv3×3 operation on M 1 to obtain the first enhanced feature map P for goggle wearing detection 1 ;
[0140] S3135: Through convolution operations and the fusion of feature maps, the multi-scale feature maps required for goggle wearing detection are enhanced, and finally the multi-scale feature maps P 1 , P 2 , P 3 and P 4 ;
[0141] In the F-Agg network, the definition of enhancing the goggle features for F 1 , F 2 , F 3 and F 4 is as follows:
[0142] F′ i = Φ Conv1×1 (F i ), 1 ≤ i ≤ 4
[0143] P j = Φ Conv3×3 (Mj ), 1 ≤ j ≤ 4
[0144] M k = M k+1 + P k+1 + F' k , 1 ≤ k ≤ 3
[0145] Wherein, F i is the input multi-scale feature map; Φ Conv1×1 represents the Conv1×1 convolution operation, which is used to adjust the number of channels through convolution; F' i represents the feature map after adjusting the channels, where F' 1 is the feature map after adjusting the channels obtained by performing the Conv1×1 convolution operation on F 1 , F' 2 is the feature map after adjusting the channels obtained by performing the Conv1×1 convolution operation on F 2 , F' 3 is the feature map after adjusting the channels obtained by performing the Conv1×1 convolution operation on F 3 , F' 4 is the feature map after adjusting the channels obtained by performing the Conv1×1 convolution operation on F 4 , F' Conv3×3 represents the Conv3×3 convolution operation, which is used to enhance the spatial information, M j represents the feature map after the fusion of goggles features, wherein, M 4 is the deepest feature map, that is, F' 4 , M 3 is the feature map obtained by fusing and adding the features of M 4 , P 4 and F' 3 , M 2 is the feature map obtained by fusing and adding the features of M 3 , P 3 and F' 2 , M 1 is the feature map obtained by fusing and adding the features of M 2 , P 2 and F' 1 ; P j is the multi-scale feature map of goggles feature enhancement, P 1 is the first layer of enhanced feature map obtained by performing the Conv3×3 convolution operation on M 1 , P 2 is the second layer of enhanced feature map obtained by performing the Conv3×3 convolution operation on M 2 , P 3 is the third layer of enhanced feature map obtained by performing the Conv3×3 convolution operation on M 3 , P4 It is the fourth enhanced feature map obtained after performing a Conv3×3 convolution operation on M 4 ;
[0146] S32: Apply the multi-scale feature map fusion module to fuse different scales of P 1 , P 2 , P 3 and P 4 to output a mask feature map F mask ; The structure of the multi-scale feature map fusion module is as shown in Figure 6 ;
[0147] The specific steps of the above-mentioned S32 are as follows:
[0148] S321: Perform a Fusion operation on P 1 , P 2 , P 3 and P 4 to obtain global semantic information and fine local detail information, and get the fused fifth enhanced feature map P 5 ;
[0149] Among them, the Fusion operation: First, perform bilinear interpolation upsampling on the multi-scale feature maps P 1 , P 2 , P 3 and P 4 to unify the resolution, then calculate the hierarchical weight coefficients through the channel attention mechanism, and finally perform weighted summation to generate the fifth enhanced feature map P 5 ;
[0150] S322: After performing a Conv1×1 convolution operation on P 5 using a convolution kernel of size 1×1, then perform batch normalization BN operation and ReLU function operation to obtain the mask feature map F mask ;
[0151] S33: Construct a Q2P-Cross module, and use the cross-attention mechanism and multi-layer perceptron MLP to perform feature fusion and enhancement processing on P 1 , P 2 , P 3 and P 4 to generate a goggle prompt embedding feature matrix;
[0152] The specific steps of the above-mentioned S33 are as follows:
[0153] S331: Use the cross-attention mechanism for P 1 , P 2 , P 3 and P 4Perform attention weight calculation to obtain the goggles query matrix, including the first goggles query matrix Q 1 , the second goggles query matrix Q 2 , the third goggles query matrix Q 3 , the fourth goggles query matrix Q 4 ;
[0154] The specific steps of S331 are as follows:
[0155] S3311: Preset Q 0 as the goggles query matrix that has been initialized to zero. The dimension of Q 0 is (N q , d), where N q represents the number of query features, and d represents the query feature dimension. Use the cross-attention mechanism to calculate the attention weight between P 1 and Q 0 . Obtain the attention weight matrix A 1 by using Softmax normalization. Add A 1 to Q 0 to generate the first goggles query matrix Q 1 ;
[0156] S3312: Use the cross-attention mechanism to calculate the attention weight between P 2 and Q 1 . Obtain the attention weight matrix A 2 by using Softmax normalization. Add A 2 to Q 1 to generate the second goggles query matrix Q 2 ;
[0157] S3313: Use the cross-attention mechanism to calculate the attention weight between P 3 and Q 2 . Obtain the attention weight matrix A 3 by using Softmax normalization. Add A 3 to Q 2 to generate the third goggles query matrix Q 3 ;
[0158] S3314: Use the cross-attention mechanism to calculate the attention weight between P 4 and Q 3 . Obtain the attention weight matrix A 4 by using Softmax normalization. Add A 4 to Q 3 to generate the fourth goggles query matrix Q 4 , and the formula is:
[0159]
[0160] Q i = A i + Q i-1 , 1 ≤ i ≤ 4
[0161] Among them, A i represents the attention weight matrix calculated through the cross-attention mechanism; Q i represents the goggle query matrix; A i and Q i both have dimensions of (N q , d), N q represents the number of features of the query; d represents the feature dimension of the query; P i is the feature map, with dimensions of (H × W, d); Φ attention (P i , Q i-1 ) represents the cross-attention function, which is used to calculate the information interaction between the goggle query matrix Q i-1 and the feature map P i ; W Q , W K and W V are linear transformation matrices that project Q i-1 and P i onto the same feature dimension d; is an N q row, H × W column matrix, representing the attention weights; Softmax represents the activation function, ensuring that the weights are between 0 and 1 and the sum is 1, normalizing the attention weight distribution; by adding A i and Q i-1 to obtain Q i , it can accumulate the information obtained from the feature map P i , enabling it to gradually learn the key features of the goggles;
[0162] Among them, the goggle query matrix: is a set of learnable vectors used to detect and identify goggles. Initially set to zero, it extracts information from the feature map through the cross-attention mechanism and is gradually updated to enhance the target perception ability; in each iteration, the query matrix calculates the attention weights with the feature map, focuses on the goggle area, and updates itself through weighted feature information, enabling it to gradually learn the position information and shape features of the goggles. Finally, the goggle query matrix updated after multiple rounds can accurately represent the target;
[0163] Cross-attention mechanism: is a method for calculating the information interaction between two different feature spaces (such as the query matrix and the feature map); it calculates the goggle query matrix Q i and the feature map P iThe similarity between them is used to obtain the attention weight matrix A through Softmax normalization i , and then the weight is used to perform weighted summation on the feature map, thereby injecting the key information of the feature map into the query matrix, achieving feature fusion, enabling the query to focus on the important regions in the feature map, and thus improving the recognition ability and expression effect of the target;
[0164] Calculation of the cross-attention mechanism weight: Calculate the goggles query matrix Q i-1 and the feature map P i The attention weight between them is calculated to extract the key features of the goggles; the specific calculation process is as follows: First, project the query matrix Q i-1 and the feature map P i onto the same feature dimension d through the linear transformation matrices W Q and W K ; then, calculate the dot product of the projected Q i-1 and P i , and divide by for scaling to prevent excessive gradients; finally, normalize the similarity score using Softmax to convert it into a probability distribution, and multiply by P i W V to aggregate features and obtain the attention weight matrix A i ;
[0165] S332: Use the multi-layer perceptron MLP to enhance the features of Q 1 , Q 2 , Q 3 and Q 4 to generate the goggles prompt embedding matrix, including the first goggles prompt embedding matrix E 1 , the second goggles prompt embedding matrix E 2 , the third goggles prompt embedding matrix E 3 and the fourth goggles prompt embedding matrix E 4 , which is used to guide the subsequent mask generation process. The formula is:
[0166] E i = Φ mlp (Q i ), 1≤i≤4
[0167] where Q i represents the goggles query matrix; Φ mlp represents the multi-layer perceptron MLP, which consists of multiple fully connected layers and ReLU activation functions; E i represents the goggles prompt embedding matrix, where E 1 is the first goggles prompt embedding matrix obtained by performing the multi-layer perceptron operation on Q 1 2 is the second goggle prompt embedding matrix obtained after performing a multi-layer perceptron operation on Q 2 ; E 3 is the third goggle prompt embedding matrix obtained after performing a multi-layer perceptron operation on Q 3 ; E 4 is the fourth goggle prompt embedding matrix obtained after performing a multi-layer perceptron operation on Q 4 ;
[0168] S333: Generate a goggle prompt embedding feature matrix by performing a sine transformation sin on E 1 , E 2 , E 3 , and E 4 , including the first goggle prompt embedding feature matrix T 1 , the second goggle prompt embedding feature matrix T 2 , the third goggle prompt embedding feature matrix T 3 , and the fourth goggle prompt embedding feature matrix T 4 . The above goggle prompt embedding feature matrices are used to provide position information and shape feature prompts for the goggles, and the calculation formula is:
[0169] T i = E i + sin(E i ), 1 ≤ i ≤ 4
[0170] where E i represents the goggle prompt embedding matrix; sin(E i ) represents performing a sine transformation on E i ; T i represents the goggle prompt embedding feature matrix, where T 1 is the first goggle prompt embedding feature matrix obtained by adding E 1 and sin(E 1 ), T 2 is the second goggle prompt embedding feature matrix obtained by adding E 2 and sin(E 2 ), T 3 is the third goggle prompt embedding feature matrix obtained by adding E 3 and sin(E 3 ), and T 4 is the fourth goggle prompt embedding feature matrix obtained by adding E 4 and sin(E 4 );
[0171] S34: Construct a B2F-Diffu module to generate a noise filter NF related to each goggle instance;
[0172] In the forward Forword stage, Gaussian noise is added to the GT samples to obtain the goggle noise bounding box NBB (Noise Bounding Box). Then, in the reverse Reverse stage, the NBB is denoised to obtain the bounding box feature BBF (Bounding Box Feature) after removing the noise. After passing through the fully connected layer FC, the noisy filter NF (Noisy Filter) related to the goggle instance is finally obtained. The structure of the B2F-Diffu module is as Figure 8 shown
[0173] The specific steps of S34 are as follows;
[0174] S341: Perturb the GT and enter the forward Forword stage to gradually add Gaussian noise to the GT samples to generate the NBB. The specific formula for the noise perturbation process is as follows:
[0175]
[0176] where q(x t |x t-1 ) represents the noise perturbation process, x 0 represents the real image without noise, t represents the t-th addition of Gaussian noise, T represents the total number of times of adding noise, represents the Gaussian distribution, β t represents the variance of the Gaussian distribution, which is used to control the magnitude of the added noise. I represents the covariance matrix of the noise, which is an identity matrix, and its size is the same as the dimension of the input data x t or x t-1 ;
[0177] S342: Sample a Gaussian independent noise variable ∈~N(0, I) from x 0 to obtain the sample of x t . The specific formula is as follows:
[0178]
[0179] where, is the weight coefficient, which represents the cumulative value of the variance at each time step; ∈ represents the Gaussian independent noise variable, which is the random noise introduced in each step of the diffusion process. As t increases, x t gets closer and closer to pure noise. When T→∞, x T is completely Gaussian noise, and the goggle noise bounding box NBB ∈ {x 1 , x 2 , …, x T};
[0180] S343: Combine the NBB with T1 , T 2 , T 3 , T 4 are all input into the reverse stage (corresponding to Reverse in Figure 8 ) for noise reduction, which is used to reverse the noise addition process and sample from p(x t-1 |x t ). The formula for the reverse process is as follows:
[0181]
[0182] During the noise reduction process, represents the Gaussian distribution; μ(x t , t) represents the mean, and Σ(x t , t) is the covariance. The mean and covariance are optimized and learned by minimizing the loss function and are parameterized by the neural network; p(x t-1 |x t ) represents inferring the state x t at the previous moment based on the current state x t-1 . The model restores the image features before noise addition by learning the parameters, and in this process, the bounding box features BBF after removing noise are obtained;
[0183] S344: Feed the BBF into the fully connected layer FC. After linear transformation and operations of the ReLU activation function, a noise filter NF related to the goggle instance is generated. The process of generating NF is as follows:
[0184] NF = η(f(x t , t))
[0185] where η represents the fully connected layer, which is used to learn the mapping relationship between the features after denoising and the filter, and f(x t , t) are the bounding box features after removing noise;
[0186] S35: Use the BMD module to construct the instance mask feature map of the power worker's goggles and obtain the wearing detection result of the power worker's goggles;
[0187] Perform a series of convolution operations on F mask using NF, and then perform a series of transformations on the convolved F mask using the ψ function to generate the goggle instance mask feature map BM (Binary Mask), as shown in the following formula:
[0188] BM = ψ(F mask ; NF) = MaxPool(BN(ReLU(F mask * NF)))
[0189] Among them, the ψ function includes the ReLU activation function, the batch normalization BN operation, and the max pooling MaxPool operation. After being transformed by the ψ function, the obtained BM can effectively represent the position and shape of each goggle instance in the image. This module is as shown in Figure 9 shown;
[0190] The specific steps of S35 are as follows:
[0191] S351: Use NF to perform a convolution operation on the F mask convolution to extract the shape, boundary, and other key features of a specific goggle instance, obtaining a feature map S 1 with bounding box information. Then, use a 1×1 Conv1×1 convolution operation to adjust the number of channels of the feature map for S 1 , thereby obtaining S 1 ; 2
[0192] S352: Apply the Sigmoid activation function operation to S 2 and constrain the output value of S 2 within [0,1]. Specifically: set a pixel determination threshold H 1 , and binarize the output value of S 2 . If the output value is less than H 1 , it indicates that the pixel belongs to the background, and the corresponding output value is 0; if the output value is greater than H 1 , it indicates that the pixel belongs to the goggle instance, and the corresponding output value is 1. Finally, a mask feature map BM of the power worker's goggles composed of 0 and 1 is obtained. In BM, 0 indicates that the pixel point belongs to the background area, and 1 indicates that the pixel point belongs to the goggle area;
[0193] S353: Set a goggle wearing detection threshold H 2 . In the bounding box area of the goggle instance, if the goggle area output in BM (i.e., the area where the pixel points in BM are 1) is greater than H 2 , it is judged that the goggles are worn. If the goggle area output in BM (i.e., the area where the pixel points in BM are 1) is less than H 2 , it is judged that the goggles are not worn, and the wearing detection result of the power worker's goggles can be obtained;
[0194] S4: Load the training set and validation set in D 2 into the EGSNet network for training to update the network parameters, and then use the test set in D 2 as the input to verify the network effect;
[0195] The specific steps of S4 are as follows:
[0196] S41: According to D 2 (D 2 (including the training set, the validation set and the test set), train and validate the EGSNet network constructed in S3. First, initialize all the neural network parameters and related hyperparameters; such as the number of training epochs, the batch size, the learning rate, the activation function, the loss function, etc.;
[0197] S42: After initializing the parameters, divide D 2 the training set in it into multiple batches according to the batch size, input them batch by batch for training, and obtain the training loss value loss of each batch. After all batches of the training set are trained, divide D 2 the validation set in it according to the batch size and input it for validation to obtain the corresponding batch loss value batch_loss;
[0198] S43: During training and validation, the algorithm will learn according to the corresponding loss value situation and adjust the parameters. The training process is carried out for multiple rounds according to the preset number of training epochs. When the algorithm is trained until the loss value tends to converge, the instance segmentation algorithm for detecting the wearing of electrician goggles ends;
[0199] S44: After training and validation are completed, apply D 2 the preprocessed test set, input the test data into the instance algorithm for detecting the wearing of electrician goggles with the optimal network parameters after training, perform the detection test on the wearing situation of electrician goggles, and use the accuracy rate and the recall rate as the verification indicators of the algorithm to verify the network effect;
[0200] S5: Apply the trained EGSNet network to the goggle detection task and integrate it into the electric power construction site monitoring system;
[0201] The specific steps of S5 are as follows:
[0202] S51: Take the trained EGSNet network as one of the modules of the electric power construction site monitoring system, and its system is connected to the camera at the electric power construction site to monitor the wearing situation of electrician goggles in real time;
[0203] S52: If there is a situation of not wearing or wearing irregularly, the system will give an alarm, and at the same time the relevant images will be recorded in the system for relevant personnel to take corresponding treatment measures.
[0204] The present invention includes but is not limited to the following embodiments:
[0205] In the power operation environment, industrial cameras are set up to collect the operation image data of power workers wearing safety goggles. After obtaining 1200 operation images of power workers wearing safety goggles, the LabelMe annotation tool is used to annotate the images. The annotation content includes the bounding box of the safety goggle wearing area and the wearing status (wearing safety goggles or not wearing safety goggles). After the annotation is completed, data preprocessing techniques are adopted to enhance the images, including image rotation and cropping. When rotating, the OpenCV library is used for image rotation processing, and cropping is performed by slicing the image through the Numpy library to ensure the diversity of the dataset. Finally, 3600 images are obtained. Then, the dataset is divided into a training set, a validation set, and a test set in the ratio of 8:1:1 to ensure that there is enough data in the training set for deep learning training.
[0206] Before starting the training, initialize the parameters and hyperparameters of the algorithm. Set an appropriate training batch size according to the hardware environment. Select the Adam optimizer and set the initial number of training epochs to 200 and the initial learning rate to 0.001. These parameters need to be tuned according to the results of multiple rounds of training until the model reaches the best effect.
[0207] After completing the initialization of the basic parameters, start the model training. The embodiment of the training process starts from the input image Y 0 and, after being processed by a series of modules, finally obtains the detection result of power workers' safety goggle wearing. First, the image Y 0 will be processed by the multi-scale feature map generation module to extract feature maps of different scales. These feature maps are enhanced by the F-Agg network to generate multi-scale feature maps P 1 to P 4 containing detailed information. Then, through the multi-scale feature map fusion module, these feature maps are fused into a mask feature map F mask for capturing global semantic information and local details. Next, the image will pass through the Q2P-Cross module to generate a prompt embedding feature matrix T 1 to T 4 using the cross-attention mechanism. This matrix provides clues for instance localization and shape. Subsequently, the B2F-Diffu module filters the generated noisy bounding boxes, removes Gaussian noise, and restores the true boundary features of the target to obtain the noise filter NF. Finally, the mask feature map F maskThe sum and the noise filter NF are fed into the BMD module to generate the final instance mask feature map BM through a convolution operation; a pixel value of 0 in BM indicates that the pixel at that position belongs to the background, and a pixel value of 1 indicates that the pixel at that position belongs to the goggle instance; by setting an appropriate goggle wearing detection threshold, in the bounding box area of the goggle instance, if the output goggle area in BM (i.e., the area where the pixel value in BM is 1) is greater than this threshold, it is determined that the goggles are worn, and if the output goggle area in BM (i.e., the area where the pixel value in BM is 1) is less than this threshold, it is determined that the goggles are not worn, thus completing the goggle wearing detection task;
[0208] The training of the detection algorithm of the present invention updates the internal parameters of the algorithm through backpropagation of the loss function. After obtaining the optimal network parameters, the EGSNet network training is completed; next, the trained EGSNet network is integrated into the power construction site monitoring system; the system will collect images in real time through the cameras installed at the power construction site and transmit them to the server for processing;
[0209] In the deployment environment, ensure that the server has the necessary software library support and computing resources (such as GPU) to accelerate model inference; after the image is processed, it is fed into the trained network model, and the model uses techniques such as multi-scale feature map fusion and noise filter to accurately detect whether the power workers are wearing goggles; if the system detects an abnormal wearing situation (such as not wearing or wearing irregularly), it will immediately send an alarm signal to remind the staff to handle it.
[0210] Although the present invention has been described above with reference to the embodiments, various improvements can be made to it and its components can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the disclosed embodiments of the present invention can be combined with each other in any way. The exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed in the text, but includes all technical solutions falling within the scope of the claims.
Claims
1. A method for detecting electrician goggles wearing based on instance segmentation, characterized in that: The specific steps are as follows: S1: Construct the original dataset D1 of images of power workers wearing goggles at work; S2: Preprocess D1 to obtain a new dataset D2 of images of power workers wearing goggles at work; S3: Design an EGSNet network and use it to process any power worker working images in D2 to obtain the detection results of power workers wearing goggles; S4: Load the training set and validation set in D2 into the EGSNet network for training to update the network parameters, and then use the test set in D2 as input to verify the network effect; S5: Apply the trained EGSNet network to the goggles detection task and integrate it into the power construction site monitoring system; The specific steps of S3 are as follows: S31: construct a multi-scale feature map generation module to generate a feature-enhanced multi-scale feature map; S32: Use the multi-scale feature map fusion module to fuse P1, P2, P3 and P4 of different scales and output a mask feature map F mask ; S33: Construct the Q2P-Cross module, use the cross attention mechanism and multi-layer perceptron MLP to fuse and enhance the features of P1, P2, P3 and P4, and generate the goggles prompt embedding feature matrix; S34: construct the B2F-Diffu module to generate the noise filter NF associated with each goggles instance; S35: Use the BMD module to construct an instance mask feature map of the electric worker goggles and obtain the wearing detection result of the electric worker goggles.
2. The method for detecting electrician goggles based on instance segmentation according to claim 1, characterized in that: The specific steps of S1 are as follows: S11: Installing surveillance cameras and cameras at the power grid construction site to collect images of power workers working under different lighting, weather, shooting angles, and different postures of power workers; S12: Among all the collected images of power workers working, the images of them wearing goggles correctly are taken as positive samples; S13: The images of not wearing, wearing incorrectly, and wearing other glasses are taken as negative samples. The number of positive and negative samples is relatively balanced, and finally the original dataset D1 of images of power workers wearing goggles at work is formed.
3. The method for detecting electrician goggles wearing based on instance segmentation according to claim 1, characterized in that: The specific steps of S2 are as follows: S21: Perform data cleaning operations on the collected D1 to remove ambiguous, low-quality and repeated samples; S22: Construct a ground truth label GT sample, wherein the GT sample includes a bounding box, an instance mask, and an instance identification annotation data of each instance in the image. Specifically, an annotation tool is used to annotate D1, and a bounding box is annotated for the goggles wearing area of the electrician in D1. The bounding box attributes include its position information and category information, wherein the category information includes two types: goggles worn and goggles not worn. In addition, an instance mask is drawn for the working image of the electrician wearing goggles in D1 using an annotation tool, and an instance mask does not need to be drawn for the working image of the electrician not wearing goggles in D1. S23: In order to enhance data diversity and improve model robustness, data augmentation techniques will be used to randomly rotate, crop, flip, and adjust the lighting of the images in D1 to obtain a data augmented dataset; S24: Segment the data set after data enhancement into a training set, a validation set, and a test set in a ratio of 8:1:1, and obtain a new data set D2 of images of power workers wearing goggles at work.
4. The method for detecting electrician goggles based on instance segmentation according to claim 1, characterized in that: The specific steps of S31 are as follows: S311: Perform a pooling operation on the image Y0∈D2 with a spatial resolution of H×W and a number of channels of C to obtain a first goggle shallow feature map Y1 with a spatial resolution of H / 2×W / 2 and a number of channels of C, and then perform two Conv1×1 convolution operations to obtain a second goggle shallow feature map Y2 with a spatial resolution of H / 2×W / 2 and a number of channels of 2C; S312: Input Y2 to the goggles wearing detection feature extraction submodule for feature extraction to obtain a multi-scale feature map, including a first-layer feature map F1, a second-layer feature map F2, a third-layer feature map F3, and a fourth-layer feature map F4; S313: Design an F-Agg network, and use the propagation method from deep semantic features to shallow features to enhance the features of F1, F2, F3 and F4, and obtain a multi-scale feature map of goggles feature enhancement, including the first layer enhanced feature map P1, the second layer enhanced feature map P2, the third layer enhanced feature map P3 and the fourth layer enhanced feature map P4.
5. The method for detecting electrician goggles wearing based on instance segmentation according to claim 4, characterized in that: The specific steps of S312 are as follows: S3121: Input Y2 to the first layer Conv_Layer1 of the goggles wearing detection feature extraction submodule, perform Conv5×5 convolution operation on Y2 with a convolution kernel of size 5×5 to capture local features, use batch normalization BN operation to reduce the gradient vanishing problem, use ReLU activation function to introduce nonlinearity, enhance the model expression ability, realize preliminary feature extraction, and obtain the first layer feature map F1 of goggles wearing detection; S3122: input F1 to the second layer Conv_Layer2 of the goggles wearing detection feature extraction submodule, perform Conv1×1 convolution operation on F1 with a convolution kernel of size 1×1 to adjust the number of channels, change the depth of the feature map, use BN operation and ReLU function operation to reduce gradient disappearance and introduce nonlinearity, then use a convolution kernel of size 3×3 to perform Conv3×3 convolution operation to capture local information in the image, use a convolution kernel of size 1×1 to perform Conv1×1 convolution operation to restore the number of channels, use BN operation and ReLU function operation on the feature map after the channel is restored, and obtain the second layer feature map F2 of goggles wearing detection; S3123: After residual addition of F1 and F2, input them into the third layer Conv_Layer3 of the goggles wearing detection feature extraction submodule, use a convolution kernel of size 1×1 to perform Conv1×1 convolution operation to adjust the number of channels, change the depth of the feature map, use BN operation and ReLU function operation to reduce gradient disappearance and introduce nonlinearity, and then use a convolution kernel of size 3×3 to perform Conv3×3 convolution operation to capture local information in the image, use a convolution kernel of size 1×1 to perform Conv1×1 convolution operation to restore the number of channels, use BN operation and ReLU function operation on the feature map after channel restoration, and obtain the third layer feature map F3 of goggles wearing detection; S3124: After performing residual addition on F2 and F3, the residuals are input to the fourth layer Conv_Layer4 of the goggles wearing detection feature extraction submodule, and a convolution kernel of size 1×1 is used to perform a Conv1×1 convolution operation to adjust the number of channels, change the depth of the feature map, and use BN operation and ReLU function operation to reduce gradient disappearance and introduce nonlinearity. Then, a convolution kernel of size 3×3 is used to perform a Conv3×3 convolution operation to capture local information in the image, and a convolution kernel of size 1×1 is used to perform a Conv1×1 convolution operation to restore the number of channels. The feature map after the channel restoration is subjected to BN operation and ReLU function operation to obtain the fourth layer feature map F4 of goggles wearing detection; The specific steps of S313 are as follows: S3131: Take F4 as input. In order to unify the number of channels, use a 1×1 convolution kernel to perform a Conv1×1 convolution operation on F4 to obtain a feature map F'4. Use a 3×3 convolution kernel to perform a Conv3×3 convolution operation on F'4 to enhance spatial information. Use the ReLU activation function to introduce nonlinearity, and obtain the fourth-layer enhanced feature map P4 for goggles wearing detection. S3132: Perform a Conv1×1 convolution operation on F3 using a convolution kernel of size 1×1 to obtain a feature map F'3, and add the feature maps F'4, F'3, and P4 to obtain M3; perform a Conv3×3 convolution operation on M3 using a convolution kernel of size 3×3 to obtain the third-layer enhanced feature map P3 for goggles wearing detection; S3133: Perform a Conv1×1 convolution operation on F2 using a convolution kernel of size 1×1 to obtain a feature map F'2, and add the feature maps M3, F'2, and P3 to obtain M2; perform a Conv3×3 convolution operation on M2 using a convolution kernel of size 3×3 to obtain the second-layer enhanced feature map P2 for goggles wearing detection; S3134: Take F1 as input, obtain feature map F'1 after Conv1×1 convolution, add feature maps M2, F'1 and P2 to obtain M1; perform Conv3×3 convolution on M1 to obtain the first layer enhanced feature map P1 for goggles wearing detection; S3135: After convolution operation and fusion of feature maps, the multi-scale feature maps required for goggles wearing detection are enhanced, and finally the multi-scale feature maps P1, P2, P3 and P4 with enhanced goggles features are obtained; In the F-Agg network, the goggle feature enhancement for F1, F2, F3 and F4 is defined as follows: F′ i =Φ Conv1×1 (F i ),1≤i≤4 P j =Φ Conv3×3 (M j ),1≤j≤4 M k =M k+1 +P k+1 +F′ k ,1≤k≤3 Among them, F i is the multi-scale feature map of the input; Φ Conv1×1 Represents the Conv1×1 convolution operation, which is used to adjust the number of channels through convolution; F' i represents the feature map after adjusting the channel, where F'1 is the feature map after adjusting the channel obtained by performing Conv1×1 convolution operation on F1, F'2 is the feature map after adjusting the channel obtained by performing Conv1×1 convolution operation on F2, F'3 is the feature map after adjusting the channel obtained by performing Conv1×1 convolution operation on F3, and F'4 is the feature map after adjusting the channel obtained by performing Conv1×1 convolution operation on F4; Φ Conv3×3 represents the Conv3×3 convolution operation, which is used to enhance spatial information. j represents the feature map after the goggles feature fusion, where M4 is the deepest feature map, i.e., F'4, M3 is the feature map obtained by fusing and adding the features of M4, P4, and F'3, M2 is the feature map obtained by fusing and adding the features of M3, P3, and F'2, and M1 is the feature map obtained by fusing and adding the features of M2, P2, and F'1; P j It is a multi-scale feature map of goggles feature enhancement. P1 is the first layer enhanced feature map obtained after performing Conv3×3 convolution operation on M1, P2 is the second layer enhanced feature map obtained after performing Conv3×3 convolution operation on M2, P3 is the third layer enhanced feature map obtained after performing Conv3×3 convolution operation on M3, and P4 is the fourth layer enhanced feature map obtained after performing Conv3×3 convolution operation on M4.
6. The method for detecting electrician goggles based on instance segmentation according to claim 1, characterized in that: The specific steps of S32 are as follows: S321: Perform a fusion operation on P1, P2, P3 and P4 to obtain global semantic information and fine local detail information, and obtain a fused fifth-layer enhanced feature map P5; S322: After performing a Conv1×1 convolution operation on P5 using a convolution kernel of size 1×1, a batch normalization BN operation and a ReLU function operation are performed to obtain a mask feature map F mask .
7. The method for detecting electrician goggles based on instance segmentation according to claim 1, characterized in that: The specific steps of S33 are as follows: S331: Use a cross attention mechanism to perform attention weight calculation on P1, P2, P3 and P4 to obtain a goggles query matrix, including a first goggles query matrix Q1, a second goggles query matrix Q2, a third goggles query matrix Q3, and a fourth goggles query matrix Q4; S332: Use a multi-layer perceptron MLP to enhance the features of Q1, Q2, Q3 and Q4, and generate a goggles prompt embedding matrix, including a first goggles prompt embedding matrix E1, a second goggles prompt embedding matrix E2, a third goggles prompt embedding matrix E3 and a fourth goggles prompt embedding matrix E4, which are used to guide the subsequent mask generation process. The formula is: From i =Φ mlp (Q i ),1≤i≤4 Among them, Q i represents the goggles query matrix; Φ mlp Represents a multi-layer perceptron MLP, which consists of multiple fully connected layers and ReLU activation functions; E i represents the goggles prompt embedding matrix, where E1 is the first goggles prompt embedding matrix obtained by performing a multi-layer perceptron operation on Q1, E2 is the second goggles prompt embedding matrix obtained by performing a multi-layer perceptron operation on Q2, E3 is the third goggles prompt embedding matrix obtained by performing a multi-layer perceptron operation on Q3, and E4 is the fourth goggles prompt embedding matrix obtained by performing a multi-layer perceptron operation on Q4; S333: E1, E2, E3 and E4 are transformed by sine to generate a goggles prompt embedding feature matrix, including a first goggles prompt embedding feature matrix T1, a second goggles prompt embedding feature matrix T2, a third goggles prompt embedding feature matrix T3, and a fourth goggles prompt embedding feature matrix T4. The goggles prompt embedding feature matrix is used to provide the position information and shape feature prompts of the goggles. The calculation formula is: T i =E i +that(E i ),1≤i≤4 Among them, E i represents the goggles prompt embedding matrix; sin(E i ) indicates that E i Perform a sine transform; T i represents the goggles prompt embedding feature matrix, where T1 is the first goggles prompt embedding feature matrix obtained by adding E1 and sin(E1), T2 is the second goggles prompt embedding feature matrix obtained by adding E2 and sin(E2), T3 is the third goggles prompt embedding feature matrix obtained by adding E3 and sin(E3), and T4 is the fourth goggles prompt embedding feature matrix obtained by adding E4 and sin(E4); The specific steps of S331 are as follows: S3311: Preset Q0 to be the goggles query matrix that has been initialized to zero, and the dimension of Q0 is (N q ,d), where N q represents the number of features of the query, d represents the feature dimension of the query, the cross attention mechanism is used to calculate the attention weight between P1 and Q0, and the attention weight matrix A1 is obtained by Softmax normalization. A1 is added to Q0 to generate the first goggles query matrix Q1; S3312: Use the cross attention mechanism to calculate the attention weight between P2 and Q1, use Softmax normalization to obtain the attention weight matrix A2, add A2 and Q1 to generate the second goggles query matrix Q2; S3313: Use the cross attention mechanism to calculate the attention weight between P3 and Q2, use Softmax normalization to obtain the attention weight matrix A3, add A3 and Q2 to generate the third goggles query matrix Q3; S3314: Use the cross attention mechanism to calculate the attention weight between P4 and Q3, use Softmax normalization to get the attention weight matrix A4, add A4 and Q3 to generate the fourth goggles query matrix Q4, the formula is: Q i =A i +Q i-1 ,1≤i≤4 Among them, A i represents the attention weight matrix calculated by the cross-attention mechanism; Q i represents the goggles query matrix; A i and Q i The dimensions are (N q ,d),N q represents the number of query features; d represents the feature dimension of the query; P i is a feature map with dimension (H×W,d); Φ attention (P i ,Q i-1 ) represents the cross attention function, which is used to calculate the goggles query matrix Q i-1 and feature map P i Information interaction between Q , W K and W V is a linear transformation matrix, Q i-1 and P i Project to the same feature dimension d; YesN q A matrix of H×W rows and columns represents the attention weights; Softmax represents the activation function, which ensures that the weights are between 0 and 1 and the sum is 1, so that the attention weight distribution is normalized; i and Q i-1 Add together to get Q i , can accumulate from the feature map P i The information obtained allows it to gradually learn the key features of the goggles.
8. The method for detecting electrician goggles based on instance segmentation according to claim 1, characterized in that: The specific steps of S34 are as follows: S341: Perturb GT and enter the forward Forword stage to gradually add Gaussian noise to GT samples to generate NBB. The specific formula of the noise perturbation process is as follows: Among them, q(x t |x t-1 ) represents the noise perturbation process, x0 represents the real image without noise, t represents the tth time of adding Gaussian noise, T represents the total number of times noise is added, represents Gaussian distribution, β t represents the variance of the Gaussian distribution, which is used to control the size of the added noise. I represents the covariance matrix of the noise, which is a unit matrix with the same size as the input data x. t or x t-1 The dimensions are consistent; S342: By sampling a Gaussian independent noise variable ∈~N0,I) for x0, we get x t The specific formula is as follows: in, is the weight coefficient, which represents the cumulative value of the variance at each time step; ∈ represents a Gaussian independent noise variable, which is the random noise introduced in each diffusion process. As t increases, x t It is getting closer and closer to pure noise. When T→, x T is a complete Gaussian noise, and the goggles noise bounding box NBB∈{x1,x2,…,x T }; S343: NBB is input into the reverse stage together with T1, T2, T3, and T4 for noise reduction, which is used to reverse the noise addition process and obtain the value from p(x t-1 |x t ) sampling, the reverse process formula is as follows: During the noise reduction process, represents Gaussian distribution; μ(x t ,t) represents the mean, Σ(x t ,t) is the covariance, the mean and covariance are optimized by minimizing the loss function and parameterized by the neural network; p(x t-1 |x t ) means that according to the current state x t Estimate the state x at the previous moment t-1 ,The model restores the image features before noise addition by learning parameters, and obtains the bounding box feature BBF after noise removal in this process; S344: Send BBF to the fully connected layer FC, and after linear transformation and ReLU activation function operation, generate a noise filter NF related to the goggles instance. The process of generating NF is as follows: NF=η(f(x t ,t)) Among them, η represents the fully connected layer, which is used to learn the mapping relationship between the denoised features and the filter, f(x t ,t) is the bounding box feature after removing noise.
9. The method for detecting electrician goggles based on instance segmentation according to claim 1, characterized in that: The specific steps of S35 are as follows: S351: Using NF to F mask Convolution is performed to extract the shape, boundary and other key features of the specific goggles instance, and obtain a feature map S1 with bounding box information. S1 contains the position and size information of the bounding box. Then, a Conv1×1 convolution operation with a size of 1×1 is used to convolve S1 to adjust the number of channels of the feature map, thereby obtaining S2. S352: Use the Sigmoid activation function to operate S2, and constrain the output value of S2 to be between [0,1]. Specifically: set a pixel judgment threshold H1, and binarize the output value of S2. If the output value is less than H1, it means that the pixel belongs to the background, and the corresponding output value is 0; if the output value is greater than H1, it means that the pixel belongs to the goggles instance, and the corresponding output value is 1. Finally, a mask feature map BM of the electrician goggles instance composed of 0 and 1 is obtained. In BM, 0 indicates that the pixel point belongs to the background area, and 1 indicates that the pixel point belongs to the goggles area; S353: Set a goggles wearing detection threshold H2. In the bounding box area of the goggles instance, if the goggles area output in the BM is larger than H2, it is judged that the goggles are worn. If the goggles area output in the BM is smaller than H2, it is judged that the goggles are not worn. The goggles wearing detection result of the electrician can be obtained.
10. The method for detecting electrician goggles based on instance segmentation according to claim 1, characterized in that: The specific steps of S4 are as follows: S41: According to D2, the EGSNet network constructed in S3 is trained and verified. First, all neural network parameters and related hyperparameters are initialized; S42: After initializing the parameters, the training set in D2 is divided into multiple batches according to the batch size, and training is performed after inputting the batches to obtain the training loss value loss of each batch. After all batches of the training set are trained, the validation set in D2 is also divided according to the batch size and inputted to obtain the corresponding batch loss value batch_loss; S43: During training and verification, the algorithm will learn and adjust parameters according to the corresponding loss value. The training process will perform multiple rounds of training according to the preset training rounds. When the algorithm is trained until the loss value tends to converge, the instance segmentation algorithm training of the electric worker goggles wearing detection is completed; S44: After the training and verification are completed, the test set preprocessed by D2 is applied, and the test data is input into the example algorithm of the detection of the wearing of goggles for power workers with the optimal network parameters after the training is completed, and the detection test of the wearing of goggles for power workers is carried out, and the accuracy and recall rate are used as the verification indicators of the algorithm to verify the network effect; The specific steps of S5 are as follows: S51: The trained EGSNet network is used as one of the modules of the power construction site monitoring system. The system is connected to the cameras at the power construction site to monitor the wearing of goggles by power workers in real time. S52: If the mask is not worn or is worn improperly, the system will give an alarm and the relevant images will be recorded in the system so that relevant personnel can take corresponding measures.
Citation Information
Patent Citations
Method and device for monitoring and identifying electric power operating personnel, and electronic equipment
CN117830939A