Dressing identification method and system for electric power operating personnel

By using the YOLOv5 model for human body detection and multi-label and multi-task dress recognition model to determine the dress situation, the traditional inspection efficiency and in real time are solved, efficient and accurate monitoring of the dress of power workers is achieved, and the safety of work is improved.

CN120220178AInactive Publication Date: 2025-06-27DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192737.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional power operators' dress inspection methods are inefficient and prone to omissions, making it difficult to meet the real-time requirements of safety management at the power operation site.

Method used

The pre-trained YOLOv5 model is used for human body detection, positioning the regional location of the power operator, and input the cropped human area image into the training multi-label multi-task dress recognition model to judge the dress status of the operator, and issue an alarm prompt when the specified equipment is not worn.

Benefits of technology

It realizes efficient and accurate identification of the dressing conditions of power workers, can monitor the dressing specifications of the work site personnel in real time, reduces the risk of accidents caused by improper dressing, and improves the safety of power operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220178A_ABST
    Figure CN120220178A_ABST
Patent Text Reader

Abstract

The invention discloses a dressing identification method and system for an electric power worker, and belongs to the technical field of computer vision and artificial intelligence. The method comprises the following steps: performing human body detection on an input to-be-recognized image by using a pre-trained YOLOv5 model, positioning a regional position of an electric power worker, and cutting out an image part only containing a human body region; inputting the cut human body area image into the trained multi-label multi-task dressing identification model; and analyzing the output result of the multi-label multi-task dressing identification model, and judging the dressing condition of the operator. The method has the advantages of high recognition accuracy, high efficiency, high expandability and the like, can be widely applied to a video monitoring system in the power industry, and is beneficial to preventing safety accidents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and more particularly to a method and system for identifying the clothing of power operation personnel. Background Art

[0002] The power industry belongs to an industry with high-risk operations. Operation personnel often face various risk factors such as high voltage, high temperature, and high altitude. Once a safety accident occurs, the consequences are often very serious. Therefore, ensuring the safety of operation personnel is the top priority of the power industry's work.

[0003] Personal protective equipment is an important means to ensure the safety of power operation personnel. For example, safety helmets can prevent the head from being hit or electrocuted, work clothes can prevent electric shock and high-temperature burns, and insulating gloves and insulating shoes can prevent electric shock accidents. However, during actual operations, due to various reasons, operation personnel sometimes forget to wear or wear personal protective equipment in an irregular manner, which greatly increases the risk of safety accidents.

[0004] Therefore, in high-risk operation environments such as the power industry, ensuring that operation personnel wear the correct personal protective equipment (such as safety helmets, work clothes, insulating gloves, and insulating shoes) is crucial for accident prevention.

[0005] However, traditional clothing inspection methods mainly rely on manual labor. The inspection requires a large amount of time and manpower, and it is difficult to meet the real-time requirements of safety management at the power operation site; manual inspection is prone to human factors such as fatigue and negligence, resulting in inaccurate inspection results and unable to detect irregular clothing in a timely manner. In addition, it is difficult to conduct data statistics and analysis for manual inspection, which is not conducive to safety management by power enterprises.

[0006] In order to overcome the limitations of traditional clothing inspection methods and improve the safety of power operations, providing an efficient and accurate automatic clothing recognition solution is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a method and system for identifying the clothing of power operation personnel, which can solve the problems of low efficiency and easy omission in traditional clothing inspection; this method can efficiently and accurately identify whether operation personnel wear the specified personal protective equipment.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] In the first aspect, an embodiment of the present invention provides a method for identifying the clothing of power operation personnel, including the following steps:

[0010] S1. Use the pre-trained YOLOv5 model to perform human detection on the input image to be recognized, locate the regional position of the power operation personnel, and crop out the image part that only contains the human body region;

[0011] S2. Input the cropped human body region image into the trained multi-label multi-task clothing recognition model;

[0012] S3. Analyze the output results of the multi-label multi-task clothing recognition model to judge the clothing situation of the operation personnel.

[0013] Furthermore, it also includes:

[0014] S4. When it is judged that the operation personnel do not wear the specified work clothes, safety helmets, insulating gloves or insulating shoes, give an alarm prompt.

[0015] Furthermore, the multi-label multi-task clothing recognition model includes:

[0016] Backbone network layer, used to extract deep semantic features from the input image;

[0017] Branch module, including multiple task branches. Each task branch shares the parameters of the backbone network layer and respectively predicts whether the operation personnel wear safety helmets, work clothes, insulating gloves on hands and insulating shoes on feet;

[0018] Output module. The output module of each branch contains at least two fully connected layers, used to output the existence or non-existence of the corresponding attribute labels.

[0019] Furthermore, the backbone network layer adopts the ResNet18 architecture and has the following optimizations:

[0020] Replace the activation function ReLU in the BasicBlock module with PRelu with learnable parameter w;

[0021] Integrate a channel attention module in each BasicBlock, used to dynamically allocate weights for different channels.

[0022] Furthermore, the channel attention module includes:

[0023] a) Global pooling module, used to perform global pooling operation on the feature map of each channel, and reduce the two-dimensional feature map to a one-dimensional vector;

[0024] b) Fully connected layer, used to perform a fully connected operation on the vector output by the global pooling module, and learn the weight of each channel through linear transformation and non-linear activation function;

[0025] c) A non - linear activation and normalization module, which is used to normalize the weights output by the fully - connected layer to ensure that the weight values of all channels are between 0 and 1, and the sum of the weights is 1.

[0026] d) An attention application module, which is used to multiply the channel attention coefficient vector obtained by the normalization process with the original feature map element - by - element by channel to achieve adaptive weighting of the features of each channel.

[0027] Furthermore, each task branch in the branch module adopts a two - layer convolution structure, specifically including:

[0028] a) A convolution layer, which is used to extract the spatial hierarchical features of the input feature map. The stride of the first - layer convolution is set to stride = 2, and the stride of the second - layer convolution is set to stride = 1.

[0029] b) A batch normalization layer, which is used to accelerate the training process by normalizing the input of each layer to reduce the internal covariate shift.

[0030] c) A non - linear activation function PRelu, which is used to adaptively learn and correct the parameters of the rectified linear unit according to the existence of a learnable parameter w.

[0031] d) A global pooling layer, which is used to reduce the spatial dimension of the feature map to a one - dimensional feature vector of a fixed length for fusing the feature information of multiple channels of each branch.

[0032] Furthermore, the training process of the multi - label multi - task clothing recognition model includes:

[0033] 1) Collect a dataset containing human body images in the power operation environment, and the human body images are with annotation content; the annotation content includes human target boxes and corresponding human target attributes: four category labels of whether wearing a safety helmet, wearing work clothes, wearing insulating gloves on hands, and wearing insulating shoes on feet.

[0034] 2) Divide the dataset into a training set and a test set according to a preset ratio; the training set is used for model training, and the test set is used for evaluating the model performance.

[0035] 3) Construct a multi - label multi - task clothing recognition model, including: a backbone network layer, a branch module, and an output module.

[0036] 4) Use the training set as the input of the multi - label multi - task clothing recognition model, and use the Adam optimizer for model training; adopt the cross - entropy loss function Cross - Entropy Loss; set training parameters, batch size, initial learning rate, and number of iterations.

[0037] 5) After each complete data iteration, calculate the accuracy of the model on the test set and save the model weights with the highest current accuracy.

[0038] 6) After the training is completed, save the model weight file and the inference program code.

[0039] In a second aspect, an embodiment of the present invention further provides a power operation personnel clothing recognition system, which uses the power operation personnel clothing recognition method described in any one of the above embodiments. This system is deployed in the monitoring system at the power operation site to achieve real-time monitoring of the clothing situation of the operation personnel. The system includes:

[0040] A human detection module that uses a pre-trained YOLOv5 model to perform human detection on the input image to be recognized, locates the regional position of the power operation personnel, and crops out the image part that only contains the human body area.

[0041] An input module for inputting the cropped human body area image into the trained multi-label multi-task clothing recognition model.

[0042] A judgment module that analyzes the output result of the multi-label multi-task clothing recognition model to judge the clothing situation of the operation personnel.

[0043] Furthermore, it further includes:

[0044] An alarm module for sending an alarm prompt when it is judged that the operation personnel do not wear the prescribed work clothes, safety helmets, insulating gloves or insulating shoes.

[0045] It can be seen from the above technical solutions that compared with the prior art, the present invention has the following technical advantages:

[0046] The implementation technology of the present invention covers multiple technical fields such as deep learning, image processing, and computer vision. This method can be widely applied to the video monitoring system in the power industry to achieve real-time monitoring and intelligent analysis of the clothing norms of the personnel at the operation site, which helps to improve the safety of power operations, ensure that the operation personnel strictly abide by the safety operation procedures, and reduce the accident risk caused by improper clothing. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0048] Figure 1 It is a flow chart of the power operation personnel clothing recognition method provided by the present invention.

[0049] Figure 2 This is the deployment flowchart of the implementation process of the clothing recognition for power operation personnel provided by the present invention.

[0050] Figure 3 This is the structure diagram of the multi-label multi-task clothing recognition model provided by the present invention.

[0051] Figure 4 This is the structure diagram of the backbone network layer provided by the present invention.

[0052] Figure 5 This is the structure diagram of each branch module provided by the present invention.

[0053] Figure 6 This is the structure block diagram of the clothing recognition system for power operation personnel provided by the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] Refer to Figure 1 As shown, an embodiment of the present invention discloses a method for clothing recognition of power operation personnel, including the following steps:

[0056] S1. Use the pre-trained YOLOv5 model to perform human detection on the input image to be recognized, locate the regional position of the power operation personnel, and crop out the image part that only contains the human body area.

[0057] Among them, YOLOv5 is an efficient and accurate object detection algorithm that has been fully trained on the COCO dataset and can identify human targets in images. According to the actual application scenario, the parameters of the YOLOv5 model can be fine-tuned, such as adjusting the detection threshold, adjusting the model size, etc., to meet the requirements of different scenarios. The YOLOv5 model will output the bounding box of the human target, accurately locate the regional position of the power operation personnel, and the image part containing the human target can be cropped out for subsequent clothing recognition.

[0058] S2. Input the cropped human body area image into the trained multi-label multi-task clothing recognition model;

[0059] In this step, the cropped human body region image is input into the trained Multi-label Multi-task Recognition Model (MMRM).

[0060] The multi-label multi-task clothing recognition model includes:

[0061] A backbone network layer for extracting deep semantic features from the input image;

[0062] A branch module, including multiple task branches. Each task branch shares the parameters of the backbone network layer and respectively predicts whether the operator wears a safety helmet, work clothes, insulating gloves on the hands, and insulating shoes on the feet;

[0063] An output module. The output module of each branch contains at least two fully connected layers for outputting the presence or absence of the corresponding attribute label.

[0064] The backbone network (Backbone) of the MMRM model extracts deep semantic features from the input image, providing a basis for subsequent clothing recognition. The four branch modules of the MMRM model respectively predict whether the operator wears a safety helmet, work clothes, insulating gloves on the hands, and insulating shoes on the feet. The output result of each branch module is either 0 or 1, where 0 indicates not wearing and 1 indicates wearing.

[0065] S3. Analyze the output results of the multi-label multi-task clothing recognition model to judge the clothing situation of the operator. For example, the judgment result is presented in a visual way, such as marking the clothing situation of the operator on the monitoring screen.

[0066] S4. When it is judged that the operator does not wear the required work clothes, safety helmet, insulating gloves or insulating shoes, an alarm prompt is issued. For example, when it is judged that the operator does not wear the required work clothes, safety helmet, insulating gloves or insulating shoes, an alarm prompt can be issued through pop-up windows, voice, text messages, etc. And the management personnel intervene in a timely manner according to the alarm prompt to ensure the safety of the operator.

[0067] This method can monitor the clothing situation of the operator in real time, discover and correct non-standard clothing behaviors in a timely manner, thus effectively preventing the occurrence of safety accidents and ensuring the safety of the operator.

[0068] In the embodiments of the present invention, the pre-trained YOLOv5 and MMRM models are involved. Therefore, during training, only the MMRM model is trained, and during deployment, the pre-trained YOLOv5 and the trained MMRM model are deployed into the monitoring system. Similarly, during inference, the obtained images are first passed through the pre-trained YOLOv5 for human body part detection, and then the cropped human body part detection images are input into the trained MMRM model for clothing detection.

[0069] Refer to Figure 2 As shown, it is the flowchart from training to deployment, and the specific process is as follows:

[0070] 1) Data collection and processing:

[0071] a. Collect a large dataset containing human body images in the power operation environment, and these data should cover various different operation environments and conditions.

[0072] b. Manually annotate each image, and the annotation content is specific to the target box of the operator's human body and the corresponding human body target attributes: four category labels of whether wearing a safety helmet, wearing work clothes, wearing insulating gloves on the hands, and wearing insulating shoes on the feet.

[0073] c. Divide the data into a training set and a test set according to a ratio of 7:3. The training set is used to train the model, and the test set is used to evaluate the model performance. These data are only used during the model training process. After the model is trained, these data are not needed for subsequent detection, recognition, and testing processes. In the actual deployment environment, only the model weight file and the corresponding inference program code need to be retained to perform the prediction task.

[0074] 2) Human body detection: It is mainly used to detect and locate the area position of the human body in the image, providing an accurate target image area for subsequent tasks such as clothing recognition. In this embodiment, the pre-trained model YOLOv5 is directly adopted. This model is an efficient real-time object detection algorithm that has been fully trained on a large dataset COCO (Common Objects in Context), covering various complex scenes and multiple object categories, including the human body. By adopting YOLOv5, the training results on the large-scale dataset can be directly utilized without training the model from scratch, thus greatly saving time and computing resources. In addition, YOLOv5 supports flexible configuration options and can be fine-tuned according to the requirements of specific application scenarios to further optimize the performance of human body detection.

[0075] 3) Clothing Recognition: The clothing situation of power industry workers includes 4 attribute categories, namely wearing a safety helmet, wearing work clothes, wearing insulating gloves on hands, and wearing insulating shoes on feet. Since the clothing attribute categories are not mutually exclusive and multiple attributes can appear in the same picture, it is impossible to directly use the method of multi-class image recognition. Currently, the commonly used deep learning CNN networks (such as MobileNet, VGG16, ResNet) are mainly used for multi-class image recognition tasks and cannot support multi-label multi-task image recognition.

[0076] To solve this problem, the method of the present invention proposes a multi-label multi-task recognition model (Multi-label Multi-task Recognition Model), abbreviated as MMRM, for clothing multi-label multi-task attribute recognition. The model network structure includes a backbone network, a branch module, and an output layer. As shown in Figure 3 shown, the details are as follows:

[0077] a. Backbone Network Layer: ResNet18 is used as the basic architecture of the Backbone network, aiming to extract deep semantic features from the input image. ResNet18 consists of a series of BasicBlock modules, and these modules build the network by introducing residual connections. The residual connection allows the input data to directly bypass one or more layers and be added to the deeper output. This not only simplifies the information flow path but also ensures that the gradient can be effectively backpropagated even when the network depth increases. This mechanism effectively alleviates the common problems of gradient disappearance and gradient explosion in the training process of deep neural networks. However, the traditional BasicBlock structure lacks a model attention mechanism, which limits its learning ability for complex patterns.

[0078] Therefore, in this embodiment, the BasicBlock of ResNet18 is optimized and enhanced. As shown in Figure 4As shown below. First, replace the original activation function ReLU of BasicBlock with PReLU with learnable parameter w. This improvement allows the activation function to adaptively adjust its slope, thereby enhancing the flexibility and expressiveness of the model. Meanwhile, it helps reduce the risk of overfitting, enabling the model to capture more complex mapping relationships. In addition, a channel attention module is integrated into each BasicBlock. This module dynamically assigns different weights according to the importance of each channel for the current task. In this way, the model can focus more on the features crucial to the task and ignore the less relevant parts, thereby enhancing the overall feature representation ability and decision-making accuracy of the model. The channel attention module mainly consists of key components such as Global Pooling, fully connected layer (FC), non-linear activation function (such as sigmoid), and attention application. The specific structure is described as follows:

[0079] a) Global Pooling: This step is responsible for performing global pooling operations on the feature maps of each channel, aiming to reduce the two-dimensional feature maps to one-dimensional vectors. This vector highly condenses the feature information of all positions within the corresponding channel and is crucial for subsequent feature weight learning. Through global pooling, the feature distribution of each channel can be grasped as a whole, providing a comprehensive feature description for the subsequent steps.

[0080] b) Fully connected layer (FC): It performs a fully connected operation on the vector output by the Global Pooling layer, learning the weights of each channel through linear transformation and non-linear activation functions. These weights essentially reflect the importance degree of different channels for the current task and constitute the core of the attention mechanism. Through the processing of the fully connected layer, the dependence relationship between channels can be captured, providing a basis for subsequent attention allocation.

[0081] c) Non-linear activation and normalization: After the fully connected layer, non-linear activation functions such as sigmoid are usually used to normalize the weights, ensuring that the weight values of all channels are between 0 and 1 and the sum of the weights is 1. This process converts the weights into attention coefficients in the form of probabilities, facilitating precise feature weighting in subsequent network layers.

[0082] d) Attention application: Multiply the obtained channel attention coefficient vector with the original feature map element-wise by channel. This step realizes the adaptive weighting of the features of each channel. The weighted feature map not only retains the original spatial structure information but also the influence of each channel is more precisely regulated. In this way, the network can focus more on the feature channels with high contribution to the task, thereby improving the overall feature representation ability and task performance.

[0083] b. Branch Module: The branch module is a multi-task branch module composed of 4 Branch branches. It shares the parameters of the Backbone backbone network and predicts 4 attribute labels, namely wearing a safety helmet, wearing work clothes, wearing insulating gloves on the hands, and wearing insulating shoes on the feet. Each Branch branch has the same structure. Referring to Figure 5 as shown, it adopts a two-layer convolution structure: Conv - BN - PRelu - Conv - BN - PRelu - Global Pooling.

[0084] a) Conv (Convolution Layer): The convolution operation is used to extract the spatial hierarchical features of the input feature map. In the two-layer convolution structure, the stride of the first convolution is set to stride = 2, that is, the moving step of the filter is 2 pixels. This setting helps to quickly reduce the size of the image, reduce the computational amount, and at the same time retain the key information in the image. The stride of the second convolution is set to stride = 1, that is, the filter moves only 1 pixel each time it moves, which helps to further extract the detailed features of the image. The first convolution layer of each branch further transforms the high-level features obtained from ResNet18 into a feature representation more suitable for this specific task.

[0085] b) BN (Batch Normalization Layer): Batch normalization helps to accelerate the training process by normalizing the input of each layer to reduce the internal covariate shift, thereby stabilizing and accelerating the training of the deep network.

[0086] c) PReLU: This is a non-linear activation function. There is a learnable parameter w in this activation function, which can adaptively learn to correct the parameters of the linear unit, improve the fitting ability of the model, and reduce the risk of overfitting, enabling the model to learn more complex mapping relationships.

[0087] d) Global Pooling (Global Pooling Layer): In a convolutional neural network, the global pooling layer is usually used to reduce the spatial dimension of the feature map and convert it into a fixed-length vector, which is very important for subsequent classification tasks. Global Average Pooling (GAP) calculates the average value for each channel of the feature map and is responsible for performing the global pooling operation on the feature map of each channel to obtain a one-dimensional feature vector, mainly used to fuse the feature information of multiple channels of each branch; this not only reduces the number of parameters but also helps to prevent overfitting.

[0088] c. Output module: The output module of each branch consists of two fully connected layers. The first fully connected layer is usually used as a feature transformation layer, which can capture the complex relationships between different features and prepare the input for the last fully connected layer. The second fully connected layer is directly associated with the classification or regression task, and the number of its neurons usually equals the number of classes of the task or the number of output attributes. Since only the existence of each attribute label needs to be judged, the number of output channels of the last fully connected layer is 2, where 0 indicates the non-existence of the attribute label and 1 indicates the existence of the attribute label.

[0089] Label 0 1 Work clothes Without work clothes Wearing work clothes Safety helmet Without safety helmet Wearing a safety helmet Insulating gloves Without wearing insulating gloves Wearing insulating gloves Insulating shoes Without wearing insulating shoes Wearing insulating shoes

[0090] d. Model training: Use the above-mentioned training set to train and optimize the MMRM model; use the above-mentioned test set to test the MMRM model and evaluate the performance of the model metrics. When training the model, the input images are uniformly scaled to 224×224. To improve the diversity of data and the generalization ability of the model, the training data image enhancement methods adopt various combinations such as random flipping, random cropping, random rotation, and random color transformation (such as adjustment of brightness, contrast, saturation, and hue). In terms of training parameters, the batch size batch_size = 64, the optimization algorithm uses the Adam optimizer, the initial learning rate is set to lr = 0.001, and the cross-entropy loss function (Cross-Entropy Loss) is adopted. The entire dataset is iterated epoch = 200 times; during the training process, after each complete data iteration of the program, the accuracy of the model on the test set is calculated once, and the model weights with the highest current accuracy are saved; in subsequent actual use, only the model weight file and the corresponding inference program code need to be retained to achieve efficient recognition of new images.

[0091] 4) System deployment: Deploy the pre-trained YOLOv5 model and the trained deep learning model to the monitoring system at the power operation site to achieve real-time monitoring of the dressing conditions of the operating personnel. When it is monitored that the operating personnel do not wear the specified personal protective equipment (such as not wearing work clothes, not wearing a safety helmet, not wearing insulating gloves, or not wearing insulating shoes), the system can capture images or videos in real time and issue warning prompts through pop-up windows, voices, etc. to assist the management personnel to intervene in a timely manner.

[0092] This method first uses the YOLOv5 model to perform human detection on the input image, accurately locates the regional position of the power operation personnel, and crops out the image part that only contains the human body area accordingly. Subsequently, a multi-label, multi-task dressing recognition model is constructed to identify whether the operating personnel wear a safety helmet, work clothes, insulating gloves on their hands, and insulating shoes on their feet. Finally, the algorithm system is deployed to the monitoring system to achieve real-time monitoring and intelligent analysis of the dressing specifications of the personnel at the operation site.

[0093] The method for identifying the clothing of power operation personnel provided by the present invention can be widely applied to multiple fields in specific implementation, including but not limited to the following aspects:

[0094] 1) Power industry supervision: At the power operation site, this method can be used to monitor the clothing situation of operation personnel in real time, ensuring that they wear personal protective equipment such as safety helmets, work clothes, insulating gloves and insulating shoes according to regulations, thus effectively preventing the occurrence of safety accidents.

[0095] 2) Work safety management: Not limited to the power industry, this method can also be applied to other industries that require strict work safety management, such as chemical industry, petroleum, mining, etc. By monitoring the clothing situation of operation personnel in real time, behaviors that do not conform to safety regulations can be discovered and corrected in time, improving the overall work safety level.

[0096] 3) Training and assessment: During the training and assessment of power operation personnel, this method can be used as an evaluation tool to check whether trainees wear clothes according to regulations. This helps to improve the safety awareness and operation skills of trainees, ensuring that they can strictly abide by safety regulations in actual work.

[0097] Based on the same inventive concept, an embodiment of the present invention also provides a system for identifying the clothing of power operation personnel, which uses the method for identifying the clothing of power operation personnel as described in the above embodiment. This system is deployed in the monitoring system at the power operation site to realize the real-time monitoring of the clothing situation of operation personnel. Referring to Figure 6 as shown, this system includes:

[0098] A human body detection module, which uses a pre-trained YOLOv5 model to perform human body detection on the input image to be recognized, locates the regional position of power operation personnel, and crops out the image part that only contains the human body area;

[0099] An input module, which is used to input the cropped human body area image into the trained multi-label multi-task clothing recognition model;

[0100] A judgment module, which analyzes the output result of the multi-label multi-task clothing recognition model to judge the clothing situation of operation personnel;

[0101] An alarm module, which is used to give an alarm prompt when it is judged that the operation personnel do not wear the prescribed work clothes, safety helmets, insulating gloves or insulating shoes.

[0102] Among them, the multi-label multi-task recognition model in the input module is used to identify whether the operator wears a safety helmet, work clothes, insulating gloves on hands and insulating shoes on feet, which solves the problem of a single model dealing with multiple non-exclusive labels and improves the flexibility and application scope of model recognition. This multi-label multi-task recognition model simultaneously predicts different attributes by sharing the backbone feature extraction layer and multiple specific task branches, improving the model resource utilization rate; and by adopting an improved network structure and optimization strategy, such as introducing a channel attention mechanism and adjusting the activation function, to improve the model performance.

[0103] This system can monitor the dressing situation of on-site operators in real time, discover and correct behaviors that do not conform to safety regulations in a timely manner. Compared with manual inspection, the system has higher efficiency and accuracy, can quickly identify multiple dressing attributes, and avoid manual omissions. With its real-time, high-efficiency, flexible and easy-to-use features, this power operator dressing recognition system can effectively improve the safety of power operations, reduce accident risks, and has broad application prospects.

[0104] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0105] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identifying the clothing of power workers, characterized in that: The following steps are involved: S1. Use the pre-trained YOLOv5 model to perform human body detection on the input image to be identified, locate the area of ​​the power workers, and crop the image part containing only the human body area; S2, input the cropped human body region image into the trained multi-label multi-task clothing recognition model; S3. Analyze the output results of the multi-label multi-task clothing recognition model to determine the clothing of the operators.

2. The method for identifying the clothing of electric power workers according to claim 1, characterized in that: Also includes: S4. When it is determined that the operator is not wearing the required work clothes, safety helmets, insulating gloves or insulating shoes, an alarm will be issued.

3. The method for identifying the clothing of electric power workers according to claim 1, characterized in that: The multi-label multi-task clothing recognition model includes: Backbone network layer, used to extract deep semantic features from the input image; The branch module includes multiple task branches. Each task branch shares the parameters of the backbone network layer and predicts whether the operator wears a helmet, work clothes, insulating gloves, and insulating shoes. Output module,The output module of each branch contains at least two fully connected layers, which are used to output the existence or non-existence of the corresponding attribute label.

4. The method for identifying the clothing of electric power workers according to claim 3, characterized in that: The backbone network layer adopts the ResNet18 architecture and performs the following optimizations: Replace the activation function ReLU in the BasicBlock module with PRelu with a learnable parameter w; A channel attention module is integrated in each BasicBlock to dynamically assign weights to different channels.

5. The method for identifying the clothing of electric power workers according to claim 4, characterized in that: The channel attention module includes: a) A global pooling module is used to perform a global pooling operation on the feature map of each channel and reduce the dimensionality of the two-dimensional feature map into a one-dimensional vector; b) Fully connected layer, used to fully connect the vector output by the global pooling module and learn the weight of each channel through linear transformation and nonlinear activation function; c) Nonlinear activation and normalization module, which is used to normalize the weights of the fully connected layer output to ensure that the weight values ​​of all channels are between 0 and 1 and the sum of the weights is 1; d) Attention application module, which is used to multiply the channel attention coefficient vector obtained by normalization by the original feature map element by channel to achieve adaptive weighting of each channel feature.

6. The method for identifying the clothing of electric power workers according to claim 3, characterized in that: Each task branch in the branch module adopts a two-layer convolution structure; specifically, it includes: a) Convolutional layer, used to extract the spatial hierarchical features of the input feature map, where the stride of the first convolution layer is set to stride=2, and the stride of the second convolution layer is set to stride=1; b) Batch Normalization layer, which is used to speed up the training process and reduce internal covariate shift by normalizing the input of each layer; c) a nonlinear activation function PRelu, which is used to adaptively learn the parameters of the rectified linear unit based on the existence of a learnable parameter w; d) Global pooling layer, which is used to reduce the spatial dimension of the feature map to a one-dimensional feature vector of fixed length and to fuse the feature information of multiple channels of each branch.

7. The method for identifying the clothing of electric power workers according to claim 1, characterized in that: The training process of the multi-label multi-task clothing recognition model includes: 1) Collect a data set containing human images in an electric power operation environment, wherein the human images are annotated; the annotated content includes a human target frame and corresponding human target attributes: four category labels: whether wearing a safety helmet, wearing work clothes, wearing insulating gloves, and wearing insulating shoes; 2) Divide the data set into a training set and a test set according to a preset ratio; the training set is used for model training, and the test set is used to evaluate model performance; 3) Construction of a multi-label multi-task clothing recognition model, including: backbone network layer, branch module and output module; 4) Use the training set as the input of the multi-label multi-task clothing recognition model and use the Adam optimizer to train the model; use the cross-entropy loss function; set the training parameters, batch size, initial learning rate and number of iterations; 5) After each complete data iteration, calculate the accuracy of the model on the test set and save the model weight with the highest current accuracy; 6) After training is completed, save the model weight file and inference program code.

8. A clothing recognition system for power workers, characterized in that: Using the method for identifying the clothing of electric power workers as described in any one of claims 1 to 7, the system is deployed in a monitoring system at a power operation site to monitor the clothing of the workers. The system includes: The human body detection module uses the pre-trained YOLOv5 model to perform human body detection on the input image to be identified, locate the area where the power workers are located, and crop the image part that only contains the human body area; An input module, used to input the cropped human body region image into the trained multi-label multi-task clothing recognition model; The judgment module analyzes the output results of the multi-label multi-task clothing recognition model to determine the clothing of the operators.

9. A clothing identification system for electric power workers according to claim 8, characterized in that: Also includes: The alarm module is used to issue an alarm prompt when it is determined that the operator is not wearing the required work clothes, safety helmets, insulating gloves or insulating shoes.