Power grid operator work clothes detection method, system and equipment based on deep learning, and medium
By constructing a deep learning-based detection model for the work clothes of power grid workers, the problem of accuracy in detecting work clothes in complex environments was solved, and stable and efficient identification was achieved in variable environments.
Patent Information
- Application Number
- CN202511473700.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-13
AI Technical Summary
Existing detection models for power grid workers' uniforms struggle to maintain stable detection accuracy in complex and ever-changing natural environments. In particular, when the uniforms are small or obscured by other objects, they are difficult to accurately identify between the uniforms and complex background information, leading to confusion during the identification process.
A deep learning-based model for detecting work clothes of power grid workers was constructed, including multiple convolutional normalized activation modules, cross-scale feature extraction modules, and deep feature fusion modules. Through multi-layer feature extraction and fusion, combined with data augmentation and training set partitioning, the model's recognition ability was improved.
It improves detection accuracy under adverse weather, low light, or blurry image conditions, enhances the ability to identify small or obscured work clothes, reduces false detection rate, and improves the accuracy and reliability of identification.
Smart Images

Figure CN121330604A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of workwear inspection technology, specifically to a method, system, equipment, and medium for inspecting workwear for power grid workers based on deep learning. Background Technology
[0002] With the rapid development of the power grid industry, the safety of power grid workers has received increasing attention. Wearing the correct and compliant work clothes is a crucial aspect of ensuring worker safety during power grid operations. However, traditional manual supervision methods suffer from low efficiency and high error rates, making it difficult to meet the high safety supervision requirements of modern power grid operations. Therefore, developing an efficient and accurate method for detecting the work clothes of power grid workers is of paramount importance.
[0003] In recent years, deep learning technology has achieved remarkable results in fields such as image recognition, speech recognition, and natural language processing due to its powerful data processing and feature extraction capabilities. As an important branch of machine learning, deep learning simulates the workings of the human brain's nervous system, utilizing multi-layered neural networks for automated data processing. Compared to traditional machine learning methods, deep learning technology can automatically learn and extract high-level features from data without manual feature extraction, thus greatly improving processing accuracy and efficiency. In power grid operation scenarios, due to the complexity and diversity of the working environment, traditional manual supervision methods often struggle to achieve comprehensive coverage and real-time monitoring of the work clothing worn by power grid workers. Furthermore, the complexity and sheer number of power grid operation scenarios make it difficult for traditional data processing methods to handle such massive datasets. Therefore, a method is needed to automatically and efficiently detect the work clothing worn by power grid workers.
[0004] This invention collects a large amount of image data of power grid workers wearing work clothes and uses deep learning algorithms for training to identify the work clothes worn by workers at power grid construction sites. After the model is trained, real-time images of power grid work sites can be input into the model for detection, thereby achieving automatic monitoring of the work clothes worn by power grid workers. Summary of the Invention
[0005] In view of the above-mentioned problems, the present invention is proposed.
[0006] Therefore, the technical problem solved by this invention is: how to solve the problem that existing power grid worker uniform detection models often fail to maintain stable detection accuracy in complex and ever-changing natural environments. For uniforms that are small in size or obscured by other objects, the detection capability of existing power grid worker uniform detection models is limited. Workers are often closely intertwined with complex background information such as surrounding transmission towers and other transmission equipment, which makes it easy for existing power grid worker uniform detection models to confuse the uniforms during the identification process and make it difficult to accurately distinguish the uniforms from this background information.
[0007] To address the aforementioned technical problems, this invention provides the following technical solution: a deep learning-based method for detecting work clothes of power grid workers, comprising: constructing an original dataset of work clothes of power grid workers; preprocessing the original dataset to obtain a processed dataset of work clothes of power grid workers, and dividing the processed dataset into a training set, a test set, and a validation set; constructing a deep learning-based detection model for work clothes of power grid workers, the detection model including multiple convolutional normalization activation modules, a cross-scale feature extraction module for work clothes, and a deep feature fusion module for work clothes; inputting images of work clothes of power grid workers into the detection model to identify the wearing status of work clothes of power grid workers; training the detection model to obtain a trained detection model for work clothes of power grid workers; and applying the trained detection model to identify the current wearing status of work clothes of power grid workers.
[0008] As a preferred embodiment of the deep learning-based method for detecting the work clothes of power grid workers described in this invention, the construction of the original power grid worker work clothes dataset includes: using a camera to photograph the work area where the power grid workers are located, and collecting image data of the work clothes of the power grid workers; collecting image data of the work clothes of the power grid workers from existing public datasets; and labeling the image data of the work clothes of the power grid workers, indicating whether they are wearing work clothes or not.
[0009] As a preferred embodiment of the deep learning-based method for detecting the work clothes of power grid workers described in this invention, the following steps are included: preprocessing the original power grid worker work clothes dataset to obtain a processed power grid worker work clothes dataset, and dividing the processed dataset into a training set, a test set, and a validation set. This includes: adjusting all power grid worker images in the original dataset to a uniform size; applying data augmentation techniques to each image to obtain multiple augmented images, which are then added to the dataset; and dividing the augmented dataset into a training set, a validation set, and a test set according to a set ratio.
[0010] As a preferred embodiment of the deep learning-based method for detecting work clothes of power grid workers described in this invention, the following steps are taken: The construction of a deep learning-based detection model for power grid workers' work clothes includes multiple convolutional normalization activation modules, a work clothes cross-scale feature extraction module, and a work clothes deep feature fusion module. The model inputs images of power grid workers' work clothes to identify the wearing status of their work clothes. This includes: using any image of power grid workers' work clothes as input to the detection model to extract shallow feature maps; performing multi-scale feature extraction based on the shallow feature maps to obtain mid-level feature maps; applying the work clothes deep feature fusion module to the mid-level feature maps to obtain deep feature maps; and performing position regression and category prediction using a detection head based on the deep feature maps to output the work clothes wearing type and position information of the power grid workers in the image.
[0011] As a preferred embodiment of the deep learning-based method for detecting work clothes of power grid workers described in this invention, the step of using any image of a power grid worker's work clothes as input to the detection model and extracting shallow feature maps includes: applying a convolutional normalization activation module to the image of the power grid worker's work clothes to perform convolution, batch normalization, and activation function operations to obtain a first shallow feature map of the work clothes; applying a convolutional normalization activation module to the first shallow feature map of the work clothes to perform convolution, batch normalization, and activation function operations to obtain a second shallow feature map of the work clothes; and applying the first work clothes cross-scale feature extraction... The module processes the shallow feature map of the second work uniform to obtain the shallow feature map of the third work uniform. The convolutional normalization activation module performs convolution, batch normalization, and activation function operations on the shallow feature map of the third work uniform to obtain the shallow feature map of the fourth work uniform. This fourth shallow feature map is then input into the second work uniform cross-scale feature extraction module for processing to obtain the shallow feature map of the fifth work uniform. The convolutional normalization activation module then performs convolution, batch normalization, and activation function operations on the shallow feature map of the fifth work uniform to obtain the shallow feature map of the sixth work uniform. This sixth shallow feature map is then input into the third work uniform cross-scale feature extraction module. The module processes the data to obtain a shallow feature map of the seventh workwear; the convolutional normalization and activation module performs convolution, batch normalization, and activation function operations on the shallow feature map of the seventh workwear to obtain a shallow feature map of the eighth workwear, and inputs the shallow feature map of the eighth workwear into the fourth workwear cross-scale feature extraction module for processing to obtain a shallow feature map of the ninth workwear; the step of performing multi-scale feature extraction based on the shallow feature map to obtain a mid-level feature map includes processing the shallow feature map of the ninth workwear using the fifth workwear cross-scale feature extraction module to obtain a mid-level feature map of the first workwear; and applying fast spatial pyramid pooling to process the first mid-level feature map. The work uniform's mid-layer feature map is used to obtain the second mid-layer feature map, which is then upsampled to obtain the third mid-layer feature map. The sixth shallow feature map is fused with the third mid-layer feature map along the channel dimension to obtain the fourth mid-layer feature map. The sixth work uniform's cross-scale feature extraction module is used to process the fourth mid-layer feature map to obtain the fifth mid-layer feature map, which is then upsampled to obtain the sixth mid-layer feature map. Finally, the fifth shallow feature map and the sixth mid-layer feature map are fused along the channel dimension to obtain the seventh mid-layer feature map.
[0012] This preferred scheme combines a multi-layer convolutional normalization activation module with multiple workwear cross-scale feature extraction modules to extract and enhance key information such as multi-scale textures and edges in shallow images layer by layer. This helps improve the model's ability to perceive differences in workwear morphology in the early feature expression stage and provides a more discriminative feature foundation for subsequent mid-level feature fusion.
[0013] As a preferred embodiment of the deep learning-based method for detecting work clothes of power grid workers according to the present invention, the step of obtaining a deep feature map based on a mid-level feature map and applying a work clothes deep feature fusion module includes: inputting a seventh work clothes mid-level feature map into a first work clothes deep feature fusion module to obtain a first work clothes deep feature map; applying a convolutional normalization activation module to perform convolution, batch normalization, and activation function operations on the first work clothes deep feature map to obtain a second work clothes deep feature map, and fusing the second work clothes deep feature map with a fifth work clothes mid-level feature map in the channel dimension to obtain a third work clothes deep feature map; applying a second work clothes deep feature fusion module to process the third work clothes deep feature map to obtain a fourth work clothes deep feature map, and applying a convolutional normalization activation module to perform convolution, batch normalization, and activation function operations on the fourth work clothes deep feature map to obtain a fifth work clothes deep feature map; and fusing the second work clothes mid-level feature map and the fifth work clothes deep feature map in the channel dimension. The process involves obtaining a sixth deep feature map of the work clothes; processing the sixth deep feature map using the third deep feature fusion module to obtain a seventh deep feature map of the work clothes; and then, based on the deep feature maps, performing position regression and category prediction through a detection head to output the type of work clothes worn by power grid workers and their location in the image. This includes inputting the first, fourth, and seventh deep feature maps of the work clothes into the first, second, and third detection heads of the power grid worker work clothes detection model, respectively, to obtain the eighth, ninth, and tenth deep feature maps of the work clothes; in the detection head, performing position regression and classification convolution on the inputs to perform target regression and target classification for the transmission line work clothes target; merging the eighth, ninth, and tenth deep feature maps of the work clothes, and simultaneously performing non-maximum suppression to remove overlapping prediction boxes to obtain the final output of the algorithm, namely the tenth deep feature map of the work clothes.
[0014] This preferred solution introduces multiple deep feature fusion modules for work clothes, combines dilated convolution with channel fusion mechanism to further enhance the expressive power of deep semantic features, improves the model's localization accuracy for targets of different scales through multi-detector parallel prediction, and finally reduces the false detection rate and improves the overall accuracy and stability of wearable recognition by fusing recognition results through nonmaximum suppression.
[0015] As a preferred embodiment of the deep learning-based method for detecting work clothes of power grid workers described in this invention, the following steps are included: training the power grid worker work clothes detection model to obtain a trained power grid worker work clothes detection model, which includes: after initializing parameters, dividing the training set and validation set data into multiple batches, inputting one batch of training set data into the power grid worker work clothes detection model for training at each time, and obtaining the training loss value of the current batch; after completing one round of training on all batches of data in the entire training set, inputting the validation set into the power grid worker work clothes detection model according to batches, and obtaining the corresponding batch loss value; the training of the power grid worker work clothes detection model ends when the batch loss value tends to converge; and applying the trained power grid worker work clothes detection model to identify the current wearing status of power grid workers' work clothes, which includes: applying the trained power grid worker work clothes detection model to identify and analyze the current remote sensing data, and outputting results including the specific type of work clothes wearing status and location identification.
[0016] This preferred solution effectively achieves automatic optimization and convergence judgment during the training process by inputting training data into the detection model in batches and combining a dynamic loss monitoring mechanism for the training and validation sets, thus avoiding overfitting. At the same time, the trained model can be applied to remote sensing image recognition tasks to achieve rapid identification of the clothing worn by power grid workers in large-scale operation scenarios, demonstrating good practicality and deployment efficiency.
[0017] This invention provides a deep learning-based system for detecting the work clothes of power grid workers.
[0018] To address the aforementioned technical problems, this invention provides the following technical solution: a deep learning-based system for detecting work clothes of power grid workers, comprising: a dataset construction module, a data processing module, a model construction module, a model training module, and a model application module; the dataset construction module is used to construct an original dataset of work clothes of power grid workers; the data processing module is used to preprocess the original dataset of work clothes of power grid workers to obtain a processed dataset of work clothes of power grid workers, and to divide the processed dataset of work clothes of power grid workers to obtain a training set, a test set, and a validation set; the model construction module is used to construct a deep learning-based model for detecting work clothes of power grid workers, the model of which includes multiple convolutional normalization activation modules, a cross-scale feature extraction module for work clothes, and a deep feature fusion module for work clothes, and inputs images of work clothes of power grid workers into the model of detecting work clothes of power grid workers to identify the wearing status of work clothes of power grid workers; the model training module is used to train the model of detecting work clothes of power grid workers to obtain a trained model of detecting work clothes of power grid workers; the model application module is used to apply the trained model of detecting work clothes of power grid workers to identify the wearing status of work clothes of current power grid workers.
[0019] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the deep learning-based method for detecting work clothes of power grid workers.
[0020] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of the deep learning-based method for detecting work clothes of power grid workers.
[0021] The beneficial effects of this invention are as follows: This invention effectively improves the detection accuracy under adverse weather, insufficient lighting, or blurry image conditions through a power grid worker work clothes detection model, ensuring stable performance in different environments; for power industry personnel who are small in size or obscured, the invention utilizes deep learning optimization strategies to enhance feature extraction capabilities, accurately identifying work clothes even in low image resolution or with obstructions, significantly improving detection capabilities; it can effectively distinguish work clothes from complex background information such as transmission towers, tree branches, and buildings, greatly reducing the false detection rate and improving the accuracy and reliability of identification. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 The above is a flowchart of a deep learning-based method for detecting work clothes of power grid workers, provided as an embodiment of the present invention.
[0024] Figure 2 This is an overall architecture diagram of a power grid worker's work clothes detection model provided in an embodiment of the present invention, which is a deep learning-based method for detecting power grid worker's work clothes.
[0025] Figure 3 This is an overall architecture diagram of the cross-scale feature extraction module for a deep learning-based method for detecting work clothes of power grid workers, provided as an embodiment of the present invention.
[0026] Figure 4 This is an overall architecture diagram of the deep feature fusion module of a deep learning-based method for detecting work clothes of power grid workers, provided in one embodiment of the present invention. Detailed Implementation
[0027] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0028] Example 1, referring to Figure 1 This is one embodiment of the present invention, which provides a deep learning-based method for detecting the work clothes of power grid workers, including: S1. Construct the original dataset of work clothes for power grid workers.
[0029] S2. Preprocess the original power grid worker work clothes dataset to obtain the processed power grid worker work clothes dataset, and divide the processed power grid worker work clothes dataset into training set, test set and validation set.
[0030] S3. Construct a deep learning-based detection model for power grid workers' work clothes. The detection model includes multiple convolutional normalization activation modules, a cross-scale feature extraction module for work clothes, and a deep feature fusion module for work clothes. Input the image of the power grid workers' work clothes into the detection model to identify the wearing status of the work clothes.
[0031] S4. Train the power grid worker's work clothes detection model to obtain a trained power grid worker's work clothes detection model.
[0032] S5. Apply the trained power grid worker uniform detection model to identify the current wearing status of power grid worker uniforms.
[0033] It should be noted that power grid workers perform construction work in various scenarios, including the field, mountains, and high-voltage environments. Their working environments are complex, their locations are not fixed, and there are many background interference factors, making it difficult for traditional image recognition algorithms to achieve stable wear recognition. This invention constructs a deep learning model through steps S1-S5 and introduces a multi-scale feature extraction and fusion structure. Combined with a dataset built from real image data, it can effectively extract key texture and edge information from workwear images in complex scenes, improving detection accuracy. Simultaneously, the model training process incorporates annotation information and optimization functions to achieve end-to-end learning, exhibiting strong robustness to changes in image details. It is suitable for the wear recognition needs of power grid workers in various scenarios and working conditions.
[0034] Example 2, refer to Figures 2-4 As an embodiment of the present invention, based on the previous embodiment, a deep learning-based method for detecting the work clothes of power grid workers is provided, comprising: In this embodiment, the power grid worker uniform dataset can be obtained by taking pictures of the work area where the power workers are working using shooting equipment (including drones, fixed-position high-definition cameras, etc.) to collect image data of the power grid worker uniforms; or by collecting image samples from existing publicly available power grid worker uniform image datasets; and by combining annotation tools to select and label the power grid worker targets in the images, with the label categories including wearing uniforms and not wearing uniforms, thereby constructing an original power grid worker uniform dataset for model training and validation.
[0035] In one alternative implementation, the power grid workers' work clothes dataset can also be formed by merging multiple datasets from different sources and unifying the image format, resolution, and annotation standards to create a multi-source fusion training dataset, thereby enhancing the coverage and diversity of the samples.
[0036] In another alternative implementation, the dataset of power grid workers' work clothes can also be constructed by extracting frame images from the video stream and performing automated detection and label generation, utilizing the differential information between consecutive frames to achieve semi-automatic annotation, thereby reducing labor costs and improving the efficiency of dataset construction.
[0037] This invention constructs a dataset of work clothes for power grid workers containing multiple sources and clear labels, which can ensure the integrity and accuracy of training samples for deep learning models, enabling the models to maintain high recognition accuracy under different work scenarios and lighting conditions, and providing reliable data support for subsequent model training and application.
[0038] Furthermore, in step S1, the original dataset of work clothes for power grid workers is constructed, including the following steps A1-A3: A1. Use a camera to photograph the work area where the power workers are working, and collect image data of the power grid workers' work clothes.
[0039] A2. Collect image data of work clothes of power grid workers from existing public datasets.
[0040] A3. Label the image data of power grid workers' work clothes, indicating whether they are wearing work clothes or not.
[0041] Specifically, in step A1, a camera (including but not limited to drones, fixed-position high-definition cameras, etc.) is used to photograph the work area where the power workers are located, and image data of the power grid workers' work clothes is collected.
[0042] Specifically, in step A2, an existing publicly available dataset of images of power grid workers' uniforms is used as the original dataset of power grid workers' uniforms.
[0043] In this embodiment of the application, the labeling in step A3 can be done manually in combination with labeling tools to select target boxes and assign categories to the collected images of power grid workers. Specifically, tools such as LabelImg and Labelme are used to label the target areas of the personnel in the images of power grid workers, determine the coordinates of the upper left and lower right corners and assign a wearing status label. The labeling results are divided into two categories: wearing work clothes and not wearing work clothes, which are used to generate labeled image samples required for training.
[0044] In one alternative implementation, labeling can also be based on a weakly supervised learning strategy, using a pre-trained model to perform preliminary identification and automatic bounding selection of the target region, combined with a small amount of manual correction to generate labels, thereby reducing the manual cost of large-scale labeling.
[0045] In another alternative implementation, labeling can also be achieved through automatic label propagation via the inter-frame coherence in the video sequence. That is, after labeling a small number of keyframe images, the labels are extended to other frame images through temporal consistency, thereby realizing semi-automatic batch labeling of video data.
[0046] This invention constructs a high-quality labeled dataset by clearly annotating the image labels of the wearing status, enabling deep learning models to accurately learn the difference features between images with and without clothing, thereby improving classification accuracy and target localization accuracy during the model training stage and providing a reliable data foundation for workwear recognition tasks.
[0047] Specifically, in step A3, the images of the power grid workers' uniforms are labeled. This involves using a labeling tool to select the location of the power workers, determining the coordinates of the top-left and bottom-right vertices of the box, and then labeling them (the labeling information is divided into two types: one indicating that they are wearing uniforms, and the other indicating that they are not wearing uniforms). Commonly used labeling tools include, but are not limited to, LabelImg, Labelme, VOTT, CVAT, SuperAnnotate, DataLoop, and RectLabel.
[0048] Furthermore, in step S2, the original power grid worker uniform dataset is preprocessed to obtain a processed power grid worker uniform dataset, and the processed power grid worker uniform dataset is divided into a training set, a test set, and a validation set, including the following steps B1-B3: B1. Adjust all images of power grid workers in the original dataset of power grid worker uniforms to a uniform size.
[0049] B2. For each image of a power grid worker, data augmentation techniques are used to obtain multiple augmented images of the power grid worker, which are then added to the dataset.
[0050] B3. The enhanced power grid worker work clothes dataset is divided into training set, validation set and test set according to the proportion.
[0051] Specifically, in step B1, all images of power grid workers in the original power grid worker work clothes dataset are adjusted to a uniform size.
[0052] Specifically, in step B2, considering the limited size of the original dataset, data augmentation techniques, including but not limited to geometric transformation, color gamut transformation, sharpness transformation, noise injection, and local erasure, are used for each image of a power grid worker to obtain multiple augmented images of power grid workers and add them to the dataset. This ensures that there are sufficient samples for training, validation, and testing, thereby enhancing the robustness of the power grid worker work clothes detection model and reducing the model's sensitivity to subtle changes in the images.
[0053] Specifically, in step B3, the enhanced power grid worker work clothes dataset is scientifically divided into a training set, a validation set, and a test set according to a ratio of 7:1.5:1.5.
[0054] In this embodiment, the power grid worker's work clothes detection model in step S3 may include a deep neural network structure comprising multiple convolutional normalized activation modules, multiple workwear cross-scale feature extraction modules (WCSFE), and multiple workwear deep feature fusion modules (WDFF). This model receives an image of the power grid worker's work clothes as input, extracts shallow, medium, and deep feature maps layer by layer, performs multi-scale feature fusion processing, and outputs the work clothes wearing type and its position in the image through multiple detection heads, thereby recognizing the wearing status of the power grid worker.
[0055] In one alternative implementation, the power grid worker uniform detection model can also employ a backbone network based on a Transformer structure or an attention-guided mechanism, combined with spatial and channel attention modules to optimize the response of key regions in the feature map, thereby improving the model's performance in recognizing targets in occluded and overlapping scenarios.
[0056] In another alternative implementation, the power grid worker uniform detection model can also construct a model cluster consisting of multiple sub-models through ensemble learning. Each sub-model focuses on different image feature dimensions (such as color, texture, or structural features), and finally outputs the recognition result through weighted voting or fusion strategies to enhance the model's generalization ability.
[0057] This invention extracts basic visual features by introducing a convolutional normalized activation module, combines it with a WCSFE module to capture multi-scale texture and edge information, and then uses a WDFF module for deep semantic enhancement. The constructed detection model can adapt to various complex power operation environments and still has high detection accuracy and stability under conditions such as changes in image quality and background interference, meeting the comprehensive requirements of engineering deployment for recognition robustness and real-time performance.
[0058] Furthermore, in step S3, a deep learning-based detection model for power grid workers' work clothes is constructed. This model includes multiple convolutional normalization activation modules, a cross-scale feature extraction module for work clothes, and a deep feature fusion module for work clothes. Images of power grid workers' work clothes are input into the detection model to identify the wearing status of their work clothes, including the following steps C1-C4: C1. Use any image of a power grid worker's work clothes as input to the power grid worker's work clothes detection model, and extract shallow feature maps.
[0059] C2. Perform multi-scale feature extraction based on shallow feature maps to obtain mid-level feature maps.
[0060] C3. Based on the mid-level feature map and applying the workwear deep feature fusion module, a deep feature map is obtained.
[0061] C4. Based on deep feature maps, position regression and category prediction are performed using a detection head to output the type of work clothes worn by power grid workers and their position in the image.
[0062] Furthermore, in step C1, an image of any power grid worker's work clothes is used as input to the power grid worker's work clothes detection model to extract shallow feature maps, including the following steps C11-C16: C11. Apply the convolution normalization activation module to perform convolution, batch normalization and activation function operations on the image of the work clothes of the power grid workers to obtain the first shallow feature map of the work clothes.
[0063] C12. Apply the convolution normalization activation module to perform convolution, batch normalization, and activation function operations on the first workwear shallow feature map to obtain the second workwear shallow feature map.
[0064] C13. Apply the first workwear cross-scale feature extraction module to process the shallow feature map of the second workwear to obtain the shallow feature map of the third workwear.
[0065] C14. Apply the convolution normalization activation module to perform convolution, batch normalization, and activation function operations on the shallow feature map of the third workwear to obtain the shallow feature map of the fourth workwear. Then, input the shallow feature map of the fourth workwear into the cross-scale feature extraction module of the second workwear for processing to obtain the shallow feature map of the fifth workwear.
[0066] C15. Apply the convolution normalization activation module to perform convolution, batch normalization and activation function operations on the shallow feature map of the fifth workwear to obtain the shallow feature map of the sixth workwear. Then, input the shallow feature map of the sixth workwear into the cross-scale feature extraction module of the third workwear for processing to obtain the shallow feature map of the seventh workwear.
[0067] C16. Apply the convolution normalization activation module to perform convolution, batch normalization and activation function operations on the shallow feature map of the seventh workwear to obtain the shallow feature map of the eighth workwear. Then, input the shallow feature map of the eighth workwear into the cross-scale feature extraction module of the fourth workwear for processing to obtain the shallow feature map of the ninth workwear.
[0068] Specifically, the overall architecture of the power grid worker's work clothes detection model is as follows: Figure 2 As shown, an arbitrary image of a power grid worker's work clothes, F1, is used as the input to the power grid worker's work clothes detection model. First, the Convolution_Batch Normalization_Rectified Linear Unit (Conv_BN_ReLU) module is applied to F1 for convolution, batch normalization, and activation function operations to obtain the first shallow feature map of the work clothes, F2. Then, the Conv_BN_ReLU module is applied to F2 for convolution, batch normalization, and activation function operations to obtain the second shallow feature map of the work clothes, F3. Next, the first WCSFE module is applied to F3 to obtain the third shallow feature map of the work clothes, F4. Finally, the Conv_BN_ReLU module is applied to F5 for convolution, batch normalization, and activation function operations to obtain the fourth shallow feature map of the work clothes, F5, which is then input into the second WCSFE module. The CSFE module processes the data to obtain the fifth shallow feature map of the workwear, F6. Then, the Conv_BN_ReLU module performs convolution, batch normalization, and activation function operations on F6 to obtain the sixth shallow feature map of the workwear, F7. F7 is then input into the third WCSFE module for processing to obtain the seventh shallow feature map of the workwear, F8. Finally, the Conv_BN_ReLU module performs convolution, batch normalization, and activation function operations on F8 to obtain the eighth shallow feature map of the workwear, F9. F9 is then input into the fourth WCSFE module for processing to obtain the ninth shallow feature map of the workwear, F10.
[0069] The Conv_BN_ReLU module consists of a 3×3 convolutional kernel with a stride of 2, a batch normalization (BN) layer, and a ReLU activation function layer, all concatenated. Conv_BN_ReLU is a neural network architecture specifically designed for processing two-dimensional image data. It extracts features such as edges, textures, and shapes from the image through convolutional layers, then normalizes these features using batch normalization layers to improve training speed and stability. Finally, the ReLU activation function layer introduces non-linearity, enabling the model to learn more complex feature representations. This process transforms the original two-dimensional image into a series of high-level feature maps. Since the initial convolutional layers may only detect simple low-level features such as edges and textures, subsequent network layers can combine these low-level features to form more complex and representative high-level features. These high-level feature maps with advanced features are crucial for subsequent image recognition, classification, or detection tasks.
[0070] The overall structure of WCSFE is as follows: Figure 3 As shown. The operation and construction process of the WCSFE module is as follows: The WCSFE module has three processing paths: L1, L2, and L3. L1 is the shallow-scale feature extraction path; L2 is the medium-scale feature extraction path; and L3 is the deep-scale feature extraction path.
[0071] The shallow feature map F3 of the second work uniform is used as the input feature map of the WCSFE module. For ease of description, F3 is denoted as X1 here.
[0072] In processing route L1, global max pooling, a 1×1 convolutional kernel, and the ReLU activation function are applied to process X1 to obtain the workwear cross-scale feature map X2; X1 and X2 are then subjected to channel weighting to obtain the workwear cross-scale feature map X3.
[0073] Channel weighting treats a 1×1×C feature map as a channel weight vector. Each channel of the H×W×C feature map is then multiplied according to its corresponding weight. This is typically achieved using a 1×1 convolutional layer, where the 1×1×C feature map acts as the kernel, convolving the H×W×C feature map. The result is a new H×W×C feature map where the pixel values of each channel are adjusted according to the corresponding weights in the 1×1×C feature map.
[0074] In processing route L2, X1 is first operated on with global max pooling and global average pooling respectively, and then channel dimension fusion is performed; then it is input into a 1×1 convolution kernel Conv1×1 for convolution operation to obtain the work clothes cross-scale feature map X4; then X3 and X4 are subjected to channel weighting operation to obtain the work clothes cross-scale feature map X6.
[0075] Channel dimensionality fusion refers to concatenating two or more feature maps with the same spatial dimensions (height and width) along the channel dimension to generate a new feature map. This new feature map has the same spatial dimensions as the input feature maps, but the number of channels is the sum of the number of channels in all input feature maps. Spatial Pyramid Pooling Fast (SPPF) is an optimization module introduced in YOLOv8. It uses a fast spatial pyramid pooling method to fuse global information at different scales to improve object detection performance while reducing computational redundancy, achieving higher efficiency and fewer floating-point operations.
[0076] In processing route L3, a convolution operation is performed on X1 using a convolution kernel of size 5×5 and a dilation rate of 3 to obtain the cross-scale feature map X5 of the work clothes.
[0077] After fusing X5 and X6 by channel dimension, the number of channels is adjusted by using a 1×1 convolution kernel Conv1×1 to obtain the workwear cross-scale feature map X7, which is the output feature map of the WCSFE module.
[0078] In this embodiment of the application, multi-scale feature extraction in step C2 can be achieved by the WCSFE module. The WCSFE module includes three feature extraction paths: shallow scale, medium scale, and deep scale. It performs 1×1 convolution, channel weighting, global pooling, and dilated convolution operations on the input feature map, respectively. After fusion, a multi-scale feature map is obtained, and the enhanced representation result is formed by concatenating the channel dimensions for subsequent feature extraction and recognition.
[0079] In an alternative implementation, multi-scale feature extraction can also be achieved by constructing a multi-layer feature fusion network based on a pyramid structure, extracting feature maps from different depths of the neural network, and then uniformly adjusting and fusing them to improve the model's ability to recognize different target sizes.
[0080] In another alternative implementation, multi-scale feature extraction can also be improved by introducing an attention mechanism and a dynamic receptive field adjustment method to dynamically select convolutional kernels with appropriate scales in the convolutional network, and adaptively extracting suitable scale information for different regions in the image, thereby improving the model's adaptability to changes in image scale.
[0081] This invention uses the WCSFE module to extract and fuse multi-scale paths in work clothes images, which significantly enhances the model's ability to recognize texture, contour and other information in target areas of different sizes (such as partially occluded work clothes or distant people), thereby improving the overall recognition accuracy and model robustness. It is particularly suitable for power grid operation scenarios with multiple environments and multiple perspectives.
[0082] Furthermore, in step C2, multi-scale feature extraction is performed based on the shallow feature map to obtain the mid-level feature map, including the following steps C21-C25: C21. The fifth workwear cross-scale feature extraction module is used to process the shallow feature map of the ninth workwear to obtain the middle feature map of the first workwear.
[0083] C22. Apply fast spatial pyramid pooling to process the mid-layer feature map of the first workwear to obtain the mid-layer feature map of the second workwear, and then perform upsampling to obtain the mid-layer feature map of the third workwear.
[0084] C23. Merge the shallow feature map of the sixth workwear with the middle feature map of the third workwear by channel dimension to obtain the middle feature map of the fourth workwear.
[0085] C24. The sixth workwear cross-scale feature extraction module is used to process the mid-layer feature map of the fourth workwear to obtain the mid-layer feature map of the fifth workwear. The mid-layer feature map of the fifth workwear is then upsampled to obtain the mid-layer feature map of the sixth workwear.
[0086] C25. Merge the shallow feature map of the fifth workwear and the middle feature map of the sixth workwear by channel dimension to obtain the middle feature map of the seventh workwear.
[0087] Specifically, firstly, the fifth WCSFE module is applied to process F10 to obtain the first workwear mid-layer feature map F11; secondly, Spatial Pyramid Pooling Fast (SPPF) is applied to process F11 to obtain the second workwear mid-layer feature map F12, which is then subjected to an upsampling operation (i.e., ... Figure 2 The upsample operation in the first step yields the third workwear mid-layer feature map F13. Then, F7 and F13 are fused along the channel dimension to obtain the fourth workwear mid-layer feature map F14. Subsequently, the sixth WCSFE module is applied to process F14 to obtain the fifth workwear mid-layer feature map F15, and F15 is upsampled to obtain the sixth workwear mid-layer feature map F16. Finally, F6 and F16 are fused along the channel dimension to obtain the seventh workwear mid-layer feature map F17.
[0088] SPPF is an optimization module introduced in YOLOv8. It uses a fast spatial pyramid pooling method to fuse global information at different scales to improve the performance of object detection, while reducing computational redundancy, achieving higher efficiency and fewer floating-point operations.
[0089] Furthermore, in step C3, based on the mid-layer feature map and applying the workwear deep feature fusion module, a deep feature map is obtained, including the following steps C31-C35: C31. Input the mid-layer feature map of the seventh work uniform into the deep feature fusion module of the first work uniform to obtain the deep feature map of the first work uniform.
[0090] C32. Apply the convolutional normalization activation module to perform convolution, batch normalization, and activation function operations on the deep feature map of the first workwear to obtain the deep feature map of the second workwear. Then, fuse the deep feature map of the second workwear with the middle feature map of the fifth workwear in terms of channel dimension to obtain the deep feature map of the third workwear.
[0091] C33. The second workwear deep feature fusion module is used to process the third workwear deep feature map to obtain the fourth workwear deep feature map. The convolution normalization activation module is then used to perform convolution, batch normalization and activation function operations on the fourth workwear deep feature map to obtain the fifth workwear deep feature map.
[0092] C34. Merge the mid-layer feature map of the second workwear and the deep feature map of the fifth workwear by channel dimension to obtain the deep feature map of the sixth workwear.
[0093] C35. The deep feature fusion module of the third workwear is used to process the deep feature map of the sixth workwear to obtain the deep feature map of the seventh workwear.
[0094] Specifically, firstly, F17 is input into the first WDFF module to obtain the first deep feature map of the workwear, F18; secondly, the Conv_BN_ReLU module is applied to F18 for convolution, batch normalization, and activation function operations to obtain the second deep feature map of the workwear, F19, and F19 is fused with F15 along the channel dimension to obtain the third deep feature map of the workwear, F20; subsequently, the second WDFF module is applied to F20 to obtain the fourth deep feature map of the workwear, F21, and the Conv_BN_ReLU module is applied to F21 for convolution, batch normalization, and activation function operations to obtain the fifth deep feature map of the workwear, F22; subsequently, F12 and F22 are fused along the channel dimension to obtain the sixth deep feature map of the workwear, F23; finally, the third WDFF module is applied to F23 to obtain the seventh deep feature map of the workwear, F24.
[0095] The overall structure of the WDFF module is as follows: Figure 4 As shown. The running and building process of the WDFF module is as follows: The feature map F17 in the middle layer of the seventh work uniform is used as the input feature map of WDFF. For ease of description, F17 is denoted as Y1 here.
[0096] First, a 3×3 convolution kernel is applied to Y1 to obtain the workwear depth feature fusion map Y2. Second, Y2 is convolved with dilated convolution kernels of size 3×3 (3), 5×5 (5), 7×7 (7), and 9×9 (9), respectively, to obtain the workwear depth feature fusion maps Y3, Y4, and Y5, and the workwear depth feature fusion map Y2. The first feature map is Y6. Then, Y3 and Y5 are fused along the channel dimension, and processed with ReLU activation and a 1×1 convolution kernel to obtain the workwear deep feature fusion map Y8. Y4 and Y6 are fused along the channel dimension, and processed with ReLU activation and a 1×1 convolution kernel to obtain the workwear deep feature fusion map Y7. Finally, Y7 and Y8 are fused along the channel dimension, and convolutional with a 1×1 kernel is performed. These are then input into the Coordinate Attention (CA) module for processing to obtain the workwear deep feature fusion map Y9, which is the output feature map of the WDFF module.
[0097] Channel attention (CA) is an attention mechanism used to enhance deep learning models' understanding of the spatial structure of input data. It aims to incorporate location information into channel attention, thereby enhancing feature representation capabilities while avoiding a significant increase in computational cost.
[0098] Furthermore, in step C4, based on the deep feature map, position regression and category prediction are performed using a detection head to output the work clothes type worn by the power grid workers and their position information in the image, including the following steps C41-C43: C41. Input the first, second, and third detection heads of the power grid worker's work clothes detection model into the first, second, and third detection heads, respectively, to obtain the eighth, ninth, and tenth work clothes deep feature maps.
[0099] C42. In the detection head, the input is used to perform position regression and target classification of the transmission line work clothes target using regression convolution and classification convolution respectively.
[0100] C43. Merge the deep feature maps of the eighth, ninth, and tenth work clothes, and simultaneously perform non-maximum suppression to remove overlapping prediction boxes, obtaining the final output of the algorithm, namely the deep feature map of the tenth work clothes.
[0101] Specifically, in the YOLOv8 detection head, regressive convolution is responsible for processing the input feature map to predict the location information of targets in power transmission line work clothes. It extracts features through a series of convolutional operations and ultimately outputs the coordinates of the bounding box (such as center point coordinates, width, and height), which are used to accurately locate the target object in the image. The regressive convolution is designed to capture the precise location features of the target, thus ensuring high accuracy in target detection.
[0102] In the YOLOv8 detector head, the classification convolution is responsible for identifying the category of targets in power line work clothes. It processes the input feature map as well, but outputs the probability distribution of the target belonging to each category. Through convolution operations and subsequent activation functions (such as softmax or sigmoid), the classification convolution can distinguish between different categories of targets and determine whether a target in the input image belongs to the power line work clothes category. This is crucial for achieving accurate target classification.
[0103] Within the deep learning-based object detection framework, the detection head handles the in-depth analysis and processing of feature maps, directly outputting the object's category, location, and presence confidence prediction. This output is typically in the form of a three-dimensional tensor, where each row contains prediction details at a specific spatial location, including the class probability of each predicted bounding box, the coordinates of the bounding box, and the object's presence confidence score. This information constitutes the key parameters necessary for the model to perform object localization and classification.
[0104] Bounding box prediction involves each cell predicting the coordinates of a set of bounding boxes and their confidence scores. These coordinates include the center point (x, y) (offset relative to the feature map cell), and the width and height (w, h) (scaled proportionally to the overall image size). These coordinate values are processed by a sigmoid function to ensure their output ranges between 0 and 1, thus enhancing training stability.
[0105] In addition to coordinate prediction, each bounding box also includes a confidence prediction value, which represents the probability that a specific category of object exists within the box.
[0106] Furthermore, in step S4, the power grid worker's work clothes detection model is trained to obtain a trained power grid worker's work clothes detection model, including the following steps D1-D3: D1. After initializing the parameters, the training set and validation set data are divided into multiple batches. Each batch of training set data is input into the power grid worker's work clothes detection model for training, and the training loss value of the current batch is obtained.
[0107] D2. After completing one round of training on all batches of data in the entire training set, input the validation set into the power grid worker's work clothes detection model according to batches to obtain the corresponding batch loss values.
[0108] D3. The training of the power grid worker's work clothes detection model ends when the batch loss value tends to converge.
[0109] Specifically, in step D1, during the training process, the power grid worker uniform detection model uses the SZIoU loss function to simultaneously optimize the location, size, confidence, and accuracy of the category prediction of the bounding box.
[0110] This invention designs an SZIoU loss function for training a model to detect the work clothes of power grid workers, in order to accurately calculate the difference between the predicted bounding box and the ground truth bounding box: in, and These represent the predicted bounding box and the ground truth bounding box, respectively. This represents the intersection-union ratio (IoU) between the predicted bounding box and the ground truth bounding box. ; and These represent the width and height of the actual bounding box, respectively. and These represent the width and height of the prediction box, respectively; and These represent the center points of the predicted bounding box and the ground truth bounding box, respectively. This represents the Euclidean distance between two pixels. and These represent the width and height of the smallest bounding box composed of the predicted box and the ground truth box, respectively.
[0111] The parameters of each layer are trained and updated on the power transmission line worker uniform detection model. First, all neural network parameters are initialized, and hyperparameters related to the power grid worker uniform detection model are set. The hyperparameters set include, but are not limited to, training epochs, batch size, optimizer selection, learning rate, weight initialization method, and Dropout ratio.
[0112] Specifically, after initializing the parameters, the training and validation sets are divided into multiple batches. Each batch of training data is input into the power grid worker uniform detection model for training, yielding a training loss value T for that batch. After one round of training on all batches of data in the entire training set, the validation set is input into the power grid worker uniform detection model in batches, yielding corresponding batch loss values T. The validation set loss values are primarily used to monitor whether the power grid worker uniform detection model is overfitting and to adjust the training strategy, such as terminating training early or adjusting the learning rate. During training and validation, the power grid worker uniform detection model automatically learns and adjusts its parameters based on each T and batch T. The training process ends when the batch T value converges after one or more rounds.
[0113] Batch L refers to the batch training loss value, which is the degree of difference between the model's predictions for all samples in a training batch and the true labels. This degree of difference is usually measured by a loss function, which calculates the error between the predicted and the true values.
[0114] Furthermore, in step S5, the trained power grid worker uniform detection model is applied to identify the current wearing status of the power grid worker's uniform, including the following steps: The trained model of power grid workers' work clothes is used to identify and analyze the current remote sensing data, and the output includes the specific type of work clothes worn and the location identification results.
[0115] Specifically, after the model training is completed, the trained model of power grid workers' work clothes is used to identify and analyze the current remote sensing data. The final output includes the specific type of work clothes wearing (there are two types: wearing work clothes and not wearing work clothes) and the location identification result.
[0116] This invention designs a cross-scale feature extraction module for work clothes (WCSFE). WCSFE innovates cross-scale feature extraction and fusion by extracting features through three feature extraction paths (L1, L2, and L3) at different levels (shallow, medium, and deep), and enriching the feature scale using global pooling, 1×1 convolution, and dilated convolution. Channel weighting enhances the sensitivity to important features, and concatenation operations fuse cross-scale features to provide comprehensive information. Dilated convolution expands the receptive field, improving the accuracy of detecting small or occluded work clothes. These innovations collectively constitute the core advantages of the model, enabling accurate detection of work clothes wearing conditions in complex environments. Cross-scale extraction and fusion of work clothes features are achieved through the three feature extraction paths (L1, L2, and L3) at different levels. Channel weighting is applied, and the weights of each channel of the feature map are adjusted using 1×1 convolution kernels, enhancing the model's sensitivity to important features and suppressing irrelevant or noisy features. This invention achieves accurate detection of workwear wearing conditions in complex and ever-changing natural environments through innovative cross-scale feature extraction and fusion mechanisms, channel weighting and fusion techniques, and the effective application of dilated convolution in the WCSFE module.
[0117] This invention designs a deep feature fusion module (WDFF) for work clothes, capturing multi-scale information through dilated convolution kernels with different dilation rates to enrich the feature maps. The feature maps are fused along the channel dimension and processed by ReLU activation and 1×1 convolution to enhance cross-scale feature complementarity. Finally, a CA attention module adaptively adjusts the channel weights to improve the representation of important features, suppress noise, and ensure accurate detection of work clothes wearing conditions in complex environments. This method can efficiently capture contextual information at different scales, providing a comprehensive and rich feature foundation for subsequent feature fusion. This invention also proposes a novel deep feature fusion method and applies it to the WDFF module. This method fuses feature maps at different scales along the channel dimension and optimizes them using ReLU activation and 1×1 convolution kernels to obtain new feature maps. This fusion method not only preserves the spatial structure of the original features but also achieves cross-scale feature complementarity and enhancement, further improving the model's detection accuracy and generalization ability.
[0118] This invention designs an SZIoU function as the loss function for training a power transmission line workwear detection model. The SZIoU loss function considers multiple aspects, including IoU, size difference, center point distance, and possible aspect ratio adjustments, to comprehensively evaluate the difference between the predicted bounding box and the ground truth bounding box, and guide the model training process. SZIoU comprehensively considers IoU optimization, size difference penalty, center point distance loss, and aspect ratio adjustment, forming an efficient and accurate loss evaluation mechanism. This function can comprehensively evaluate the difference between the predicted bounding box and the ground truth bounding box, guide model training, and significantly improve the accuracy and robustness of object detection. By optimizing IoU calculation, penalizing size inconsistency, quantifying positional offset, and adjusting aspect ratio differences, it provides strong technical support for improving the performance of object detection algorithms.
[0119] Example 3 is an embodiment of the present invention, which provides a method for detecting the work clothes of power grid workers based on deep learning. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0120] All images of power grid workers in the original dataset of power grid worker uniforms were adjusted to a uniform size of H×W×C, where H is the height of the feature map, W is the width of the feature map, and C is the number of channels of the feature map, with H being 512, W being 512, and C being 3.
[0121] Image F1 of the power grid worker's work clothes is used as input to the power grid worker's work clothes detection model. F1 has a size of 512×512×3. First, the Conv_BN_ReLU module is applied to F1 to obtain the first shallow feature map F2 of the work clothes, with a size of 512×512×64. Then, the Conv_BN_ReLU module is applied to F2 for convolution, batch normalization, and activation function operations to obtain the second shallow feature map F3 of the work clothes, with a size of 128×128×128. Next, the first WCSFE module is applied to F3 to obtain the third shallow feature map F4 of the work clothes, with a size of 128×128×256. Then, the Conv_BN_ReLU module is applied to F5 to obtain the fourth shallow feature map F5 of the work clothes, with a size of 64×64×512. F5 is then input into the second WCSFE module for processing to obtain the fifth work clothes. The shallow feature map F6 has a size of 64×64×512. Then, the Conv_BN_ReLU module is applied to F6 for convolution, batch normalization, and activation function operations to obtain the sixth workwear shallow feature map F7, with a size of 32×32×512. F7 is then input into the third WCSFE module for processing to obtain the seventh workwear shallow feature map F8, with a size of 32×32×512. Finally, the Conv_BN_ReLU module is applied to F8 for convolution, batch normalization, and activation function operations to obtain the eighth workwear shallow feature map F9, with a size of 16×16×512. F9 is then input into the fourth WCSFE module for processing to obtain the ninth workwear shallow feature map F10, with a size of 16×16×512.
[0122] The fifth WCSFE module is applied to process F10 to obtain the first workwear mid-layer feature map F11, which has a size of 16×16×512. Next, SPPF is applied to process F11 to obtain the second workwear mid-layer feature map F12, which also has a size of 16×16×512. Then, an upsampling operation (i.e., ...) is performed. Figure 2 The process involves several steps: First, an upsampling operation is performed to obtain the third mid-layer feature map F13, with a size of 32×32×512. Then, F7 and F13 are fused along the channel dimension to obtain the fourth mid-layer feature map F14, with a size of 32×32×1024. Next, the sixth WCSFE module is applied to process F14 to obtain the fifth mid-layer feature map F15, with a size of 32×32×1024. F15 is then upsampled to obtain the sixth mid-layer feature map F16, with a size of 64×64×512. Finally, F6 and F16 are fused along the channel dimension to obtain the seventh mid-layer feature map F17, with a size of 64×64×1024.
[0123] F17 is input into the first WDFF module to obtain the first deep feature map of the workwear, F18, with a size of 64×64×512. Next, the Conv_BN_ReLU module is applied to F18 for convolution, batch normalization, and activation function operations to obtain the second deep feature map of the workwear, F19, with a size of 32×32×512. F19 is then fused with F15 by channel dimension to obtain the third deep feature map of the workwear, F20, with a size of 32×32×1536. Subsequently, the second WDFF module is applied to process F20 to obtain the fourth deep feature map of the workwear. The layer feature map F21, with a size of 32×32×512, is processed by applying the Conv_BN_ReLU module to perform convolution, batch normalization, and activation function operations to obtain the fifth deep feature map F22, with a size of 16×16×512. Subsequently, F12 and F22 are fused by channel dimension to obtain the sixth deep feature map F23, with a size of 16×16×1024. Finally, the third WDFF module is applied to process F23 to obtain the seventh deep feature map F24, with a size of 16×16×512. Inputting F18, F21, and F24 into the first, second, and third detection heads of the algorithm, respectively, yields the eighth, ninth, and tenth deep feature maps of the work clothes, F25, F26, and F27. The output format is (x, y, w, h, conf1, conf2, ..., confi, ..., confn), where x represents the horizontal coordinate of the predicted box center, y represents the vertical coordinate of the predicted box center, w represents the width of the predicted box, h represents the height of the predicted box, confi represents the confidence score of the category corresponding to the i-th index, and n is the number of defect categories of the transmission tower base. At this point, the head recognition part of the power grid worker work clothes identification algorithm is completed. F25, F26, and F27 are merged, and non-maximum suppression is performed to remove overlapping predicted boxes, resulting in the final output F28 of the algorithm, which also has the format (x, y, w, h, conf1, conf2, ..., confi, ..., confn).
[0124] Example 4 is an embodiment of the present invention. This embodiment provides a deep learning-based system for detecting the work clothes of power grid workers, including a dataset construction module, a data processing module, a model construction module, a model training module, and a model application module.
[0125] The dataset building module is used to build a dataset of original power grid workers' work clothes.
[0126] The data processing module is used to preprocess the original power grid worker work clothes dataset to obtain the processed power grid worker work clothes dataset, and then divide the processed power grid worker work clothes dataset to obtain the training set, test set and validation set.
[0127] The model building module is used to construct a deep learning-based detection model for power grid workers' work clothes. The detection model includes multiple convolutional normalization activation modules, a work clothes cross-scale feature extraction module, and a work clothes deep feature fusion module. The model is used to identify the wearing status of power grid workers' work clothes by inputting images of their work clothes.
[0128] The model training module is used to train the detection model of power grid workers' work clothes, and obtain a trained power grid workers' work clothes detection model.
[0129] The model application module is used to apply the trained power grid worker uniform detection model to identify the current wearing status of power grid worker uniforms.
[0130] This embodiment also provides an electronic device applicable to a deep learning-based method for detecting the work clothes of power grid workers, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the deep learning-based method for detecting the work clothes of power grid workers as proposed in the above embodiment.
[0131] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements a deep learning-based method for detecting work clothes of power grid workers as proposed in the above embodiment.
[0132] The storage medium proposed in this embodiment belongs to the same inventive concept as the method for detecting the work clothes of power grid workers based on deep learning proposed in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0133] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0134] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for detecting work clothes of power grid workers based on deep learning, characterized in that: include, Construct a dataset of original power grid workers' work clothes; The original power grid worker work clothes dataset is preprocessed to obtain the processed power grid worker work clothes dataset. The processed power grid worker work clothes dataset is then divided into training set, test set and validation set. A deep learning-based detection model for power grid workers' work clothes was constructed. The model includes multiple convolutional normalization activation modules, a cross-scale feature extraction module for work clothes, and a deep feature fusion module for work clothes. Images of power grid workers' work clothes are input into the detection model to identify the wearing status of their work clothes. The detection model for power grid workers' work clothes was trained to obtain a well-trained detection model for power grid workers' work clothes. A pre-trained model for detecting the work clothes of power grid workers is used to identify the current wearing status of their work clothes.
2. The method for detecting work clothes of power grid workers based on deep learning as described in claim 1, characterized in that: The construction of the original power grid worker work clothes dataset includes, The camera was used to photograph the work area where the power workers were working, and the image data of the power grid workers' work clothes was collected. Image data of power grid workers' work clothes were collected from existing publicly available datasets; Label the images of power grid workers' work clothes, indicating whether they are wearing work clothes or not.
3. The method for detecting work clothes of power grid workers based on deep learning as described in claim 2, characterized in that: The original power grid worker uniform dataset is preprocessed to obtain a processed power grid worker uniform dataset. This processed dataset is then divided into training, testing, and validation sets, including... Adjust all images of power grid workers in the original dataset of power grid worker uniforms to a uniform size; For each image of a power grid worker, data augmentation techniques are used to obtain multiple augmented images of the power grid worker, which are then added to the dataset. The enhanced dataset of power grid workers' work clothes was divided into training set, validation set and test set according to proportion.
4. The method for detecting work clothes of power grid workers based on deep learning as described in claim 3, characterized in that: The aforementioned deep learning-based model for detecting the work clothes of power grid workers includes multiple convolutional normalized activation modules, a cross-scale feature extraction module for work clothes, and a deep feature fusion module for work clothes. Images of power grid workers' work clothes are input into the model to identify the wearing status of their work clothes. include, Use any image of a power grid worker's work clothes as input to the power grid worker's work clothes detection model to extract shallow feature maps; Multi-scale feature extraction is performed based on the shallow feature map to obtain the mid-level feature map; Based on the mid-level feature map and by applying the workwear deep feature fusion module, a deep feature map is obtained; Based on deep feature maps, a detection head is used to perform position regression and category prediction, outputting the type of work clothes worn by power grid workers and their location information in the image.
5. The method for detecting work clothes of power grid workers based on deep learning as described in claim 4, characterized in that: The step of using any image of a power grid worker's uniform as input to the power grid worker's uniform detection model and extracting shallow feature maps includes: The convolutional normalization activation module is applied to the image of the work clothes of power grid workers to perform convolution, batch normalization and activation function operations to obtain the first shallow feature map of the work clothes. The convolutional normalization activation module is applied to the shallow feature map of the first workwear to perform convolution, batch normalization and activation function operations to obtain the shallow feature map of the second workwear. The first workwear cross-scale feature extraction module is used to process the shallow feature map of the second workwear to obtain the shallow feature map of the third workwear. The convolutional normalization activation module is applied to the shallow feature map of the third workwear to perform convolution, batch normalization and activation function operations to obtain the shallow feature map of the fourth workwear. The shallow feature map of the fourth workwear is then input into the cross-scale feature extraction module of the second workwear for processing to obtain the shallow feature map of the fifth workwear. The convolutional normalization activation module is applied to the shallow feature map of the fifth workwear to perform convolution, batch normalization and activation function operations to obtain the shallow feature map of the sixth workwear. The shallow feature map of the sixth workwear is then input into the cross-scale feature extraction module of the third workwear for processing to obtain the shallow feature map of the seventh workwear. The convolutional normalization activation module is applied to the shallow feature map of the seventh workwear to perform convolution, batch normalization and activation function operations to obtain the shallow feature map of the eighth workwear. The shallow feature map of the eighth workwear is then input into the cross-scale feature extraction module of the fourth workwear for processing to obtain the shallow feature map of the ninth workwear. The process of performing multi-scale feature extraction based on shallow feature maps to obtain mid-level feature maps includes, The shallow feature map of the ninth workwear was processed by the cross-scale feature extraction module of the fifth workwear to obtain the middle feature map of the first workwear. The first workwear mid-layer feature map is processed by fast spatial pyramid pooling to obtain the second workwear mid-layer feature map, and then an upsampling operation is performed to obtain the third workwear mid-layer feature map. The shallow feature map of the sixth work uniform and the middle feature map of the third work uniform are fused along the channel dimension to obtain the middle feature map of the fourth work uniform. The sixth workwear cross-scale feature extraction module is used to process the mid-layer feature map of the fourth workwear to obtain the mid-layer feature map of the fifth workwear. The mid-layer feature map of the fifth workwear is then upsampled to obtain the mid-layer feature map of the sixth workwear. The shallow feature map of the fifth workwear and the middle feature map of the sixth workwear are fused along the channel dimension to obtain the middle feature map of the seventh workwear.
6. The method for detecting work clothes of power grid workers based on deep learning as described in claim 5, characterized in that: The process of obtaining a deep feature map based on the mid-layer feature map and applying the workwear deep feature fusion module includes, The middle layer feature map of the seventh work uniform is input into the deep feature fusion module of the first work uniform to obtain the deep layer feature map of the first work uniform; The convolutional normalization activation module is applied to the deep feature map of the first workwear to perform convolution, batch normalization and activation function operations to obtain the deep feature map of the second workwear. The deep feature map of the second workwear is then fused with the middle feature map of the fifth workwear by channel dimension to obtain the deep feature map of the third workwear. The deep feature fusion module of the second work clothes is used to process the deep feature map of the third work clothes to obtain the deep feature map of the fourth work clothes. Then, the convolution normalization activation module is used to perform convolution, batch normalization and activation function operations on the deep feature map of the fourth work clothes to obtain the deep feature map of the fifth work clothes. The middle layer feature map of the second work uniform and the deep layer feature map of the fifth work uniform are fused along the channel dimension to obtain the deep layer feature map of the sixth work uniform. The deep feature map of the sixth workwear is processed using the third workwear deep feature fusion module to obtain the deep feature map of the seventh workwear. The method, based on deep feature maps, uses a detection head to perform position regression and category prediction, outputting information such as the type of work clothes worn by power grid workers and their location in the image. The first, fourth, and seventh deep feature maps of the work clothes are input into the first, second, and third detection heads of the power grid worker's work clothes detection model, respectively, to obtain the eighth, ninth, and tenth deep feature maps of the work clothes. In the detection head, the input is used to perform position regression and target classification of the transmission line work clothes target using regression convolution and classification convolution respectively; The deep feature maps of the eighth, ninth, and tenth work clothes are merged, and non-maximum suppression is performed to remove overlapping prediction boxes, resulting in the final output of the algorithm, namely the deep feature map of the tenth work clothes.
7. The method for detecting work clothes of power grid workers based on deep learning as described in claim 6, characterized in that: The process of training the power grid worker's work clothes detection model to obtain a trained power grid worker's work clothes detection model includes, After initializing the parameters, the training set and validation set data are divided into multiple batches. Each batch of training set data is input into the power grid worker's work clothes detection model for training, and the training loss value of the current batch is obtained. After completing one round of training on all batches of data in the entire training set, the validation set is input into the power grid worker's work clothes detection model according to batches to obtain the corresponding batch loss values; The training of the power grid worker's work clothes detection model ends when the batch loss value tends to converge. The application uses a pre-trained model to detect the work clothes of power grid workers, identifying the current wearing status of their work clothes. include, The trained model of power grid workers' work clothes is used to identify and analyze the current remote sensing data, and the output includes the specific type of work clothes worn and the location identification results.
8. A deep learning-based system for detecting work clothes of power grid workers, employing the deep learning-based method for detecting work clothes of power grid workers as described in any one of claims 1 to 7, characterized in that, include: The system includes a dataset construction module, a data processing module, a model construction module, a model training module, and a model application module. The dataset construction module is used to construct the original dataset of work clothes for power grid workers. The data processing module is used to preprocess the original power grid worker work clothes dataset to obtain the processed power grid worker work clothes dataset, and to divide the processed power grid worker work clothes dataset to obtain the training set, test set and validation set. The model building module is used to build a deep learning-based detection model for power grid workers' work clothes. The detection model includes multiple convolutional normalization activation modules, a cross-scale feature extraction module for work clothes, and a deep feature fusion module for work clothes. The model inputs images of power grid workers' work clothes to identify the wearing status of their work clothes. The model training module is used to train the power grid worker's work clothes detection model to obtain a trained power grid worker's work clothes detection model; The model application module is used to apply a trained power grid worker uniform detection model to identify the current wearing status of power grid worker uniforms.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the deep learning-based method for detecting the work clothes of power grid workers as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the deep learning-based method for detecting the work clothes of power grid workers as described in any one of claims 1 to 7.