A method, system, equipment, and medium for identifying workers at power grid construction sites.
By using the CSII-Net network for data preprocessing and establishing identity detection algorithms at power construction sites, the problems of low efficiency and insufficient accuracy in identity detection of construction site workers have been solved, achieving efficient and accurate identity recognition.
Patent Information
- Application Number
- CN202510828717.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The identification of workers at power construction sites is inefficient and inaccurate, and traditional methods are difficult to achieve fast and accurate identification in complex and ever-changing environments.
A method for identifying workers at power grid construction sites is proposed, including data preprocessing, establishment and deployment of identity detection algorithms, and the use of an identity detection network (CSII-Net) for identity detection. The network includes an information perception module, a dual-branch identity detection head module, and a spatial pyramid pooling module. Data quality and adaptability are improved through data filtering, labeling, enhancement, and scenario simulation.
It achieves efficient and accurate worker identification in complex environments, improves the model's generalization and detection accuracy, and can identify the identity of workers at construction sites in real time.
Smart Images

Figure CN120356243B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, and in particular to a method, system, equipment, and medium for identifying workers at power grid construction sites. Background Technology
[0002] In power construction sites, accurate and efficient identification of workers is crucial for ensuring operational safety and improving work efficiency. Traditional identification methods, such as manual verification of identification documents or reliance on fixed surveillance camera systems, have many limitations. For example, due to high personnel mobility and changing scenarios, traditional methods often struggle to achieve rapid and accurate identification of workers.
[0003] In environments such as construction sites, worker identity verification and security monitoring are crucial. However, these environments are typically characterized by complex and variable conditions and high population density, posing significant challenges to identity verification. Currently, some scenarios rely on manual verification or simple access control systems, but these methods are not only inefficient but also susceptible to human error, making it difficult to achieve comprehensive and reliable identity management.
[0004] Currently, the identification of personnel at construction sites or substations along power transmission lines requires manual screening and inspection of image data, or on-site inspections, which incurs significant time and professional manpower costs. Summary of the Invention
[0005] In view of the aforementioned existing problems, the present invention is proposed.
[0006] Therefore, the present invention provides a method, system, equipment and medium for identifying workers at power grid construction sites, which can solve the problems of low efficiency and insufficient accuracy in identifying workers at power construction sites.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0008] In a first aspect, the present invention provides a method for identifying workers at power grid construction sites, comprising:
[0009] First data of the target power construction site is acquired, and the first data is preprocessed to obtain second data;
[0010] The first preprocessing includes data filtering, data annotation, data augmentation, and data partitioning;
[0011] A first identity detection algorithm is established, and the second data is used as the training set and verification set for the first identity detection algorithm.
[0012] The first identity detection algorithm includes an identity detection network, which includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module.
[0013] The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the location regression of the target person, and the second branch is used to calculate the identity classification of the target person.
[0014] The identity of the workers at the target power grid construction site is verified using the first identity detection algorithm after the verification is completed.
[0015] As a preferred embodiment of the power grid construction site worker identification method of the present invention, the first branch and the second branch include:
[0016] Both the first branch and the second branch are used to perform several convolution operations on the feature map that has been input into the second dual-branch identity detection head module and has undergone channel compression.
[0017] The first branch includes several convolution operations, including target regression convolution operations.
[0018] The second branch includes several convolution operations, including object classification convolution operations.
[0019] As a preferred embodiment of the power grid construction site worker identity detection method described in this invention, the third spatial pyramid pooling module is used to perform downsampling in the identity detection network;
[0020] The third spatial pyramid pooling module includes channel number compression operation, splicing operation, normalization operation, and several max pooling operations.
[0021] In a preferred embodiment of the power grid construction site worker identification method of the present invention, the first information sensing module includes:
[0022] Compress the second data by increasing the number of channels;
[0023] The second data, after channel compression, undergoes segmentation, concatenation, and several fusion convolution and self-attention mechanism operations.
[0024] As a preferred embodiment of the method for identifying workers at power grid construction sites according to the present invention, the first preprocessing of the first data includes:
[0025] The first data consists of images of workers at the target power construction site;
[0026] The first data is then filtered and labeled;
[0027] The first data after filtering and labeling is augmented to obtain the second data.
[0028] As a preferred embodiment of the power grid construction site worker identification method of the present invention, the step of data augmentation of the first data after screening and labeling includes:
[0029] Based on the safety helmet categories of construction workers at the work site in the first data, use image annotation tools to annotate the category and coordinate frame information of the personnel targets in all worker images;
[0030] The categories are divided into management personnel, technical personnel, construction and inspection personnel, supervision personnel, and other personnel.
[0031] After obtaining the filtered and labeled small sample dataset, data augmentation is performed on any staff image in the small sample dataset using data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition techniques. All data-augmented images are then combined with the small sample dataset to form an enhanced staff identity detection dataset.
[0032] The enhanced staff identity detection dataset is augmented with scenario simulation to obtain the second dataset.
[0033] As a preferred embodiment of the method for detecting the identity of workers at power grid construction sites according to the present invention, the scenario includes at least one or more of the following: rainy weather scenario and foggy weather scenario.
[0034] Secondly, the present invention provides a system for identifying workers at power grid construction sites, comprising:
[0035] The data acquisition and processing module is used to acquire first data from the target power construction site and perform first preprocessing on the first data to obtain second data.
[0036] The first preprocessing includes data filtering, data annotation, data augmentation, and data partitioning;
[0037] The algorithm establishment module is used to establish a first identity detection algorithm, and uses the second data as the training set and verification set of the first identity detection algorithm;
[0038] The first identity detection algorithm includes an identity detection network, which includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module.
[0039] The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the location regression of the target person, and the second branch is used to calculate the identity classification of the target person.
[0040] The detection module is used to detect the identity of the workers at the target power grid construction site based on the first identity detection algorithm after verification.
[0041] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.
[0042] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0043] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention proposes a method, system, equipment, and medium for detecting the identity of workers at power grid construction sites. It acquires first data from a target power construction site, performs a first preprocessing on the first data to obtain second data, establishes a first identity detection algorithm, and uses the second data as both a training set and a validation set for the first identity detection algorithm, and then performs identity detection of workers at the target power grid construction site based on the validated first identity detection algorithm. This invention establishes a highly generalizable, accurate, and automated network for detecting the identity of construction site workers, which can efficiently identify the identities of construction site workers. After training, the network designed in this invention can be deployed at construction sites or substations along various transmission lines to achieve real-time observation of the identities of construction site workers.
[0044] Specifically, the first information perception module enhances the ability to perceive the features of workers at super-resolution pixels in construction scenarios. By adaptively and rationally allocating feature map data, this module can effectively fit the fine positional information within the super-resolution feature map while ensuring detection efficiency. The second dual-branch identity detection head module enhances the detection effect of workers from multiple angles in construction site feature images. This detection head can adapt to the dynamic changes of workers and effectively eliminate the interference of background noise at the construction site, thereby improving the detection accuracy of the model. The third spatial pyramid pooling module enhances the trainable offset and the fitting ability of worker identity targets during downsampling in the network. Furthermore, this structure can improve the generalization of the network and amplify global interactive features, thereby avoiding the loss of key information and feature space distortion. Attached Figure Description
[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 A flowchart of a method for detecting the identity of workers at a power grid construction site, provided as an embodiment of the present invention.
[0047] Figure 2 This is a schematic diagram of the CSII-Net network for detecting the identity of workers at a power grid construction site, provided as an embodiment of the present invention.
[0048] Figure 3 This is a structural diagram of an SRIC module for a method of identifying workers at a power grid construction site, provided as an embodiment of the present invention.
[0049] Figure 4 This is a structural diagram of the DDB-Head module of a method for detecting the identity of workers at a power grid construction site, provided as an embodiment of the present invention.
[0050] Figure 5 The diagram shows the structure of a DDSPPF module for a method of identifying workers at a power grid construction site, as provided in one embodiment of the present invention.
[0051] Figure 6 This is an internal structural diagram of a computer device for a method of identifying workers at a power grid construction site, provided as an embodiment of the present invention. Detailed Implementation
[0052] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0053] Example 1, referring to Figures 1-6 This is the first embodiment of the present invention, which provides a method for detecting the identity of workers at a power grid construction site, including:
[0054] Before detailing the embodiments of this application, some related concepts will be explained for clarity.
[0055] The object-aware mixer CAMixer, as an existing technology, uses a learnable predictor to generate multiple guides, including offsets for window warping, masks for classifying windows, and convolutional attention to impart dynamic properties to the convolutions. This attention can adaptively adjust to include more useful textures and improve the representational power of the convolutions. Other operations such as convolution, batch normalization, ReLU activation functions, and concatenation are common in the deep learning industry.
[0056] Normalized Batch Normalization (BN) is a technique used during the training of a neural network to normalize the input data for each mini-batch. It reduces internal covariate shift by normalizing the input data to a distribution with a mean of 0 and a variance of 1 (or close to this distribution), thereby accelerating the training of the neural network and improving its performance.
[0057] ReLU (Rectified Linear Unit): refers to the activation function, which is a commonly used non-linear activation function in deep learning.
[0058] ELA: An existing technology, it is an efficient local attention (ELA) method that can accurately locate the region of interest without dimensionality reduction by effectively encoding two one-dimensional location feature maps, while allowing for lightweight implementation.
[0059] Deformable Conv layer: This is an existing technology that enables the network to adapt to changes in the shape, scale, and pose of objects by dynamically adjusting the sampling position of the convolution kernel, thereby improving the robustness of the model in complex scenes.
[0060] Convolution (Conv) is a one-dimensional convolution operation where the convolution kernel slides along the time or sequence axis of the feature map to extract local features from sequential data. The convolution operation operates simultaneously on all input channels and integrates cross-channel information through shared weights. It efficiently captures local patterns in sequential data while preserving global contextual information. It is an existing technology.
[0061] SRIC (Super Resolution Information Capture): A lightweight super-resolution information perception module, namely the first information perception module in this application, is used to enhance the perception of worker features under super-resolution pixels in construction scenarios. By adaptively allocating feature map data, it enhances detection efficiency while ensuring fine position fitting within the super-resolution feature map.
[0062] DDB-Head (Deformable Double Branch Head): A dynamic double-branch identity detection head, namely the second double-branch identity detection head module in this application, is used to enhance the identity detection effect of workers from multiple angles in the feature image of the construction site and adapt to the dynamic changes of workers; at the same time, it can also better eliminate the interference of background noise at the construction site and enhance the detection accuracy of the model.
[0063] DDSPPF (Deformable DAT Spatial Pyramid Pooling Factorization): A spatial pyramid pooling structure, namely the third spatial pyramid pooling module in this application, is used to downsample the network, enhance the trainable offsets and the fitting ability to the worker identity target, improve the generalization of the network, amplify global interaction features, and avoid the loss of key information and feature space distortion.
[0064] CSII-Net (Construction-Site Staff Identity Identification Net): This is a network for identifying the identity of construction site staff, as described in this application, used to identify staff targets at power construction sites.
[0065] Existing technologies have some problems, such as the complex environment of power construction sites making image recognition difficult; or inaccurate worker identification tags affecting the accuracy of identity detection.
[0066] This application provides a method that can effectively solve the problems mentioned above. The following will describe in detail how to implement the method for detecting the identity of workers at the power grid construction site in conjunction with several embodiments.
[0067] Figure 1 A method for identifying workers at power grid construction sites is shown, including:
[0068] S101, acquire the first data of the target power construction site, and perform the first preprocessing on the first data to obtain the second data;
[0069] In an optional embodiment, the target power construction site can be any power construction site that needs to be monitored, such as a construction site along a transmission line, a substation construction site, a power distribution room construction site, etc. The specific scenario is not limited in this application.
[0070] In an optional embodiment, the first data is an image of workers at the target power construction site, which can be obtained through image acquisition devices such as cameras and drones.
[0071] It should be noted that the first preprocessing involves performing a series of operations on the first data to prepare it suitable for subsequent training and validation of the identity detection algorithm. Specific steps in the first preprocessing may include, but are not limited to, image scaling, cropping, grayscale conversion, denoising, contrast enhancement, and data augmentation techniques such as rotation, flipping, and translation to improve the model's generalization ability. These preprocessing steps aim to eliminate redundant information in the image and highlight key features, thereby ensuring that the subsequent identity detection algorithm can accurately and efficiently identify the target person. A carefully designed preprocessing workflow can significantly improve the accuracy and robustness of identity detection.
[0072] In this embodiment of the application, the first preprocessing includes data filtering, data annotation, data augmentation, and data partitioning.
[0073] In this embodiment of the application, the first preprocessing of the first data includes:
[0074] The first data is images of workers at the target power construction site;
[0075] The first set of data is filtered and labeled;
[0076] Data augmentation is performed on the first set of filtered and labeled data to obtain the second set of data.
[0077] In an optional embodiment, screening and labeling can be performed by professionals to ensure the accuracy and reliability of the data. The screening process may involve removing blurry, occluded, or poor-quality images, while labeling may include assigning a unique identification tag to each worker in the image and annotating their location information. Such processing steps are crucial for training a high-accuracy identity detection algorithm.
[0078] In another alternative embodiment, screening and labeling may be performed in a semi-automatic or fully automatic manner to improve efficiency. In a semi-automatic manner, professionals can manually label workers and their identification tags in images using image annotation tools. In a fully automatic manner, advanced computer vision technologies, such as deep learning models, can be used to automatically detect and label workers.
[0079] In one optional embodiment, data augmentation can employ various techniques, including but not limited to color dithering, brightness adjustment, edge enhancement, and blurring, to increase the diversity and complexity of the dataset. These augmentation techniques can simulate different lighting conditions, weather conditions, and shooting angles, thereby improving the model's adaptability and robustness in real-world application scenarios. By training on the augmented dataset, the identity detection algorithm can better learn the characteristics of workers in various environments, thus improving the accuracy and stability of identity detection.
[0080] In this embodiment of the application, data augmentation of the first data after filtering and labeling includes:
[0081] Based on the safety helmet categories of construction workers at the work site in the first data, use image annotation tools to annotate the category and coordinate frame information of the personnel targets in all worker images;
[0082] The categories are divided into management personnel, technical personnel, construction and inspection personnel, supervision personnel, and other personnel.
[0083] After obtaining the filtered and labeled small sample dataset, data augmentation techniques such as data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition are applied to any staff image in the small sample dataset. All data-augmented images are then combined with the small sample dataset to form an augmented staff identity detection dataset.
[0084] Data augmentation through scenario simulation was performed on the enhanced staff identity detection dataset to obtain the second dataset.
[0085] Specifically, images of workers at power construction sites are acquired using imaging equipment (such as drones and fixed equipment), and these images are then filtered for identity verification. After the identity verification images are filtered, image annotation tools are used to label the categories and bounding boxes of all worker images according to the type of safety helmet worn by the workers. The categories are divided into management personnel (red safety helmets), technicians (blue safety helmets), construction and inspection personnel (yellow safety helmets), supervisors (white safety helmets), and other personnel (other colored safety helmets or no safety helmets). Finally, a small sample dataset of filtered and labeled worker identities is obtained. .
[0086] It should be noted that in most cases, due to the many limitations in collecting images of workers at construction sites, the number of images of workers collected is generally less than 1,000, which is insufficient to cover the diversity and complexity required for training the detection task, making it difficult for the model to be trained effectively.
[0087] It should also be noted that commonly used image annotation tools include LabelImg, LabelME, etc.
[0088] Furthermore, based on the small sample dataset of staff identities after filtering and labeling... ,right Data augmentation is performed on any one of the staff images using techniques such as data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition. All data-augmented images are then compared with... Together, they form a dataset for enhanced staff identity detection. .
[0089] Furthermore, based on Data augmentation is performed by simulating different scenarios to increase the robustness of the dataset used to train the network.
[0090] In one optional embodiment, the scenario can include different weather conditions, lighting conditions, construction equipment layout, and personnel distribution. By simulating these scenarios, the dataset can be further enriched, enabling the model to better adapt to various real-world construction environments and improving the accuracy and stability of identity detection.
[0091] In the embodiments of this application, the scenarios include at least one or more of the following: rainy weather scenarios and foggy weather scenarios.
[0092] For example, the following simulations of rainy and foggy environments are used to increase the robustness of the network training on the dataset, focusing on the rainy scenario:
[0093] Firstly, regarding Any image to be enhanced Generate a random noise image of the same size as the original staff member's identity image. This simulates the random effect of raindrops falling on an image. (Random noise graph) The calculation formula is as follows:
[0094] (1)
[0095] in, Indicated in random noise graph middle The noise value at that location, and This represents a random number in the interval [0,1].
[0096] Then, to Generate blur effect Based on the length, angle, and width of the raindrops, a motion blur effect is applied to make the raindrops appear stretched and blurred in the image. The formula for generating it is as follows:
[0097] (2)
[0098] in, Indicates the position of the pixel to be processed. The fuzzy noise value at that location It is a fuzzy standard deviation that controls the length and width of raindrops.
[0099] Subsequently, a motion blur kernel is applied to blur raindrops in specific directions and lengths, making them appear stretched and simulating the motion of real raindrops. Motion blur kernel The formula is as follows:
[0100] (3)
[0101] in, Indicates the location Motion blur kernel at the location, For the angle of the raindrops, The length of the raindrop, This is the Diracdelta function.
[0102] It should be noted that the motion blur kernel is a mathematical model that describes the blurring effect produced by camera or object motion in an image. Its result is typically a two-dimensional matrix representing the direction and intensity of the blur.
[0103] It should also be noted that data rotation involves rotating the image by a certain angle to increase data diversity. Flipping involves flipping the image horizontally or vertically. Cropping involves removing a portion of the image. Random occlusion involves randomly occluding portions of the image. Contrast adjustment involves adjusting the image's contrast to make bright areas brighter and dark areas darker. Random scaling involves randomly adjusting the image's size. Noise addition techniques involve adding random noise to the image.
[0104] Then, image weighting is performed: the generated raindrop layer is compared with the original image. Weighted overlay is used to simulate the effect of raindrops superimposed on an image. The formula is as follows:
[0105] (4)
[0106] in, Indicates the location The image pixel values at which the raindrop effect is applied. Indicates the location The original image pixel values at that location, This is a weighting factor used to adjust the transparency of the raindrop effect.
[0107] Through the above steps, staff identification images can be obtained. Images after rain simulation For the dataset Each image was used to simulate a rainy day scene, resulting in a dataset of simulated rainy day scenes. .
[0108] For foggy scenarios: First, regarding Any image to be enhanced ,for Each pixel in Its brightness under foggy conditions can be calculated using the following formula:
[0109] (5)
[0110] (6)
[0111] In the formula: Indicates the original image at position pixel values, This indicates the location after atomization. pixel values, It is the transmittance of light through the fog, which is related to the distance d from the pixel to the center of the fog, where d is the distance from the pixel to the center of the fog. , () are the coordinates of the atomization center. This indicates the adjusted background brightness. This indicates transmittance.
[0112] It should be noted that the above steps are used to... Each pixel in the image is processed to obtain the image after simulating a foggy scene. For the dataset All images were simulated for foggy scenes, resulting in a dataset with simulated foggy scenes. .
[0113] Furthermore, and By merging the samples, we can obtain smaller sample data. Sample augmentation dataset after image transformation data augmentation, rain simulation, and fog simulation. .
[0114] Furthermore, regarding the dataset The training data, validation data, and test data are randomly divided proportionally to obtain training data. Validation data and test data .
[0115] It should be noted that acquiring the first data from the target power construction site and performing initial preprocessing on this data can significantly improve the training effect and practical application performance of the identity detection model. By filtering, labeling, enhancing, and reasonably dividing the original image data, not only is the quality and accuracy of the data ensured, but the diversity and complexity of the dataset are also greatly enriched. This preprocessing process helps the model learn more feature information from different scenarios, thereby improving the accuracy and robustness of identity detection. Especially when facing the complex and ever-changing environment of power grid construction sites, a carefully preprocessed dataset enables the model to better adapt to various challenges and achieve more reliable and efficient identity detection.
[0116] S102, establish a first identity detection algorithm, and use the second data as the training set and verification set for the first identity detection algorithm;
[0117] In an optional embodiment, the identity detection algorithm can employ deep learning techniques, such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs), to build an efficient and accurate identity detection model. These algorithms can automatically extract features from input images and learn unique feature representations of different individuals through training. During training, an augmented dataset (i.e., second data) is used as both the training and validation sets, and the model parameters are iteratively optimized to gradually converge the model to a high-performance state.
[0118] In an alternative embodiment, the identity detection algorithm can also employ lightweight neural network models, such as MobileNet or EfficientNet, to meet the needs of real-time identity detection in resource-constrained environments. These lightweight models significantly reduce computational load and memory consumption while maintaining detection accuracy, enabling the algorithm to run efficiently on embedded or mobile devices.
[0119] It should be noted that the identity detection algorithms established in the above ways cannot achieve optimal performance solely relying on limited labeled data, especially in the complex and ever-changing environment of power grid construction sites. Therefore, this application designs a novel first identity detection algorithm.
[0120] In this embodiment, the first identity detection algorithm includes an identity detection network, which includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module.
[0121] In this embodiment of the application, the first information sensing module includes:
[0122] Compress the second data by increasing the number of channels;
[0123] The second data, after channel compression, undergoes segmentation, concatenation, and several fusion convolution and self-attention mechanism operations.
[0124] In the embodiments of this application, such as Figure 3 The diagram shows the overall structure of the Super-Resolution Information Perception Module (SRIC), which is also the overall structure of the first information perception module.
[0125] The input staff identity feature map, i.e., the second data, is compressed using a Conv layer with a 1×1 convolutional kernel to adjust the spatial dimension of the input, thus obtaining the first information-perceived feature map. ; then, The object-aware mixer CAMixer, through the fusion of convolution and self-attention mechanisms, Simple computational information is allocated to convolution, while complex computational information is allocated to the self-attention mechanism. The feature map output by CAMixer is sequentially input into a BN layer and a ReLU layer to obtain the second information-aware feature map. .
[0126] Furthermore, Segmentation is performed along the channel dimension to obtain third-information perception feature maps. and the fourth information perception feature map , and The number of input channels is Half of it.
[0127] Furthermore, The input is fed into a target-aware mixer CAMixer, and then sequentially fed into a BN layer and a ReLU layer to obtain the fifth information-aware feature map. .
[0128] Furthermore, The input is fed into an ELA (Efficient Local Attention) module. Within the ELA module, the input is first processed... Perform adaptive pooling in both horizontal and vertical directions to obtain and ,Will and The horizontal and vertical positional information of the feature map are respectively input through one-dimensional convolution; then, group normalization and sigmoid activation functions are applied to obtain the feature map with enhanced positional information in the vertical direction. Feature maps with enhanced positional information in the horizontal direction .
[0129] (7)
[0130] (8)
[0131] (9)
[0132] (10)
[0133] In equation (1), for The feature map obtained after performing adaptive pooling in the horizontal direction is shown in equation (2). for The feature map obtained after performing adaptive pooling in the vertical direction. Indicates altitude, To represent the width, in equations (3) and (4), and Represents one-dimensional convolution. Indicates group normalization, This represents the sigmoid function.
[0134] Furthermore, , and Multiplying them together yields the sixth information perceptual feature map. .
[0135] (11)
[0136] Furthermore, Perform 3×3 convolution to extract features and obtain the seventh information perception feature map. .
[0137] Furthermore, the features , and By concatenating the features along the channel dimension, the eighth information perception feature map is obtained. , The number of channels is , and The sum of the number of channels.
[0138] Furthermore, the features Perform a 3×3 convolution, then sequentially input it into a BN layer and a ReLU layer to obtain the ninth information-aware feature map. .
[0139] It should be noted that the first information perception module can enhance the perception of worker features under super-resolution pixels in construction scenarios. By adaptively and reasonably allocating feature map data, this module can effectively fit the fine position information within the super-resolution feature map while ensuring detection efficiency.
[0140] In this embodiment of the application, the second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the location regression of the target person, and the second branch is used to calculate the identity classification of the target person.
[0141] In this embodiment of the application, the first branch and the second branch include:
[0142] Both the first and second branches are used to perform several convolution operations on the feature maps that have been input into the second dual-branch identity detection head module and have undergone channel compression.
[0143] The first branch includes several convolution operations, including target regression convolution operations.
[0144] The second branch includes several convolution operations, including object classification convolution operations.
[0145] In the embodiments of this application, such as Figure 4 The diagram shown is the structure of the DDB-Head module, which is the second dual-branch identity detection header module in this application, wherein:
[0146] The input staff identity feature map is compressed using a Conv layer with a 1×1 convolutional kernel to adjust the spatial dimension of the input, resulting in the first bi-branch detection feature map. ,Will The data is then input into the first and second branches to perform regression and classification of staff identities.
[0147] First branch: The input is fed into a 3×3 deformable convolutional (Deformable Conv) layer, and then the output of the Deformable Conv layer is fed into a ReLU layer to obtain the second dual-branch detection feature map. ;Will The input is fed into a 3×3 Deformable Conv layer, and then into a BN layer to obtain the third bi-branch detection feature map. Finally, The input undergoes a convolution for target regression (Locate), resulting in an output with 2 channels (representing the top-left and bottom-right coordinates of the network's predicted bounding box). The output is the fourth dual-branch detection feature map. .
[0148] Second branch: The input is fed into a 3×3 deformable convolutional (Deformable Conv) layer, and the output of the Deformable Conv layer is then fed into a ReLU layer to obtain the fifth bi-branch detection feature map. ,Will The input is fed into a 3×3 deformable convolutional layer, and then into a batch normalization (BN) layer to obtain the sixth bi-branch detection feature map. Finally, Input a convolutional function for object classification (Sort), output a seventh-branch detection feature map. .
[0149] Furthermore, and Concat along the channel dimension to obtain the eighth bi-branch detection feature map. .
[0150] It should be noted that the second dual-branch identity detection head module can enhance the identity detection effect of workers from multiple angles in the feature images of the construction site. This detection head can adapt to the dynamic changes of workers and effectively eliminate the interference of background noise at the construction site, thereby improving the detection accuracy of the model.
[0151] In this embodiment, the third-space pyramid pooling module is used for downsampling in the identity detection network;
[0152] The third-space pyramid pooling module includes channel number compression, splicing, normalization, and several max pooling operations.
[0153] In an optional embodiment, such as Figure 5 The diagram shown is a schematic of the overall structure of DDSPPF (Deformable DAT Spatial Pyramid Pooling Factorization), which is the schematic diagram of the third-space pyramid pooling module structure in this application, wherein:
[0154] The input worker identity feature map is compressed using a Conv layer with a 1×1 convolutional kernel to adjust the spatial dimension of the input, resulting in the first spatial pyramid pooling feature map. ,Will The input is fed into the Deformable Conv layer, and then into the BN layer to obtain the second-space pyramid pooling feature map. .
[0155] Furthermore, The input is fed into the Deformable Attention (DAT) (Global Attention Mechanism) module, and the calculation process for this attention is as follows:
[0156] (12)
[0157] in, This represents the output projection matrix of the deformable attention DAT. The first representing the deformable attention The attention output of each head is calculated as follows:
[0158]
[0159] in, , and Representing the first The query vector, key-value vector, and value vector under each header. The dimension representing each head, M represents the total number of heads, calculated by dividing the number of input channels by the number of heads. represent The function, assuming the input is { The calculation formula is as follows:
[0160]
[0161] Furthermore, after the above calculation process, we obtain Output third-space pyramid pooling feature map .
[0162] Furthermore, The input is fed into a max pooling layer to obtain a fourth-space pyramid pooling feature map. Then, two more max-pooling layers are input sequentially to obtain the fifth-space pyramid pooling feature map. and the sixth space pyramid pooling feature map .
[0163] Furthermore, , , and By stitching along the channel dimension, the seventh-space pyramid pooled feature map is obtained. .
[0164] Furthermore, The input is fed into the Deformable Conv layer, and then into the batch normalization layer to obtain the DDSPPF output eighth-space pyramid pooling feature map. .
[0165] It should be noted that when the third-space pyramid pooling module performs downsampling in the network, it can enhance the trainable offset and the fitting ability to the worker identity target. Secondly, through this structure, the generalization of the network can be improved and the global interaction features can be amplified, thereby avoiding the loss of key information and the distortion of the feature space.
[0166] In this embodiment, based on the first information sensing module, the second dual-branch identity detection head module, and the third spatial pyramid pooling module, a personnel identity detection network CSII-Net is designed to realize the identity detection of workers at power grid construction sites. Figure 2 As shown.
[0167] This application first inputs the initial image into a convolutional layer (Conv) and transforms the number of channels to adjust the spatial dimension of the input, thereby obtaining the first identity detection feature map. Then The second identity detection feature map is obtained by sequentially inputting the convolutional layer and the batch normalization layer (BN). .
[0168] Next, this application designs a two-layer feature extraction structure in this section. These two layers are a combination of an SRIC module, a convolutional layer, and an SRIC module, and a combination of an SRIC module, a convolutional layer, an SRIC module, and a DDSPPF. After inputting into the SRIC module, convolutional layer, and combined SRIC module layer, the third identity detection feature map is obtained after these two layer structures. ; then The input is combined with the SRIC module, convolutional layer, SRIC module and DDSPPF to obtain the fourth identity detection feature map. .Will The input is processed in two parts. The first part is input into the SRIC module to obtain the fifth identity detection feature map. In the second part, Upsampling is performed to obtain the output sixth identity detection feature map. Then, and By stitching along the channel dimension, we obtain the seventh identity detection feature map of the stitched result. Next The input is processed in the SRIC module to extract features, resulting in the eighth identity detection feature map. Subsequently, Inputting the SRIC module and the SRIC module sequentially yields the ninth identity detection feature map. Next and By splicing the images together, we obtain the tenth identity detection feature map. Finally, output the processing results of the upper and lower branches. and These correspond to the two-branch feature processing results for small and large targets in the image, respectively.
[0169] Next, a DDB-Head detection head was designed, and the DDB-Head detection head was connected to... and The identification process is then completed in both the upper and lower DDB-Head detection heads. and Then, the detection head recognition results of the two branches are output. and After that, and The process involves merging the detection results from the two branches of the small and large target detection heads to ultimately obtain the worker identity target detection result from the CSII-Net network. The output contains a tensor of prediction information, with each row corresponding to a prediction, including bounding box coordinates, class label, and confidence score.
[0170] It should be noted that the training and validation data are input into the CSII-Net network for training until the network loss function converges.
[0171] Initialize the various parameters and hyperparameters of CSII-Net, including the number of training iterations, batch size, optimizer, learning rate type, and initial learning rate.
[0172] After completing the above preparations, input the training and validation data into CSII-Net for training until the network loss function converges. In each training round, the network uses employee identity data from the training set in batches with preset values to update the loss value and various parameters. After each training round, the network uses employee identity data from the validation set in batches with preset values to verify the effectiveness of each training round. Training of the CSII-Net network ends when the loss value tends to converge.
[0173] Furthermore, after the CSII-Net network is trained, test data is used. To conduct the test, the data is input into the optimally trained CSII-Net network to evaluate its performance. The metric for evaluating the CSII-Net network's test results is accuracy, which is the percentage of staff members correctly identified from the total test data. The network test is complete when the accuracy reaches the user's expected accuracy threshold.
[0174] It should be noted that establishing a first identity detection algorithm and using the second set of data as both the training and validation sets for the first algorithm can fully utilize existing labeled data, improving the generalization ability and accuracy of the identity detection model. By learning from this data, the first identity detection algorithm can grasp the identity characteristics of construction site workers, thus achieving fast and accurate identity recognition in subsequent practical applications. Furthermore, using the second set of data as both the training and validation sets can help adjust and optimize algorithm parameters, further improving detection performance.
[0175] S103, perform identity verification of the target power grid construction site personnel based on the first identity detection algorithm after verification.
[0176] In practical applications, the system is deployed by inputting worker identification images into CSII-Net, which outputs detection results for worker identities. This result contains a tensor of detection information, with each row corresponding to a prediction, including worker bounding box coordinates, worker category label, and confidence score. Worker identification images can be obtained using various imaging devices or extracted as needed from videos of power grid construction sites.
[0177] In summary, this invention proposes a method for identifying workers at power grid construction sites. The method involves acquiring first data of the target power construction site, performing a first preprocessing step to obtain second data, establishing a first identity detection algorithm, and using the second data as both the training and validation set for the first identity detection algorithm. The method then performs identity detection of the workers at the target power grid construction site based on the validated first identity detection algorithm. This invention establishes a highly generalizable, accurate, and automated network for identifying construction site workers, enabling efficient identification of their identities. After training, the network designed in this invention can be deployed at construction sites or substations along various transmission lines to achieve real-time monitoring of the identities of construction site workers.
[0178] Example 2, in a preferred embodiment, employs as follows Figures 2-5 The structure shown is applied in a real-world scenario, and the specific steps are as follows: Images containing workers at a power construction site are acquired using imaging equipment (such as drones, fixed equipment, etc.), and the identity data is filtered. To ensure sufficient data volume, data containing workers of various identities are filtered, and 300 images containing the identities of various workers are cropped to 640×640 pixels. Subsequently, the LabelImg image annotation tool was used to annotate the category and coordinate bounding box information of all staff identity images: a pre-set VOC annotation format was used to select and mark the staff identity area, creating bounding boxes and category labels. For the category labels, this application divides the staff identity categories into {management personnel (red), technical personnel (blue), construction and inspection personnel (yellow), supervision personnel (white), and other personnel (other colors or no label)}, corresponding to category indices. After annotation is completed, the annotated image data is generated. One-to-one correspondence annotation file The data format is .xml. The labeled dataset... , Perform data augmentation operations, including data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition. Each of these operations has a significant impact on... and Each data augmentation operation will generate twice the amount of augmented data, ultimately resulting in a total dataset including the original data and all the augmented data. and .
[0179] based on and Data augmentation is performed to simulate scenarios, and the following simulations are conducted for rainy and foggy weather.
[0180] right and Perform data augmentation operations to simulate rainy scenes, and obtain the total dataset after the rainy scene simulation. , .
[0181] right and Perform data augmentation operations to simulate foggy scenes, and obtain the total dataset after the foggy scene simulation. , .
[0182] Will , , , and , Data merging is performed to obtain the total augmented dataset. , .
[0183] After the preprocessing of the operator dataset is completed, To divide the dataset, this application randomly divides the training set, validation set, and test set in a ratio of 7:2:1, resulting in 5040 training data sets, 1440 validation data sets, and 720 test data sets.
[0184] The detection network designed in this application is used to train the dataset after data partitioning. First, the various parameters and hyperparameters of the CSII-Net network for detecting the identity of construction site workers are initialized. This embodiment explains several important parameters in network training: the corresponding training batch size is initialized according to the hardware environment, the optimizer is initialized to optimizer, the batch size is set to 4, the number of training rounds is initialized to 200, the learning rate type is cosine learning rate, and the initial learning rate is 0.001. These parameters need to be adjusted according to the effect of multiple training sessions until the network converges.
[0185] After initializing the basic parameters for network training, model training begins. Within this network, the initial image is first input into a convolutional layer (Conv) and the number of channels is transformed to adjust the spatial dimension of the input, resulting in the first identity detection feature map. Then The second identity detection feature map is obtained by sequentially inputting the convolutional layer and the batch normalization layer (BN). Next is the core of the model. This application designs a two-layer feature extraction structure in this part, namely, a combination of an SRIC module, a convolutional layer, and an SRIC module, and a combination of an SRIC module, a convolutional layer, an SRIC module, and a DDSPPF module. After inputting these two layers of structures in sequence, the feature extraction results, specifically the third identity detection feature map, are output respectively. and Subsequently, The input is processed in two parts. The first part is input into the SRIC module to obtain the fifth identity detection feature map. In the second part, Upsampling is performed to obtain the output sixth identity detection feature map. Then, and By stitching along the channel dimension, we obtain the seventh identity detection feature map of the stitched result. Next The input is processed in the SRIC module to extract features, resulting in the eighth identity detection feature map. Subsequently, Inputting the SRIC module and the SRIC module sequentially yields the ninth identity detection feature map. Next and By splicing the images together, we obtain the tenth identity detection feature map. Finally, output the processing results of the upper and lower branches. and These correspond to the two-branch feature processing results for small and large targets in the image, respectively. Next, this application designs the detection head part of the network, connecting the designed DDB-Head detection head to... and The identification process is then completed in both the upper and lower DDB-Head detection heads. and Then, the detection head recognition results of the two branches are output. and Subsequently, this application will and The process involves merging the detection results from the two-branch detection heads for small and large targets to ultimately obtain the worker identity target detection result from the CSII-Net network. The result contains a tensor of detection information, with each row corresponding to a prediction, including the worker bounding box coordinates, worker category label, and confidence score.
[0186] CSII-Net training updates network parameters through backpropagation using a loss function. In this invention, the bounding box coordinate loss for the worker identity target uses DFL loss, the bounding box confidence loss uses WIoU loss, and the classification loss uses cross-entropy loss. These loss functions are updated epoch-by-epoch during gradient backpropagation during network training, allowing the network to gradually converge and improve detection accuracy with each training epoch. After each training epoch, the network uses the partitioned validation data to evaluate its accuracy. Recall rate The mean accuracy (mAP) is used as a criterion for evaluating the effectiveness of the training. After the CSII-Net of this invention is trained, 720 pre-divided worker identity image datasets are used as test data. This data is input into the optimal network parameters of the trained CSII-Net to test the effect.
[0187] In practical applications, pilot applications were deployed by inputting staff identity image data into CSII-Net and outputting the extraction results of the staff identity image data. The output results contain tensors of prediction information, with each row corresponding to a prediction, including bounding box coordinates, class label and confidence score.
[0188] Example 3: This example also provides a system for detecting the identity of workers at power grid construction sites, including:
[0189] The data acquisition and processing module is used to acquire the first data of the target power construction site and perform the first preprocessing on the first data to obtain the second data;
[0190] The first preprocessing step includes data filtering, data annotation, data augmentation, and data partitioning.
[0191] The algorithm establishment module is used to establish the first identity detection algorithm, and the second data is used as the training set and verification set of the first identity detection algorithm.
[0192] The first identity detection algorithm includes an identity detection network, which includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module.
[0193] The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the location regression of the target person, and the second branch is used to calculate the identity classification of the target person.
[0194] The detection module is used to detect the identity of the workers at the target power grid construction site based on the first identity detection algorithm after verification.
[0195] The above-mentioned unit modules can be embedded in the processor of the computer device in hardware form or independent of it, or they can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.
[0196] This embodiment also provides a computer device, which may be a terminal, and its internal structure diagram may be as follows. Figure 6 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a method for identifying workers at a power grid construction site. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0197] This embodiment also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, it performs the following steps:
[0198] First data of the target power construction site is obtained, and first preprocessing is performed on the first data to obtain second data;
[0199] The first preprocessing step includes data filtering, data annotation, data augmentation, and data partitioning.
[0200] A first identity detection algorithm is established, and the second data is used as the training set and validation set for the first identity detection algorithm.
[0201] The first identity detection algorithm includes an identity detection network, which includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module.
[0202] The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the location regression of the target person, and the second branch is used to calculate the identity classification of the target person.
[0203] The identity of the workers at the target power grid construction site is verified using the first identity detection algorithm after the verification is completed.
[0204] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
[0205] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages.
[0206] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0207] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0208] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0209] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0210] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for identifying workers at power grid construction sites, characterized in that, include: First data of the target power construction site is acquired, and the first data is preprocessed to obtain second data; The first data is images of workers at the target power construction site; The first preprocessing includes data filtering, data annotation, data augmentation, and data partitioning; A first identity detection algorithm is established, and the second data is used as the training set and verification set for the first identity detection algorithm. The first identity detection algorithm includes an identity detection network, which includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module. The first information sensing module is a super-resolution information sensing module (SRIC), and its overall structure is as follows: The input staff identity feature map, i.e., the second data, is compressed using a Conv layer with a 1×1 convolutional kernel to adjust the spatial dimension of the input, thus obtaining the first information-perceived feature map. ; then, The object-aware mixer CAMixer, through the fusion of convolution and self-attention mechanisms, Simple computational information is assigned to convolution, while complex computational information is assigned to the self-attention mechanism. The feature map output by CAMixer is sequentially input into a BN layer and a ReLU layer to obtain the second information-aware feature map. ; Will Segmentation is performed along the channel dimension to obtain third-information perception feature maps. and the fourth information perception feature map , and The number of input channels is Half of; Will The input is fed into a target-aware mixer CAMixer, and then sequentially fed into a BN layer and a ReLU layer to obtain the fifth information-aware feature map. ; Will The input is fed into an ELA attention mechanism module, where the ELA module first performs... Perform adaptive pooling in both horizontal and vertical directions to obtain and ,Will and The horizontal and vertical positional information of the feature map are respectively input through one-dimensional convolution; then, group normalization and sigmoid activation functions are applied to obtain the feature map with enhanced positional information in the vertical direction. Feature maps with enhanced positional information in the horizontal direction ; ; ; ; ; in, for The feature map obtained after performing adaptive pooling in the horizontal direction. for The feature map obtained after performing adaptive pooling in the vertical direction. Indicates altitude, Indicates width, and Represents one-dimensional convolution. Indicates group normalization, Represents the sigmoid function; Will , and Multiplying them together yields the sixth information perceptual feature map. ; ; Will 3×3 convolution is performed to extract features, resulting in the seventh information perception feature map. ; Features , and The features are concatenated along the channel dimension to obtain the eighth information perception feature map. , The number of channels is , and The sum of the number of channels; feature map Perform a 3×3 convolution, then sequentially input it into a BN layer and a ReLU layer to obtain the ninth information-aware feature map. ; The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the location regression of the target person, and the second branch is used to calculate the identity classification of the target person. The number of channels in the staff identity feature map input to the second dual-branch identity detection head module is compressed to obtain the first dual-branch detection feature map; The first branch is used to input the first bi-branch detection feature map into a deformable convolutional layer, and then into an activation function layer to obtain a second bi-branch detection feature map; the second bi-branch detection feature map is input into a deformable convolutional layer, and then into a normalization layer to obtain a third bi-branch detection feature map; the third bi-branch detection feature map is input into a convolution for target regression to obtain a fourth bi-branch detection feature map. The second branch is used to input the first bi-branch detection feature map into a deformable convolutional layer, and then into an activation function layer to obtain a fifth bi-branch detection feature map. The fifth bi-branch detection feature map is then input into a deformable convolutional layer, and then into a normalization layer to obtain a sixth bi-branch detection feature map. The sixth bi-branch detection feature map is then input into a convolutional layer for target classification to output a seventh bi-branch detection feature map. The identity of the workers at the target power grid construction site is verified using the first identity detection algorithm after the verification is completed.
2. The method for detecting the identity of workers at power grid construction sites as described in claim 1, characterized in that, The first branch and the second branch include: Both the first branch and the second branch are used to perform several convolution operations on the feature map that has been input into the second dual-branch identity detection head module and has undergone channel compression. The first branch includes several convolution operations, including target regression convolution operations. The second branch includes several convolution operations, including target classification convolution operations.
3. The method for detecting the identity of workers at power grid construction sites as described in claim 2, characterized in that, The third spatial pyramid pooling module is used to perform downsampling in the identity detection network; The third spatial pyramid pooling module includes channel number compression operation, splicing operation, normalization operation, and several max pooling operations.
4. The method for detecting the identity of workers at power grid construction sites as described in claim 3, characterized in that, The first preprocessing of the first data includes: The first data consists of images of workers at the target power construction site; The first data is then filtered and labeled; The first data after filtering and labeling is augmented to obtain the second data.
5. The method for detecting the identity of workers at power grid construction sites as described in claim 4, characterized in that, The data augmentation of the first data after filtering and labeling includes: Based on the safety helmet categories of construction workers at the work site in the first data, image annotation tools were used to annotate the category and coordinate frame information of the personnel targets in all worker images; The categories are divided into management personnel, technical personnel, construction and inspection personnel, supervision personnel, and other personnel. After obtaining the filtered and labeled small sample dataset, data augmentation is performed on any staff image in the small sample dataset using data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition techniques. All data-augmented images are then combined with the small sample dataset to form an enhanced staff identity detection dataset. The enhanced staff identity detection dataset is augmented with scenario simulation to obtain the second dataset.
6. The method for detecting the identity of workers at power grid construction sites as described in claim 5, characterized in that, The scenarios include at least one or more of the following: rainy weather scenarios and foggy weather scenarios.
7. A system for detecting the identity of workers at a power grid construction site, using the method described in any one of claims 1 to 6, characterized in that, include: The data acquisition and processing module is used to acquire first data from the target power construction site and perform first preprocessing on the first data to obtain second data. The first preprocessing includes data filtering, data annotation, data augmentation, and data partitioning; The algorithm establishment module is used to establish a first identity detection algorithm, and uses the second data as the training set and verification set of the first identity detection algorithm; The first identity detection algorithm includes an identity detection network, which includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module. The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the location regression of the target person, and the second branch is used to calculate the identity classification of the target person. The detection module is used to detect the identity of the workers at the target power grid construction site based on the first identity detection algorithm after verification.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Substation infrared image mouse identification method based on deep learning
CN118314532A
Industrial pure iron metallographic specimen defect detection optimization method based on YOLOv8obb
CN119831964A