Power grid construction site worker identity detection method, system, equipment and medium
By obtaining data at the power construction site for pre-processing and establishing an identity detection algorithm, the problems of inefficiency and insufficient accuracy in the existing technology are solved, and efficient and accurate staff identity detection is achieved to adapt to the complex and changeable construction site environment.
Patent Information
- Application Number
- CN202510828717.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
It is difficult to achieve efficient and accurate staff identity detection at power construction sites in the prior art. Especially in complex and changeable environments, traditional methods are inefficient and susceptible to human interference, making it difficult to achieve comprehensive and reliable identity management.
An identity detection method for workers on the power grid construction site is adopted. By obtaining on-site data for pre-processing, an identity detection algorithm including an information perception module, a dual-branch identity detection head module and a spatial pyramid pooling module is established, and identity detection is performed using deep learning technology.
It realizes efficient and accurate automated identity detection in complex environments, improves the generalization and detection accuracy of the model, and can identify the identity of construction site staff in real time.
Smart Images

Figure CN120356243A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a method, system, device and medium for detecting the identity of construction site workers in the power grid construction site. Background Art
[0002] In the power construction site, accurately and efficiently identifying the identity of workers is the key to ensuring operation safety and improving work efficiency. Traditional identity recognition methods, such as manually verifying identity documents or relying on fixed surveillance camera systems, have many limitations. For example, due to the large mobility of personnel and the changing scenarios, traditional methods often have difficulty in quickly and accurately detecting the identity of workers.
[0003] In environments such as construction sites, the identity verification and safety supervision of workers are crucial. However, these environments usually have complex, changeable environmental conditions and dense personnel, which pose great challenges to identity detection. At present, some scenarios rely on manual verification or simple access control systems, but these methods are not only inefficient but also easily interfered by human factors and are difficult to achieve comprehensive and reliable identity management.
[0004] At the current stage, in the power system, the identity detection of workers at the construction sites along the transmission lines or in substations requires manual screening of image data and inspection, or on-site patrols, which require a large amount of time cost and professional manpower. Summary of the Invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] Therefore, the present invention provides a method, system, device and medium for detecting the identity of construction site workers in the power grid, which can solve the problems of low efficiency and insufficient accuracy in the identity detection of construction site workers in the power grid.
[0007] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for detecting the identity of construction site workers in the power grid, including: Obtaining first data of a target power construction site and performing first preprocessing on the first data to obtain second data; The first preprocessing includes data screening, data annotation, data augmentation and data partitioning; Establishing a first identity detection algorithm and using the second data as the training set and validation set of the first identity detection algorithm; The first identity detection algorithm includes an identity detection network, and the identity detection network includes a first information perception module, a second dual-branch identity detection head module and a third spatial pyramid pooling module; The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the position regression of the target person, and the second branch is used to calculate the identity classification of the target person. Perform identity detection on the staff at the target power grid construction site according to the first identity detection algorithm after verification.
[0008] As a preferred solution of the method for detecting the identity of the staff at the power grid construction site according to the present invention, wherein: the first branch and the second branch include: Both the first branch and the second branch are used to perform a number of convolution operations on the feature map that is input into the second dual-branch identity detection head module and has been channel-compressed. The number of convolution operations of the first branch includes target regression convolution operations. The number of convolution operations of the second branch includes target classification convolution operations.
[0009] As a preferred solution of the method for detecting the identity of the staff at the power grid construction site according to the present invention, wherein: the third spatial pyramid pooling module is used for downsampling in the identity detection network. The third spatial pyramid pooling module includes channel number compression operations, splicing operations, normalization operations, and a number of max pooling operations.
[0010] As a preferred solution of the method for detecting the identity of the staff at the power grid construction site according to the present invention, wherein: the first information perception module includes: Compress the number of channels of the second data. Perform segmentation operations, splicing operations, and a number of fusion convolution and self-attention mechanism operations on the second data after channel number compression.
[0011] As a preferred solution of the method for detecting the identity of the staff at the power grid construction site according to the present invention, wherein: the first preprocessing of the first data includes: The first data is an image containing operating personnel at the target power construction site. Screen and label the first data. Perform data augmentation on the first data after screening and labeling to obtain second data.
[0012] As a preferred solution of the method for detecting the identity of the staff at the power grid construction site according to the present invention, wherein: the data augmentation of the first data after screening and labeling includes: According to the safety helmet category of the construction personnel at the operation site in the first data, use an image annotation tool to label the category and coordinate frame information of the personnel targets in all the staff images. Among them, the categories are divided into management personnel, technical personnel, construction and inspection personnel, supervision personnel, and other personnel; A small sample data set after screening and annotation is obtained. For any staff image in the small sample data set, data augmentation is performed using techniques such as data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition. All the images after data augmentation and the small sample data set together constitute an enhanced staff identity detection data set; Data augmentation of scene simulation is performed on the enhanced staff identity detection data set to obtain a second data.
[0013] As a preferred solution of the method for detecting the identity of staff at the power grid construction site according to the present invention, wherein: the scene includes at least one or more of the following: rainy day environment scene and foggy day environment scene.
[0014] In a second aspect, the present invention provides a system for detecting the identity of staff at a power grid construction site, including: A data acquisition and processing module for acquiring first data of a target power construction site and performing a first preprocessing on the first data to obtain second data; The first preprocessing includes data screening, data annotation, data augmentation, and data partitioning; An algorithm establishment module for establishing a first identity detection algorithm and using the second data as the training set and verification set of the first identity detection algorithm; The first identity detection algorithm includes an identity detection network, and the identity detection network includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module; The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the position regression of the target person, and the second branch is used to calculate the identity classification of the target person; A detection module for detecting the identity of staff at the target power grid construction site according to the first identity detection algorithm after verification is completed.
[0015] In a third aspect, the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described above are implemented.
[0016] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described above are implemented.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a method, system, device and medium for detecting the identity of workers at the power grid construction site. First data of the target power grid construction site is acquired and subjected to first preprocessing to obtain second data. A first identity detection algorithm is established, and the second data is used as the training set and verification set of the first identity detection algorithm. According to the first identity detection algorithm after verification, the identity of workers at the target power grid construction site is detected. The present invention establishes an identity detection network for workers at the construction site with strong generalization ability, high accuracy and automation, which can efficiently identify the identity of workers at the construction site. The network designed by the present invention can be deployed at the construction sites or substations along each transmission line after training to realize real-time observation of the identity of workers at the construction site.
[0018] Specifically, the first information perception module can enhance the perception ability of the characteristics of workers under super-resolution pixels in the construction scenario. By adaptively and reasonably allocating the feature map data, while ensuring the detection efficiency, this module can effectively fit the fine position information in the super-resolution feature map. The second dual-branch identity detection head module can enhance the identity detection effect of workers at multiple angles in the construction site feature image. This detection head can adapt to the dynamic changes of workers and effectively exclude the interference of background noise at the construction site, thereby improving the detection accuracy of the model. When downsampling in the network, the third spatial pyramid pooling module can enhance the trainable offset and the fitting ability for the identity target of workers. Secondly, through this structure, the generalization ability of the network can be improved, the global interaction features can be amplified, and thus the loss of key information and the distortion of the feature space can be avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 It is a flowchart of a method for detecting the identity of workers at the power grid construction site provided by an embodiment of the present invention.
[0021] Figure 2 It is a schematic diagram of the construction site worker identity detection network CSII-Net for a method for detecting the identity of workers at the power grid construction site provided by an embodiment of the present invention.
[0022] Figure 3 It is a structural diagram of the SRIC module for a method for detecting the identity of workers at the power grid construction site provided by an embodiment of the present invention.
[0023] Figure 4 The DDB-Head module structure diagram of a method for detecting the identity of on-site power grid construction workers provided in an embodiment of the present invention.
[0024] Figure 5 The DDSPPF module structure diagram of a method for detecting the identity of on-site power grid construction workers provided in an embodiment of the present invention.
[0025] Figure 6 The internal structure diagram of a computer device for a method for detecting the identity of on-site power grid construction workers provided in an embodiment of the present invention. Detailed implementation manners
[0026] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] Example 1, referring to Figures 1 - 6 , which is the first embodiment of the present invention. This embodiment provides a method for detecting the identity of on-site power grid construction workers, including: Before introducing the embodiments of the present application in detail, for clarity, some related concepts will be explained first.
[0028] Target-aware mixer CAMixer: It is a prior art. CAMixer uses a learnable predictor to generate multiple guides, including offsets for window warping, masks for classifying windows, and convolutional attention for endowing convolutions with dynamic properties, which can adaptively adjust the attention to include more useful textures and improve the representation ability of convolutions. Other operations such as convolution, batch normalization, ReLU activation function, and splicing operations are common operations in the deep learning industry.
[0029] Normalization BN: It refers to a technique for normalizing the input data of each mini-batch during the training process of a neural network. It normalizes the input data to a distribution with a mean of 0 and a variance of 1 (or close to this distribution) to reduce internal covariate shift, thereby accelerating the training of the neural network and improving its performance.
[0030] ReLU (Rectified Linear Unit): It refers to an activation function, which is a commonly used non-linear activation function in deep learning.
[0031] ELA: is a prior art, which is an Efficient Local Attention (ELA) method. This method can accurately locate the region of interest without dimensionality reduction by effectively encoding two one-dimensional position feature maps, while allowing for lightweight implementation.
[0032] Deformable Conv layer: is a prior art. It enables the network to adapt to the shape, scale, and pose changes of objects by dynamically adjusting the sampling positions of the convolution kernels, thereby improving the robustness of the model in complex scenarios.
[0033] Conv: is a one-dimensional convolution operation. Its convolution kernel slides along the time axis or sequence axis of the feature map to extract local features in the sequence data. The convolution operation acts on all input channels simultaneously and achieves cross-channel information integration through shared weights. It can efficiently capture local patterns in the sequence data while retaining global context information. It is a prior art.
[0034] SRIC (Super Resolution Information Capture): a lightweight super-resolution information perception module, which is the first information perception module in this application. It is used to enhance the perception of the characteristics of workers under super-resolution pixels in the construction scene. By adaptively allocating feature map data, it strengthens the detection efficiency on the premise of ensuring fine position fitting within the super-resolution feature map.
[0035] DDB-Head (Deformable Double Branch Head): a dynamic double-branch identity detection head, which is the second double-branch identity detection head module in this application. It is used to enhance the identity detection effect of workers at multiple angles in the feature images of the construction site and adapt to the dynamic changes of workers. At the same time, it can better exclude the interference of background noise at the construction site and strengthen the detection accuracy of the model.
[0036] DDSPPF (Deformable DAT Spatial Pyramid Pooling Factorization): a spatial pyramid pooling structure, which is the third spatial pyramid pooling module in this application. It is used to perform downsampling in the network, enhance the trainable offset and the fitting ability for the worker identity target, improve the generalization of the network, amplify the global interaction features, and avoid the loss of key information and feature space distortion.
[0037] CSII-Net (Construction-Site Staff Identity Identification Net): The construction site staff identity detection network, that is, the identity detection network in this application, is used to identify the staff targets at the power construction site.
[0038] In the existing related technologies, there are some problems. For example, the environment at the power construction site is complex, resulting in difficult image recognition; or the identity tags of the staff are inaccurate, affecting the accuracy of identity detection, etc.
[0039] This application provides a method that can effectively solve the above-mentioned problems. Next, multiple embodiments will be combined to elaborate in detail on how to implement the method for detecting the identity of the staff at the power grid construction site; Figure 1 A method for detecting the identity of the staff at the power grid construction site is shown, including: S101, obtain the first data of the target power construction site, and perform a first preprocessing on the first data to obtain the second data; In an optional embodiment, the target power construction site can be any power construction site that needs to be monitored. For example, the construction sites along the transmission line, the substation construction site, the distribution room construction site, etc. The specific scenarios are not limited in this application.
[0040] In an optional embodiment, the first data is an image containing operating personnel at the target power construction site, which can be obtained through image acquisition devices such as cameras and drones.
[0041] It should be noted that the first preprocessing is to perform a series of operations on the first data to prepare the data so that it is suitable for subsequent identity detection algorithm training and verification. The specific steps of the first preprocessing may include but are not limited to operations such as image scaling, cropping, grayscaling, denoising, contrast enhancement, etc., as well as data augmentation techniques such as rotation, flipping, translation, etc., to improve the generalization ability of the model. These preprocessing steps are designed to eliminate redundant information in the image and highlight key features, so as to ensure that the subsequent identity detection algorithm can accurately and efficiently identify the target personnel. Through a carefully designed preprocessing process, the accuracy and robustness of identity detection can be significantly improved.
[0042] In the embodiment of this application, the first preprocessing includes data screening, data annotation, data augmentation, and data partitioning.
[0043] In the embodiment of this application, performing the first preprocessing on the first data includes: The first data is an image containing operating personnel at the target power construction site; Screen and label the first data; Perform data augmentation on the first data after screening and annotation to obtain the second data.
[0044] In an alternative embodiment, screening and annotation can be performed by professionals to ensure the accuracy and reliability of the data. The screening process may involve removing blurred, occluded, or low-quality images, while annotation may include assigning a unique identity label to each worker in the image and annotating their location information. Such processing steps are crucial for training a high-precision identity detection algorithm.
[0045] In another alternative embodiment, screening and annotation may be performed in a semi-automatic or fully automatic manner to improve work efficiency. In the semi-automatic mode, professionals can use image annotation tools to manually annotate the workers and their identity labels in the images. In the fully automatic mode, advanced computer vision technologies, such as deep learning models, can be used to automatically detect and annotate the workers.
[0046] In an alternative embodiment, data augmentation can employ a variety of techniques, including but not limited to color jittering, brightness adjustment, edge enhancement, and blurring, to increase the diversity and complexity of the dataset. These augmentation techniques can simulate different lighting conditions, weather conditions, and shooting angles, thereby improving the adaptability and robustness of the model in actual application scenarios. By training on the augmented dataset, the identity detection algorithm can better learn the characteristics of workers in various environments, thereby enhancing the accuracy and stability of identity detection.
[0047] In the embodiment of the present application, performing data augmentation on the first data after screening and annotation includes: According to the safety helmet categories of the construction workers at the job site in the first data, use an image annotation tool to annotate the category and coordinate box information of the personnel targets in all the worker images; Among them, the categories are divided into management personnel, technical personnel, construction and inspection personnel, supervision personnel, and other personnel; Obtain the small sample dataset after screening and annotation. For any worker image in the small sample dataset, perform data augmentation using data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition techniques, and together with all the images after data augmentation and the small sample dataset, form an enhanced worker identity detection dataset; Perform data augmentation by simulating scenarios on the enhanced worker identity detection dataset to obtain the second data.
[0048] Specifically, images of workers at the power construction site are obtained through a photographing device (such as a drone, a fixed device, etc.), and the identity detection images are screened. After the screening of the identity detection images is completed, according to the safety helmet categories of the construction workers at the work site, an image annotation tool is used to annotate the category and coordinate box information of the personnel targets in all the worker images. The categories are divided into management personnel (red safety helmets), technical personnel (blue safety helmets), construction and inspection personnel (yellow safety helmets), supervision personnel (white safety helmets), and other personnel (safety helmets of other colors or no safety helmets). Finally, a small sample dataset of the identities of the screened and annotated workers is obtained. 。
[0049] It should be noted that in most cases, due to many limitations in collecting images of workers at the construction site, the number of worker images collected is generally less than 1000, which is not enough to cover the diversity and complexity required for the training of the detection task, resulting in the difficulty of effectively training the model.
[0050] It should also be noted that common image annotation tools include LabelImg, LabelME, etc.
[0051] Furthermore, based on the small sample dataset of the identities of the screened and annotated workers ,for any one of the worker images, data augmentation is performed using techniques such as data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition, and all the images after data augmentation are combined with to form an enhanced dataset for worker identity detection 。
[0052] Furthermore, based on data augmentation for scene simulation is performed, and different scenes are simulated respectively to increase the robustness of the dataset training network.
[0053] In an optional embodiment, the scenes may include different weather conditions, lighting conditions, construction equipment layouts, and personnel distributions, etc. By simulating these scenes, the dataset can be further enriched, enabling the model to better adapt to various actual construction environments and improving the accuracy and stability of identity detection.
[0054] In the embodiment of the present application, the scenes include at least one or more of the following: rainy day environment scene and foggy day environment scene.
[0055] Exemplarily, the rainy day and foggy day environments are simulated respectively below to increase the robustness of the dataset training network. For the rainy day scene: First, for any one of the images to be enhanced , generate a random noise map with the same size as the original staff identity image , simulating the random effect produced by raindrops falling on the image. The random noise map has the following calculation formula: (1) where represents the noise value at position in the random noise map , and represents a random number in the interval [0, 1].[[]]END]]
[0056] Then, apply a blur effect : According to the length, angle, and width of the raindrop, apply a motion blur effect to make the raindrop appear stretched in the image. The generation formula for the blur effect is as follows: (2) (2) where represents the blurred noise value at the position of the pixel to be processed , is the standard deviation of the blur, controlling the length and width of the raindrop.
[0057] Subsequently, apply a motion blur kernel. By blurring in a specific direction and length, make the raindrops in the image look stretched, simulating the real raindrop motion effect. The motion blur kernel at position has the following formula: (3) where represents the motion blur kernel at position , is the angle of the raindrop, is the length of the raindrop, is the Dirac delta function.
[0058] It should be noted that the Motion Blur Kernel is a mathematical model that describes the blur effect caused by camera or object motion in an image. Its operation result is usually a two-dimensional matrix, representing the direction and intensity of the blur.
[0059] It should also be noted that data rotation means rotating the image by a certain angle to increase data diversity. Flipping means flipping the image horizontally or vertically. Cropping means intercepting a part of the region from the image. Random occlusion means randomly occluding part of the region in the image. Contrast adjustment means adjusting the contrast of the image to make the bright parts brighter and the dark parts darker. Random scaling means randomly adjusting the size of the image. Noise addition technology means adding random noise to the image.
[0060] Subsequently, image weighting is performed: the generated raindrop layer is weighted and superimposed with the original image to simulate the effect of raindrops superimposed on the image , and the formula is as follows: (4) where, represents the pixel value of the image with raindrop effect at position , represents the pixel value of the original image at position , is the weighting coefficient, used to adjust the transparency of the raindrop effect.
[0061] Through the above steps, the image of the staff member after simulation in rainy weather can be obtained. For each image in the dataset , rainy weather scene simulation is performed, and finally a dataset after rainy weather scene simulation is obtained.
[0062] For the foggy weather scene: First, for any image to be enhanced in , for each pixel point in , its brightness under the foggy weather effect can be calculated by the following formula: (5) (6) In the formula: represents the pixel value of the original image at position , represents the pixel value after fogging treatment at position , is the transmittance of light through the fog, related to the distance d from the pixel point to the fog center, d is the distance from the pixel point to the fog center, ([[]] , ) is the coordinate of the fog center, represents the adjusted background brightness, represents the transmittance.
[0063] It should be noted that by processing each pixel point in , an image after foggy weather scene simulation is obtained . For all the images in the dataset , foggy weather scene simulation is performed, and finally a dataset after foggy weather scene simulation is obtained .
[0064] Furthermore, and are merged to obtain small sample data in the sample enhanced dataset after image transformation data augmentation, rainy day simulation, and foggy day simulation .
[0065] Furthermore, the dataset is randomly divided into training data, validation data, and test data in proportion, and training data , validation data , and test data are obtained respectively.
[0066] It should be noted that obtaining the first data of the target power construction site and performing the first preprocessing on the first data can significantly improve the training effect and practical application performance of the identity detection model. By screening, annotating, augmenting, and reasonably dividing the original image data, not only the quality and accuracy of the data are ensured, but also the diversity and complexity of the dataset are greatly enriched. Such a preprocessing process helps the model learn more feature information in different scenarios, thereby improving the accuracy and robustness of identity detection. Especially when facing the complex and changeable power grid construction site environment, the carefully preprocessed dataset can make the model more adaptable to various challenges and achieve more reliable and efficient identity detection.
[0067] S102. Establish a first identity detection algorithm, and use the second data as the training set and validation set of the first identity detection algorithm; In an optional embodiment, the identity detection algorithm can adopt deep learning technologies, such as convolutional neural network (CNN) or recurrent neural network (RNN), etc., to construct an efficient and accurate identity detection model. These algorithms can automatically extract features from the input image and learn the unique feature representations of different identity personnel through training. During the training process, the enhanced dataset (i.e., the second data) is used as the training set and validation set, and the model parameters are optimized through repeated iterations, so that the model can gradually converge to a high-performance state.
[0068] In an optional embodiment, the identity detection algorithm can also adopt lightweight neural network models, such as MobileNet or EfficientNet, etc., to meet the requirements of real-time identity detection in resource-constrained environments. These lightweight models greatly reduce the computational amount and memory occupancy while ensuring the detection accuracy, enabling the algorithm to operate efficiently on embedded devices or mobile devices.
[0069] It should be noted that the identity detection algorithms established by the above methods cannot achieve optimal performance solely relying on limited labeled data, especially in the complex and changeable power grid construction site environment. Therefore, a brand-new first identity detection algorithm is designed in this application.
[0070] In the embodiment of this application, the first identity detection algorithm includes an identity detection network, and the identity detection network includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module.
[0071] In the embodiment of this application, the first information perception module includes: Compress the number of channels of the second data; Perform splitting operation, splicing operation, and several fusion convolution and self-attention mechanism operations on the second data after channel number compression.
[0072] In the embodiment of this application, as Figure 3 shown is the overall structure of the super-resolution information perception module SRIC, that is, the overall structure of the first information perception module, where: Use a Conv layer with a 1×1 convolution kernel to compress the number of channels of the input staff identity feature map, that is, the second data, to adjust the input spatial dimension and obtain the first information perception feature map ; Subsequently, Pass through a target-aware mixer CAMixer, which distributes the simple calculation information in to convolution through fusion convolution and self-attention mechanism, and distributes the complex calculation information to the self-attention mechanism. Input the feature map output by CAMixer into a BN layer and a ReLU layer in sequence to obtain the second information perception feature map .
[0073] Furthermore, is split in the channel dimension to respectively obtain the third information perception feature map and the fourth information perception feature map , and both have an input channel number that is half of that of
[0074] Furthermore, It is input into a target-aware mixer CAMixer, and then successively input into a BN layer and a ReLU layer to obtain the fifth information-aware feature map .
[0075] Furthermore, is input into an ELA (Efficient Local Attention) attention mechanism module. In the ELA module, first, is respectively subjected to adaptive pooling in the horizontal and vertical directions to obtain and . Then, and are respectively input into one-dimensional convolutional enhanced feature maps to enhance the position information in the horizontal and vertical directions; subsequently, they are respectively input into a group normalization and a sigmoid activation function for operation to obtain the feature map with enhanced position information in the vertical direction and the feature map with enhanced position information in the horizontal direction
[0076] (7) (8) (9) (10) In Equation (1), is the feature map obtained after horizontal direction adaptive pooling. In Equation (2), is the feature map obtained after vertical direction adaptive pooling. represents the height, represents the width. In Equations (3) and (4), and represent one-dimensional convolution, represents group normalization, represents the sigmoid function
[0077] Furthermore, , and are multiplied to obtain the sixth information-aware feature map .
[0078] (11) Furthermore, is subjected to 3×3 convolution for feature extraction to obtain the seventh information-aware feature map .
[0079] Further, the features , and are concatenated in the channel dimension to obtain the eighth information perception feature map . The number of channels of is the sum of the number of channels of , and .
[0080] Further, the feature is subjected to 3×3 convolution, and then sequentially input into the BN layer and the ReLU layer to obtain the ninth information perception feature map .
[0081] It should be noted that the first information perception module can enhance the perception ability of the characteristics of the operators under the super-resolution pixels in the construction scene. By adaptively and reasonably allocating the feature map data, this module can effectively fit the fine position information in the super-resolution feature map while ensuring the detection efficiency.
[0082] In the embodiment of the present application, the second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the position regression of the target person, and the second branch is used to calculate the identity classification of the target person.
[0083] In the embodiment of the present application, the first branch and the second branch include: Both the first branch and the second branch are used to perform several convolution operations on the feature map that is input into the second dual-branch identity detection head module and has been compressed in channels; The several convolution operations of the first branch include target regression convolution operations; The several convolution operations of the second branch include target classification convolution operations.
[0084] In the embodiment of the present application, as Figure 4 shown is the structure diagram of the DDB-Head module, that is, the second dual-branch identity detection head module in the present application, where: The input staff identity feature map is compressed in the number of channels using a Conv layer with a 1×1 convolution kernel to adjust the input spatial dimension, obtaining the first dual-branch detection feature map . is respectively input into the first and second branches for the regression and classification of the staff identity.
[0085] The first branch: is input into a 3×3 deformable convolution Deformable Conv layer, and then the output result of the DeformableConv layer is input into the ReLU layer to obtain the second dual-branch detection feature map ; Input into a 3×3 Deformable Conv layer, and then input into a BN layer to obtain the third dual-branch detection feature map , finally, Input is convolved for object regression (Locate), and a fourth dual-branch detection feature map with an output channel number of 2 (representing the upper-left and lower-right coordinate positions of the network prediction box) is obtained .
[0086] Second branch: Input into a 3×3 deformable convolution Deformable Conv layer, and then the output result of the Deformable Conv layer is input into a ReLU layer to obtain the fifth dual-branch detection feature map , Input into a 3×3 deformable convolution Deformable Conv layer, and then input into a BN layer to obtain the sixth dual-branch detection feature map , finally, Input is convolved for object classification (Sort), and a seventh dual-branch detection feature map is output .
[0087] Furthermore, and are concatenated (Concat) along the channel dimension to obtain an eighth dual-branch detection feature map .
[0088] It should be noted that the second dual-branch identity detection head module can enhance the identity detection effect of multi-angle workers in the construction site feature image. This detection head can adapt to the dynamic changes of workers and effectively exclude the interference of construction site background noise, thereby improving the detection accuracy of the model.
[0089] In the embodiment of the present application, the third spatial pyramid pooling module is used for downsampling in the identity detection network; The third spatial pyramid pooling module includes an operation for compressing the number of channels, a concatenation operation, a normalization operation, and several maximum pooling operations.
[0090] In an optional embodiment, as Figure 5 shown is the overall structure diagram of DDSPPF (Deformable DAT Spatial PyramidPooling Factorization), that is, the structure diagram of the third spatial pyramid pooling module in the present application, where: The input worker identity feature map is compressed in terms of the number of channels using a Conv layer with a 1×1 convolution kernel to adjust the input spatial dimension, and the first spatial pyramid pooling feature map is obtained , input into the Deformable Conv layer, and then into the BN layer to obtain the second spatial pyramid pooling feature map .
[0091] Furthermore, input into the deformable attention DAT (Global Attention Mechanism) module. The calculation process of this attention is as follows: (12) Among them, represents the output projection matrix of the deformable attention DAT, represents the attention output result of the th head of the deformable attention. Its calculation process is as follows:
[0092] Among them, , and respectively represent the query vector, key-value vector, and value vector under the th head. represents the dimension of each head. M represents the total number of heads, and the calculation method is the number of input channels divided by the number of heads. represents function. Assuming the input is { }, its calculation formula is:
[0093] Furthermore, after the above calculation process, obtain 's output third spatial pyramid pooling feature map .
[0094] Furthermore, input into the max pooling layer to obtain the fourth spatial pyramid pooling feature map , and then input it into two max pooling layers in sequence to obtain the fifth spatial pyramid pooling feature map and the sixth spatial pyramid pooling feature map .
[0095] Furthermore, splice , , and along the channel dimension to obtain the seventh spatial pyramid pooling feature map .
[0096] Furthermore, input It is input into the Deformable Conv layer, and then into the batch normalization layer to obtain the output eighth spatial pyramid pooling feature map of DDSPPF. .
[0097] It should be noted that when the third spatial pyramid pooling module performs downsampling in the network, it can enhance the trainable offset and the fitting ability for the staff identity target. Secondly, through this structure, the generalization of the network can be improved, and the global interaction features can be amplified, thus avoiding the loss of key information and the distortion of the feature space.
[0098] In the embodiment of the present application, based on the above-mentioned first information perception module, second dual-branch identity detection head module, and third spatial pyramid pooling module, a staff identity detection network CSII-Net is designed to implement the staff identity detection based on the power grid construction site, as Figure 2 shown.
[0099] In the present application, the initial image is first input into the convolutional layer (Conv) to perform the transformation of the number of channels to adjust the input spatial dimension, and the first identity detection feature map is obtained. , and then is successively input into the convolutional layer and the batch normalization layer (BN) to obtain the second identity detection feature map. .
[0100] Next, in this part of the present application, a two-layer feature extraction structure is designed, and these two layers are respectively a combination of the SRIC module, the convolutional layer, and the SRIC module, and a combination of the SRIC module, the convolutional layer, the SRIC module, and the DDSPPF. After is input into the combined layer of the SRIC module, the convolutional layer, and the SRIC module, after obtaining these two layer structures, the third identity detection feature map ; then is input into the combination of the SRIC module, the convolutional layer, the SRIC module, and the DDSPPF to obtain the fourth identity detection feature map . is respectively input into two parts for processing. The first part is input into the SRIC module to obtain the fifth identity detection feature map ; in the second part, is upsampled to obtain the output sixth identity detection feature map , then, and are concatenated along the channel dimension to obtain the concatenated result seventh identity detection feature map , then is input into the SRIC module for feature extraction to obtain the eighth identity detection feature map . Then, Input the SRIC module and the SRIC module in sequence to obtain the ninth identity detection feature map , then and are concatenated to obtain the tenth identity detection feature map . Finally, the processing results of the upper and lower branches are output and , corresponding to the two-branch feature processing results of small-sized and large-sized targets in the image respectively
[0101] Then a DDB-Head detection head is designed, and the DDB-Head detection head is connected to and respectively and recognized. After the recognition is completed in the upper and lower DDB-Head detection head parts and , the recognition results of the two-branch detection head are output and . Then and are merged, that is, the recognition results of the two-branch detection heads of small and large targets are merged, and finally the staff identity target detection result of the CSII-Net network is obtained , and the output result contains a tensor of prediction information, and each row corresponds to a prediction, including bounding box coordinates, class labels, and confidence scores
[0102] It should be noted that the training data and validation data are input into the CSII-Net network for training until the network loss function converges
[0103] Initialize the various parameters and hyperparameters of CSII-Net, including the number of training times, batch size, optimizer, learning rate type, and initial learning rate, etc
[0104] After completing the above preparations, the training data and validation data are input into CSII-Net for training until the network loss function converges. In each round of training, the network grabs the staff identity data in the training set according to the preset batch processing value for training, updates the loss value and various parameters in the network. After each round of training is completed, the staff identity data in the validation set is also grabbed according to the preset batch processing value for verification to check the effectiveness of each round of training. When the network training reaches the point where the loss value tends to converge, the training of the CSII-Net network ends
[0105] Furthermore, after the CSII-Net network training is completed, test data is used Perform a test by inputting this data into the optimal network parameters CSII-Net after training to conduct an effect test. The determination index for the test result of the CSII-Net network is the accuracy rate, that is, the ratio of the number of staff identities correctly recognized from the test data to the total amount of test data. When the accuracy rate reaches the user's expectation (i.e., the accuracy threshold set by the user), the network test is completed.
[0106] It should be noted that establishing the first identity detection algorithm and using the second data as the training set and validation set of the first identity detection algorithm can make full use of the existing labeled data, improve the generalization ability and accuracy of the identity detection model. By learning these data, the first identity detection algorithm can master the identity characteristics of the staff at the construction site, so as to achieve fast and accurate identity recognition in subsequent practical applications. In addition, using the second data as the training set and validation set can also help adjust and optimize the algorithm parameters and further improve the detection performance.
[0107] S103, perform identity detection on the staff at the target power grid construction site according to the first identity detection algorithm after verification.
[0108] In actual application, deploy and apply it. Input the staff identity image into CSII-Net and output the detection result of the staff identity. This result contains a tensor of detection information, and each row corresponds to a prediction, including the coordinates of the staff bounding box, the staff category label, and the confidence score. The staff identity image can be obtained by using various shooting devices or intercepted as needed from the video at the power grid construction site.
[0109] In summary, the present invention proposes a method for detecting the identity of staff at a power grid construction site, which obtains the first data of the target power construction site, performs the first preprocessing on the first data to obtain the second data; establishes the first identity detection algorithm, and uses the second data as the training set and validation set of the first identity detection algorithm; performs identity detection on the staff at the target power grid construction site according to the first identity detection algorithm after verification. The present invention establishes an identity detection network for staff at the construction site with strong generalization, high accuracy and automation, and can efficiently identify the identity of staff at the construction site. The network designed by the present invention can be deployed at the construction site or substation along each transmission line after training to realize real-time observation of the identity of staff at the construction site.
[0110] Embodiment 2, in a preferred embodiment, adopt as Figures 2 - 5The actual scenario application of the shown structure is as follows: Obtain images containing workers at the electric power construction site through a shooting device (such as a drone, a fixed device, etc.), and screen the identity data. To ensure a sufficient amount of data, screen the data of workers with various identities, and crop 300 images containing the identities of various workers according to the size of 640×640 pixels. , Subsequently, use the LabelImg image annotation tool to annotate the category and coordinate box information of the identity targets of all worker identity images: Preset the VOC annotation format, frame and mark the worker identity area, create a bounding box and a category label. For the category label, in this application, the worker identity categories are divided into {management personnel (red), technical personnel (blue), construction and inspection personnel (yellow), supervision personnel (white), and other personnel (other colors or not wearing)}, corresponding to the category indexes respectively. . After the annotation is completed, generate annotation image data corresponding one by one to the annotation files , and the data format is.xml. The annotated data set is and subjected to data augmentation operations, including data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition. Each data augmentation operation for and will generate an equal amount of augmented data. Finally, obtain the total data set including the original data and all augmented data
[0111] . Based on and , perform data augmentation for scene simulation. The following respectively simulate rainy and foggy environments.
[0112] Perform data augmentation operations for rainy scene simulation on and to obtain the total data set after rainy scene simulation , .
[0113] Perform data augmentation operations for foggy scene simulation on and to obtain the total data set after foggy scene simulation , .
[0114] Combine , , , and , Perform data merging to obtain the total dataset for sample enhancement , .
[0115] After the preprocessing of the operator dataset is completed, perform dataset partitioning. In this application, the training set, validation set, and test set are randomly partitioned in a ratio of 7:2:1. After partitioning, 5040 training data, 1440 validation data, and 720 test data can be obtained.
[0116] Train the dataset after data partitioning according to the detection network designed in this application: First, initialize the various parameters and hyperparameters for the training of the construction site worker identity detection network CSII-Net. In this embodiment, several important parameters in network training are described: Initialize the corresponding training batch size according to the hardware environment, initialize the optimizer as the optimizer optimizer, set Batchsize to 4, initialize the number of training epochs to 200, the learning rate type is cosine learning rate, and the initial learning rate is 0.001. These parameters need to be adjusted according to the effects of multiple network trainings until the network converges.
[0117] After completing the initialization of the basic parameters for network training, start model training. In this network, this application first inputs the initial image into the convolutional layer (Conv) to perform channel number transformation to adjust the input spatial dimension and obtain the first identity detection feature map , and then is sequentially input into the convolutional layer and the batch normalization layer (BN) to obtain the second identity detection feature map . Next, it is the backbone part of the model. In this application, a two-layer feature extraction structure is designed in this part. The two-layer structure is respectively a combination of SRIC modules, convolutional layers, and SRIC modules, and a combination of SRIC modules, convolutional layers, SRIC modules, and DDSPPF modules. After sequentially inputting these two-layer structures, the results of feature extraction, the third identity detection feature map and are respectively output. Subsequently, are respectively input into two parts for processing. The first part is input into the SRIC module to obtain the fifth identity detection feature map ; in the second part, is upsampled to obtain the output sixth identity detection feature map . Subsequently, is concatenated with along the channel dimension to obtain the concatenation result, the seventh identity detection feature map . Then, is input into the SRIC module for feature extraction to obtain the eighth identity detection feature map . Subsequently, are input into the SRIC module and the SRIC module in sequence to obtain the ninth identity detection feature map . Then, is concatenated with to obtain the tenth identity detection feature map . Finally, the processing results of the upper and lower branches are output and , which respectively correspond to the two-branch feature processing results of small-sized and large-sized targets in the image. Then, the detection head part of the network is designed in this application. The designed DDB-Head detection head is respectively connected to and and recognized. After the recognition in the upper and lower DDB-Head detection head parts is completed and , the detection head recognition results of the two branches are output and . Subsequently, this application combines and , that is, combines the detection head recognition results of small-sized and large-sized targets, and finally obtains the staff identity target detection result of the CSII-Net network . This result contains a tensor of detection information, and each row corresponds to a prediction, including the staff bounding box coordinates, staff class labels, and confidence scores.
[0118] The training of CSII-Net updates the internal parameters of the network through backpropagation of the loss function. In the loss function of the present invention, the bounding box coordinate loss of the staff identity target adopts the DFL loss, the confidence loss of the bounding box adopts the WIoU loss, and the classification loss adopts the cross-entropy loss. The above loss function is updated round by round during the gradient backpropagation process of network training. The network gradually converges during the round-by-round training, and the detection accuracy gradually improves. After the end of each round of training, the network will call the divided validation data to verify the accuracy , recall rate and mean average precision mAP of the network, which are used as the evaluation criteria for whether the training effect is effective. After the training of CSII-Net of the present invention is completed, 720 divided staff identity image data are used as test data, and this data is input into the optimal network parameter CSII-Net after training for effect testing.
[0119] In practical applications, it is deployed for pilot applications. The identity image data of the staff is input into CSII-Net, and the extraction result of the identity image data of the staff is output. The output result contains a tensor of prediction information, and each row corresponds to a prediction, including bounding box coordinates, class labels, and confidence scores.
[0120] Embodiment 3. In this embodiment, a staff identity detection system for a power grid construction site is further provided, including: A data acquisition and processing module, configured to acquire first data of a target power construction site and perform first preprocessing on the first data to obtain second data; The first preprocessing includes data screening, data annotation, data augmentation, and data partitioning; An algorithm establishment module, configured to establish a first identity detection algorithm, and use the second data as the training set and verification set of the first identity detection algorithm; The first identity detection algorithm includes an identity detection network, and the identity detection network includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module; The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the position regression of the target person, and the second branch is used to calculate the identity classification of the target person; A detection module, configured to perform identity detection on the staff of the target power grid construction site according to the first identity detection algorithm after verification is completed.
[0121] The above unit modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0122] This embodiment also provides a computer device. The computer device can be a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for detecting the identity of workers at the power grid construction site. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, trackball, or touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0123] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the following steps are implemented: Obtain the first data of the target power grid construction site and perform a first preprocessing on the first data to obtain the second data; The first preprocessing includes data screening, data annotation, data enhancement, and data partitioning; Establish a first identity detection algorithm and use the second data as the training set and verification set of the first identity detection algorithm; The first identity detection algorithm includes an identity detection network, and the identity detection network includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module; The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the position regression of the target person, and the second branch is used to calculate the identity classification of the target person; Perform the identity detection of the workers at the target power grid construction site according to the first identity detection algorithm after verification.
[0124] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
[0125] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages.
[0126] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0129] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0130] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these changes and modifications.
Claims
1. A method for detecting the identity of on-site workers at a power grid construction site, characterized in that, Including: Obtain the first data of the target power construction site, and perform a first preprocessing on the first data to obtain second data; The first preprocessing includes data screening, data annotation, data augmentation, and data partitioning; Establish a first identity detection algorithm, and use the second data as the training set and validation set of the first identity detection algorithm; The first identity detection algorithm includes an identity detection network, and the identity detection network includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module; The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the position regression of the target person, and the second branch is used to calculate the identity classification of the target person; Compress the number of channels of the staff identity feature map input to the second dual-branch identity detection head module to obtain a first dual-branch detection feature map; The first branch is used to input the first dual-branch detection feature map into a deformable convolutional layer, and then input it into an activation function layer to obtain a second dual-branch detection feature map; Input the second dual-branch detection feature map into a deformable convolutional layer, and then input it into a normalization layer to obtain a third dual-branch detection feature map. Input the third dual-branch detection feature map into a convolution for target regression to obtain a fourth dual-branch detection feature map; The second branch is used to input the first dual-branch detection feature map into a deformable convolutional layer, and then input it into an activation function layer to obtain a fifth dual-branch detection feature map. Input the fifth dual-branch detection feature map into a deformable convolutional layer, and then input it into a normalization layer to obtain a sixth dual-branch detection feature map. Input the sixth dual-branch detection feature map into a convolution for target classification and output a seventh dual-branch detection feature map; Perform identity detection on the staff at the target power grid construction site according to the first identity detection algorithm after verification is completed.
2. The method for detecting the identity of on-site staff at the power grid construction site according to claim 1, characterized in that, The first branch and the second branch include: Both the first branch and the second branch are used to perform a number of convolution operations on the feature map that is input into the second dual-branch identity detection head module and has its channels compressed; The number of convolution operations of the first branch includes target regression convolution operations; The number of convolution operations of the second branch includes target classification convolution operations.
3. The method for detecting the identity of on-site workers at the power grid construction site according to claim 2, wherein The third spatial pyramid pooling module is used to perform downsampling in the identity detection network; The third spatial pyramid pooling module includes an operation for compressing the number of channels, a splicing operation, a normalization operation, and a number of max pooling operations.
4. The method for detecting the identity of on-site staff at the power grid construction site according to claim 3, wherein, The first information perception module includes: Compress the number of channels of the second data; Perform segmentation operations, splicing operations, and a number of fusion convolution and self-attention mechanism operations on the second data after the number of channels is compressed.
5. The method for detecting the identity of on-site workers at the power grid construction site according to claim 4, characterized in that, The performing the first preprocessing on the first data includes: The first data is an image containing operating personnel at the target power construction site; Screen and label the first data; Perform data augmentation on the first data after screening and labeling to obtain second data.
6. The method for detecting the identity of on-site staff at the power grid construction site according to claim 5, characterized in that, The performing data augmentation on the first data after screening and labeling includes: According to the safety helmet categories of construction workers at the job site in the first data, use an image annotation tool to annotate the category and coordinate box information of the personnel targets in all the staff images; Among them, the categories are divided into management personnel, technical personnel, construction and inspection personnel, supervision personnel, and other personnel; Obtain the small sample data set after screening and annotation. For any staff image in the small sample data set, use data rotation, flipping, cropping, random occlusion, contrast adjustment, random scaling, and noise addition techniques for data augmentation, and combine all the images after data augmentation with the small sample data set to form an enhanced staff identity detection data set; Perform data augmentation on the enhanced staff identity detection data set for scene simulation to obtain the second data.
7. The method for detecting the identity of on-site workers at the power grid construction site according to claim 6, wherein, The scenarios include at least one or more of the following: rainy day environment scenario and foggy day environment scenario.
8. An identity detection system for on-site workers at a power grid construction site, which applies the method according to any one of claims 1 to 7, characterized in that, Include: A data acquisition and processing module, configured to acquire the first data of the target power construction site and perform a first preprocessing on the first data to obtain the second data; The first preprocessing includes data screening, data annotation, data augmentation, and data division; An algorithm establishment module, configured to establish a first identity detection algorithm, and use the second data as the training set and validation set of the first identity detection algorithm; The first identity detection algorithm includes an identity detection network, and the identity detection network includes a first information perception module, a second dual-branch identity detection head module, and a third spatial pyramid pooling module; The second dual-branch identity detection head module includes a first branch and a second branch. The first branch is used to calculate the position regression of the target personnel, and the second branch is used to calculate the identity classification of the target personnel; A detection module, configured to perform identity detection on the staff at the target power grid construction site according to the first identity detection algorithm after verification is completed.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Personnel identification method in electric power operation scene
CN114821486A
Substation infrared image mouse identification method based on deep learning
CN118314532A
Industrial pure iron metallographic specimen defect detection optimization method based on YOLOv8obb
CN119831964A
Person re-identification method combining reverse attention and multi-scale deep supervision
US20210232813A1
Cited By
Power distribution cabinet instrument intelligent identification method and system
CN121392309A