Power marketing on-site operation clothes safety standardization perception method and system
By adopting multi-angle detection and deep neural network analysis methods at the power marketing site, the safety accident problem caused by workers' neglect of safety helmets is solved, and automated monitoring and efficient violation analysis are achieved.
Patent Information
- Application Number
- CN202411938132.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-27
AI Technical Summary
In on-site operations in the power industry, workers ignore the importance of safety helmets, resulting in frequent safety accidents. The existing manual monitoring methods consume a lot of manpower and are prone to missed inspections.
A standardized perception method for clothing safety of on-site power marketing operations is adopted. Violation analysis and key point extraction are carried out by detecting the full-body dress imaging of the operator from four angles, from forward, reverse and two lateral directions, and inputting them into a pre-trained deep neural network, including the human segmentation network and the clothing safety standard analysis network.
Automatic monitoring is realized to minimize false alarms and personnel intervention, improve supervision efficiency, reduce labor costs, and improve the accuracy of clothing violation analysis.
Smart Images

Figure CN120047964A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image analysis, and in particular to a method and system for sensing safety standardization of clothing for power marketing field operations. Background Art
[0002] Safety clothing is an important labor protection tool in the power industry. It is widely used and very important. However, in actual scenarios, many workers still ignore the importance of safety helmets. At the same time, due to inadequate supervision, countless safety accidents have been caused by not wearing safety clothing properly. Therefore, it is very important to detect the wearing status of safety clothing for operators.
[0003] Manual monitoring of safety clothing wearing not only consumes a lot of manpower but also often creates the risk of missed detection. With the development and progress of computer vision technology in recent years, the target detection algorithm based on intelligent deep learning has become one of the application scenarios for safety clothing wearing detection. Through automatic monitoring and intelligent learning of power industry knowledge, the number of false alarms and human intervention can be minimized, which is conducive to enterprises implementing standardized production management, ensuring production safety, improving supervision efficiency, and reducing labor costs. It can play an important role in enterprise safety production supervision scenarios. Summary of the invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a method and system for perceiving safety standardization of clothing for power marketing field operations.
[0005] In a first aspect, the present invention provides a method for sensing safety standardization of clothing for power marketing field operations, comprising:
[0006] Detect the whole body clothing imaging of the operator from four angles: forward, reverse and two lateral directions;
[0007] The whole-body clothing image is input into a pre-trained deep neural network for sensing the safety standardization of work clothing, wherein the deep neural network includes: a human body segmentation network for segmenting the human body, and a clothing safety standard analysis network for sensing the safety standardization of work clothing foreground; the human body foreground is extracted from the whole-body clothing image by the human body mask predicted by the human body segmentation network; the clothing safety standard analysis network predicts the bounding box, category label and violation category label based on the image data in the human body foreground, the clothing safety standard analysis network predicts the clothing mask based on the image data in the human body foreground, and the clothing safety standard analysis network predicts the key points based on the image data in the human body foreground;
[0008] The violation category is analyzed based on the relationship between the clothing details and the key points between the clothing details and the human body parts, and the final violation situation is obtained by integrating the predicted violation category label and the classified category obtained based on the key point analysis.
[0009] Furthermore, the human segmentation network consists of an encoder network and a corresponding decoder network;
[0010] The encoder network consists of 13 convolutional layers and 5 pooling layers. In the encoder network, the pooling layers are connected after the 2nd convolutional layer, the 4th convolutional layer, the 7th convolutional layer, the 10th convolutional layer, and the 13th convolutional layer respectively.
[0011] The decoder network is implemented by 13 convolutional layers and 5 upsampling layers. Each convolutional layer of the decoder network corresponds to a convolutional layer in the encoder network. In the decoder network, upsampling is set before the 1st convolutional layer, the 4th convolutional layer, the 7th convolutional layer, the 10th convolutional layer, and the 12th convolutional layer respectively. The last convolution operation of the decoder network produces a single-channel feature map. The 3rd, 6th, 9th, and 11th convolutional layers of the decoder network are connected to 1×1 convolutional layers respectively. A deconvolution layer is added after the 1×1 convolutional layer to process the feature map to adapt to the original size of the input. The outputs of all deconvolutional layers and the single-channel feature map of the last convolutional layer of the encoder network are merged, and the merged results are fused by a 1×1 convolutional layer to combine features of different scales. The features are processed by the activation function to show the probability of each pixel belonging to the foreground or background of the human body, and the probability is binarized to obtain the human body mask.
[0012] Furthermore, the clothing safety standard analysis network includes: a convolution layer for extracting human foreground features, the convolution layer is connected to a feature pyramid network, the feature pyramid network includes: a multi-level forward feature encoder and a reverse feature encoder based on Resnet or VGG, the forward feature encoder includes five forward feature encoding layers based on Resnet or VGG, the first four forward feature encoding layers are followed by a pooling layer for downsampling, the channel of the output feature of the forward feature encoding layer at any level is twice the channel of the output feature of the feature encoding layer at the previous level, and the output of the forward feature encoding layer at any level is The width and height of the feature are 1 / 2 of the width and height of the output feature of the forward feature encoding layer of the previous level; the reverse feature encoder includes five reverse feature encoding layers based on Resnet or VGG, the first four reverse feature encoding layers are followed by upsampling, the channel of the output feature of the reverse feature encoding layer at any level is 1 / 2 times the channel of the output feature of the reverse feature encoding layer of the previous level, the width and height of the output feature of the reverse feature encoding layer at any level are twice the width and height of the output feature of the reverse feature encoding layer of the previous level, and a 1×1 convolution layer is set between the corresponding forward feature encoder and the reverse feature encoder to fuse the features;
[0013] The features output by the feature pyramid network are spatially transformed through three Roi Align operations, and the three features after spatial transformation are respectively input into a first head module for regression prediction of bounding boxes, category labels and violation category labels, a second head module for prediction of clothing masks, and a third head module for prediction of key points. The first head module, the second head module and the third head module respectively output predicted bounding boxes, categories and violation categories, clothing masks and key points.
[0014] Furthermore, the clothing safety standard analysis network is trained using training data, and the loss function involved in the training is a weighted sum of the cross entropy loss of clothing categories and clothing violation categories, the detection box regression loss, the clothing mask cross loss, and the key point cross entropy loss. The parameters of the deep neural network are updated by stochastic gradient descent with the goal of minimizing the loss function.
[0015] Furthermore, the construction of the training data includes:
[0016] Collect full-body images of power marketing field workers, and adjust the proportion of people in the full-body images, occlusions, image size, and viewing angle of the full-body images to enrich the diversity of full-body images.
[0017] Manually annotate full-body clothing images, including key point labels, human body masks, and labels of staff clothing;
[0018] The labels of the staff’s clothing include: bounding box, category label, clothing mask, and violation category label.
[0019] Furthermore, based on the standard specifications for safety wear in marketing field operations, a set of standards for wear violation levels and violation items are formulated: 1. Major violations include: not wearing a safety helmet, not wearing a work jacket, not wearing work pants, not wearing insulated shoes, and not wearing safety gloves; 2. General violations include: the safety helmet is not fastened with a chin strap, and the work jacket and work pants are damaged; 3. Minor violations include: the chin strap of the safety helmet is not on both sides of the ears, the chin strap buckle of the safety helmet is not standardized, and the buttons of the work clothes are not fully fastened;
[0020] Construct labels for staff clothing based on the level of wearing violation and violation criteria: To determine whether the staff clothing is a major violation, the manual annotator needs to draw a bounding box for each piece of clothing of the staff and assign a category label. The bounding box indicates the clothing detected in the full-body clothing image, and the bounding box is determined by the coordinates of the two end points of the rectangular diagonal; the category label indicates the type of clothing the staff is wearing; set clothing masks for each type of clothing in the full-body clothing image, and mark each pixel mask to mark whether each element in the full-body clothing image belongs to any type of clothing;
[0021] When determining that the clothing is work clothing, it is also necessary to determine whether the staff's clothing falls into general violations and minor violations. The manual annotator needs to configure a violation category label for each work clothing of the staff. The violation category label indicates the type of violation of the work clothing worn by the staff; and, configure key point labels for the ears and safety helmets, and configure key point labels for other clothing except safety helmets, so as to perform anomaly analysis based on the relationship between key point labels.
[0022] Furthermore, the process of setting up the human body mask and clothing mask annotation includes: obtaining the instance contour through contour detection, obtaining a preliminary instance mask based on the instance contour, and correcting the instance contour by a manual annotator to improve the instance mask.
[0023] Furthermore, based on the relationship between clothing details and key points between clothing details and human body parts, the violation categories include:
[0024] Determine whether the chin strap of the helmet is worn on both sides of the ears based on the relationship between the key points of the ear and the key points of the chin strap of the helmet in the lateral full-body clothing imaging;
[0025] Determine whether all buttons are fastened based on the relationship between the key points of the sub-buttons and the key points of the mother-buttons in the forward full-body clothing imaging;
[0026] Determine whether the chin strap buckle is standardized based on the key point relationship of the chin strap buckle and the buckle of the helmet.
[0027] In a second aspect, the present invention provides a device for perceiving safety standardization of clothing for power marketing field operations, comprising: at least one processing unit, wherein the processing unit is connected to a storage unit and a collection unit via a bus unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the method for perceiving safety standardization of clothing for power marketing field operations is implemented.
[0028] In a third aspect, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method for perceiving safety standardization of clothing for power marketing field operations.
[0029] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art:
[0030] The present application detects the full-body clothing imaging of the working human body from four angles: forward, reverse and two lateral directions; the full-body clothing imaging is input into the pre-trained deep neural network for the perception of safety standardization of working clothing for violation analysis and key point extraction, and the deep neural network includes: a human segmentation network and a clothing safety standard analysis network; the human foreground is extracted from the full-body clothing imaging by the human mask predicted by the human segmentation network; the human mask is obtained from the full-body clothing image, and the human body part of the foreground is extracted using the human mask, which is helpful for the subsequent clothing analysis. The clothing safety standard analysis network predicts the bounding box, category and violation category, clothing mask and key points based on the image data in the human foreground; the violation category is analyzed based on the relationship between the clothing details and the key points between the clothing details and the human body parts, and the violation category predicted by the deep neural network and the violation category obtained based on the key point analysis are integrated to obtain the final violation situation. The present application combines the violation category obtained by the key point analysis and the violation category obtained directly based on the features in the human foreground, and performs violation analysis based on different features to improve the accuracy of the clothing violation analysis at the marketing operation site. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0033] Figure 1 A flowchart of a method for sensing safety standardization of clothing for power marketing field operations provided by an embodiment of the present invention;
[0034] Figure 2 A schematic diagram of a method for sensing safety standardization of clothing for on-site operations in power marketing provided by an embodiment of the present invention;
[0035] Figure 3 A structural diagram of a human body segmentation network provided by an embodiment of the present invention;
[0036] Figure 4 A structural diagram of a clothing safety standard analysis network provided by an embodiment of the present invention;
[0037] Figure 5 An example of key points of a top provided by an embodiment of the present invention;
[0038] Figure 6 A schematic diagram of a device for sensing clothing safety standardization for power marketing field operations provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0041] Example 1
[0042] By using computer vision and machine learning technology, the technology of the present invention realizes a method for perceiving safety standardization of work clothing at power marketing sites. The present application constructs a deep neural network for perceiving safety standardization of work clothing, and creates training data specifically for training the deep neural network. After the deep neural network is trained using the training data, the trained deep neural network extracts a human foreground from full-body clothing imaging through a human mask predicted by a human segmentation network; a clothing safety standard analysis network predicts bounding boxes, categories and violation categories, clothing masks, and key points based on image data within the human foreground; violation categories are analyzed based on the relationship between clothing details and between clothing details and key points of human body parts, and the violation categories predicted by the deep neural network and the violation categories obtained based on key point analysis are integrated to obtain the final violation situation.
[0043] In the training phase, we first construct the training data. The process includes:
[0044] Full-body images of on-site workers in power marketing are collected from four angles: forward, reverse, and two lateral angles. Adjustments are made to the proportion of people in the full-body images, occlusions, image size, and viewing angle to enrich the diversity of the full-body images.
[0045] Manually annotate full-body clothing images. The process includes:
[0046] The present application sets a human body mask for the human body in the full-body clothing image, and the human body mask marks whether each element in the full-body clothing image belongs to the human body; the setting process of the human body mask includes: obtaining the human body contour through contour detection, and obtaining a preliminary human body mask based on the human body contour. In this process, the human body contour is not accurate, and the human body contour is corrected by a manual annotator to improve the human body mask.
[0047] The annotation of the full-body clothing image needs to meet the requirements of judging the safety wear standards for marketing field operations. In the present invention, based on the safety wear standards for marketing field operations, a set of standards for wear violation levels and violation items are formulated: 1. Major violations, including: not wearing a safety helmet, not wearing a work jacket, not wearing work pants, not wearing insulating shoes, and not wearing safety gloves; 2. General violations, including: the safety helmet is not fastened with a chin strap, the work jacket and work pants are damaged; 3. Minor violations: including the safety helmet chin strap is not on both sides of the ears, the safety helmet chin strap buckle is not standardized, and the work clothes buttons are not fully fastened.
[0048] Construct labels for staff clothing based on the level of wearing violations and standards for violations: In order to determine whether the staff clothing is a major violation, it is actually necessary to identify whether the clothing worn by the staff is work clothing or non-work clothing. Work clothing includes: safety helmets, work clothes tops, work clothes pants, insulating shoes and safety gloves. Therefore, in this application, the manual annotator needs to draw a bounding box for each piece of clothing of the staff and assign a category label. The bounding box indicates the clothing detected in the full-body clothing image, and the bounding box is determined by the coordinates of the two end points of the rectangular diagonal; the category label indicates the type of clothing worn by the staff. This application sets clothing masks for various types of clothing in the full-body clothing image, and the mask that marks each pixel marks whether each element in the full-body clothing image belongs to any type of clothing; the mask setting process includes: obtaining clothing contours through contour detection, and obtaining a preliminary clothing mask based on the contour. In this process, the clothing contour is not accurate, and the clothing contour is corrected by the manual annotator to improve the clothing mask.
[0049] When determining that the clothing is work clothing, it is also necessary to determine whether the staff's clothing falls into the general violation and minor violation items. This requires determining the abnormal situation of the work clothing. Therefore, the manual annotator needs to configure a violation category label for each work clothing of the staff, and the violation category label indicates the violation type of the work clothing worn by the staff.
[0050] Simply using violation categories to perform classification and regression training on the neural network model will result in poor prediction results. To further improve the accuracy of violation classification, it is necessary to further analyze the relationship between clothing details and between clothing details and human body parts. For example, the relationship between clothing details and human body parts includes the relationship between the ear and the chin strap of a safety helmet. To meet the need to determine the relationship between the ear and the chin strap of a safety helmet, key point labels are assigned to the ear and the safety helmet. To meet the need to determine the relationship between clothing details, key point labels are assigned to other clothing items except the safety helmet. Violation analysis is performed based on the relationship between key point labels. For example Figure 5 As shown in the figure, the key point labels of the work clothes top are based on the existing clothing key points, and the key points of the sub-buttons and mother buttons represented by uppercase and lowercase letters are added. The key points of the ear use the inner and outer contour points of the auricle and the key points of the earlobe in the human posture estimation task.
[0051] Then build a deep neural network for the perception of safety standardization of work clothing, such as Figure 2 As shown, the deep neural network includes:
[0052] Human segmentation network for segmenting human bodies.
[0053] like Figure 3As shown in Figure 1, the human body segmentation network consists of an encoder network and a corresponding decoder network; the encoder network consists of 13 convolutional layers and 5 pooling layers. In the encoder network, pooling layers are connected after the 2nd convolutional layer, the 4th convolutional layer, the 7th convolutional layer, the 10th convolutional layer, and the 13th convolutional layer. The decoder network is implemented by 13 convolutional layers and 5 upsampling, and each convolutional layer of the decoder network corresponds to a convolutional layer in the encoder network. In the decoder network, upsampling is set before the 1st convolutional layer, before the 4th convolutional layer, before the 7th convolutional layer, before the 10th convolutional layer, and before the 12th convolutional layer. After each convolution operation in the decoder network, batch normalization is performed on the output feature map. The last convolution operation of the decoder network produces a single-channel feature map. The 3rd, 6th, 9th and 11th convolutional layers of the decoder network are connected to 1×1 convolutional layers respectively. After the 1×1 convolutional layer, a deconvolutional layer is added to process the feature map to adapt to the original size of the input; the output of all deconvolutional layers and the single-channel feature map of the last convolutional layer of the encoder network are merged, and the merged result is fused by a 1×1 convolutional layer to combine features of different scales. The features are processed by the Sigmoid activation function to show the probability of each pixel belonging to the foreground or background of the human body, and the probability is binarized to obtain the human body mask.
[0054] The human segmentation network is trained using the full-body dressed images and their human masks in the training data until the intersection-over-union loss between the human mask predicted by the human segmentation network and the real human mask in the training data is lower than the set threshold. The human segmentation network integrates multi-scale convolutional features in the fully convolutional network to achieve accurate human segmentation. The human foreground is extracted from the full-body dressed imaging through the human mask predicted by the human segmentation network. In most cases, the clothing in the full-body dressed image only occupies a part, and the full-body dressed image has a large background, which will have a great impact on the analysis of clothing. Deep neural networks will be misled by the background during analysis. Obtaining the human mask from the full-body dressed image and using the human mask to extract the foreground human part will help to achieve subsequent clothing analysis.
[0055] The deep neural network includes: a clothing safety standard analysis network for sensing the safety standardization of work clothing for human foreground, and the clothing safety standard analysis network performs feature learning and analysis on the segmented human foreground. The clothing safety standard analysis network includes:
[0056] like Figure 4As shown, a convolutional layer for extracting human foreground features, the convolutional layer is connected to a feature pyramid network, the feature pyramid network comprises: a multi-level forward feature encoder based on Resnet or VGG, the forward feature encoder comprises five forward feature encoding layers based on Resnet or VGG, the first four forward feature encoding layers are followed by a pooling layer for downsampling, the channel of the output feature of the forward feature encoding layer at any level is twice the channel of the output feature of the feature encoding layer at the previous level, and the width and height of the output feature of the forward feature encoding layer at any level are 1 / 2 of the width and height of the output feature of the forward feature encoding layer at the previous level. A multi-level reverse feature encoder based on Resnet or VGG, wherein the reverse feature encoder comprises five reverse feature encoding layers based on Resnet or VGG, the first four reverse feature encoding layers are followed by upsampling, the channels of the output features of the reverse feature encoding layer at any level are 1 / 2 times the channels of the output features of the reverse feature encoding layer at the previous level, the width and height of the output features of the reverse feature encoding layer at any level are twice the width and height of the output features of the reverse feature encoding layer at the previous level, and a 1×1 convolution layer is set between the corresponding forward feature encoder and the reverse feature encoder to fuse features.
[0057] The features output by the feature pyramid network are spatially transformed through the Roi Align operation, and the spatially transformed features are respectively input into a first head module for regression prediction of bounding boxes, category labels and violation category labels, a second head module for prediction of clothing masks, and a third head module for prediction of key points.
[0058] In the specific implementation process, the first head module contains three fully connected networks, one fully connected network for clothing category classification, one fully connected network for violation category classification, and one fully connected network for bounding box regression. The second head module contains N convolutional layers and two deconvolutional layers connected in sequence, which are used to predict key points of clothing and human body parts; the third head module has the same structure as the second head module and is used to predict clothing masks.
[0059] The clothing safety standard analysis network is trained by training data. In a specific implementation process, the full-body clothing images at four angles, namely, forward, reverse and two lateral angles, associated in the training data are input into the clothing safety standard analysis network. The loss function involved in the training is a weighted sum of the cross entropy loss of clothing category and clothing violation category, the detection box regression loss, the clothing mask cross loss, and the key point cross entropy loss. The parameters of the deep neural network are updated by stochastic gradient descent with the goal of minimizing the loss function.
[0060] In traditional single-model detection, the entire image needs to be processed, resulting in reduced accuracy, low efficiency, and high training complexity. The multi-angle and regional multi-model detection technology used in this device can significantly improve detection efficiency and accuracy. Multi-angle can distinguish the wearing situation from the front and side images, making the detection more accurate; regional multi-model detection designs special models for different regions to improve the detection capability of specific targets, and samples lightweight models to reduce the model calculation burden and optimize resource allocation. The Roi Align operation extracts features of regions of interest, such as the head, upper body, hands, legs, and feet, and divides them into different regional features; then, based on the analysis of each regional feature, it focuses on the detection tasks within the region, such as head model detection of helmets, hand model detection of gloves, and foot model detection of insulating shoes.
[0061] In the application stage, Figure 1 As shown, including:
[0062] Detect the whole body clothing imaging of the operator from four angles: forward, reverse and two lateral directions;
[0063] Input the whole body clothing image into a pre-trained deep neural network for standardization of work clothing safety perception to perform violation analysis and key point extraction, the deep neural network includes: a human body segmentation network for segmenting the human body, a clothing safety standard analysis network for standardization of work clothing safety perception of the human body foreground; extract the human body foreground from the whole body clothing image through the human body mask predicted by the human body segmentation network; the clothing safety standard analysis network performs regression prediction of the bounding box, category label and violation category label based on the image data in the human body foreground, the clothing safety standard analysis network performs regression prediction of the clothing mask based on the image data in the human body foreground, and the clothing safety standard analysis network performs key point prediction based on the image data in the human body foreground;
[0064] The violation categories are analyzed based on the relationship between the clothing details and the key points between the clothing details and the human body parts. The violation categories predicted by the deep neural network and the violation categories obtained based on the key point analysis are integrated to obtain the final violation situation.
[0065] Violation categories based on the relationship between clothing details and key points between clothing details and human body parts include:
[0066] Based on the relationship between the ear key points and the chin strap key points of the helmet in the lateral full-body clothing imaging, it is determined whether the chin strap of the helmet is worn on both sides of the ears; when the chin strap of the helmet covers the ear key points and the ears cover the chin strap key points of the helmet, it is determined whether the chin strap of the helmet is worn on both sides of the ears.
[0067] Based on the relationship between the key points of the sub-button and the key points of the mother button in the forward full-body clothing imaging, it is determined whether all the buttons are fastened; when the sub-button is buttoned on the corresponding mother button, the sub-button will block the key points of the mother button, and the positions of the sub-button and the mother button are determined according to the key points of the sub-button and the mother button, and whether all the buttons are fastened is determined according to whether the key points of the mother button are blocked by the sub-button.
[0068] Based on the relationship between the key points of the parent and child buckles of the chin strap of the helmet, determine whether the chin strap buckle is standardized; based on the distribution of the key points of the parent and child buckles when the chin strap buckle is fully fastened, the distribution of the key points of the parent and child buckles when the chin strap buckle is not fully fastened, and the distribution of the key points of the parent and child buckles when the chin strap buckle is not fastened in each reference view, determine whether the actual chin strap buckle in each actual view is standardized.
[0069] In the detection process, interactivity and feedback mechanisms are introduced to increase the interactive function between operators and equipment, so that the equipment can intelligently prompt violations. Operators can check again after making corrections, and ultimately judge whether they are qualified according to the violation level and violation item standards.
[0070] Example 2
[0071] See also Figure 6 As shown, an embodiment of the present invention provides a device for perceiving safety standardization of clothing for on-site work in power marketing, comprising: at least one processing unit, the processing unit being connected to a storage unit and a collection unit via a bus unit, the storage unit being a computer-readable storage medium, and being used to store software programs, computer executable programs, and modules, such as the software programs, computer executable programs, and modules corresponding to a method for perceiving safety standardization of clothing for on-site work in power marketing in an embodiment of the present invention. The processing unit implements the above-mentioned method for perceiving safety standardization of clothing for on-site work in power marketing by running the software programs, computer executable programs, and modules stored in the storage unit, comprising:
[0072] Detect the whole body clothing imaging of the operator from four angles: forward, reverse and two lateral directions;
[0073] Input the whole body clothing image into a pre-trained deep neural network for standardization of work clothing safety perception to perform violation analysis and key point extraction, the deep neural network includes: a human body segmentation network for segmenting the human body, a clothing safety standard analysis network for standardization of work clothing safety perception of the human body foreground; extract the human body foreground from the whole body clothing image through the human body mask predicted by the human body segmentation network; the clothing safety standard analysis network performs regression prediction of the bounding box, category label and violation category label based on the image data in the human body foreground, the clothing safety standard analysis network performs regression prediction of the clothing mask based on the image data in the human body foreground, and the clothing safety standard analysis network performs key point prediction based on the image data in the human body foreground;
[0074] The violation categories are analyzed based on the relationship between the clothing details and the key points between the clothing details and the human body parts. The violation categories predicted by the deep neural network and the violation categories obtained based on the key point analysis are integrated to obtain the final violation situation.
[0075] Of course, the computer program stored in the storage unit of the device for perceiving safety standardization of clothing for power marketing on-site operations provided by an embodiment of the present invention is not limited to the operations of the method described above, but can also execute related operations of the method for perceiving safety standardization of clothing for power marketing on-site operations provided by any embodiment of the present invention.
[0076] Example 3
[0077] An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed, the method for sensing safety standardization of clothing for power marketing field operations is implemented, including:
[0078] Detect the whole body clothing imaging of the operator from four angles: forward, reverse and two lateral directions;
[0079] Input the whole body clothing image into a pre-trained deep neural network for standardization of work clothing safety perception to perform violation analysis and key point extraction, the deep neural network includes: a human body segmentation network for segmenting the human body, a clothing safety standard analysis network for standardization of work clothing safety perception of the human body foreground; extract the human body foreground from the whole body clothing image through the human body mask predicted by the human body segmentation network; the clothing safety standard analysis network performs regression prediction of the bounding box, category label and violation category label based on the image data in the human body foreground, the clothing safety standard analysis network performs regression prediction of the clothing mask based on the image data in the human body foreground, and the clothing safety standard analysis network performs key point prediction based on the image data in the human body foreground;
[0080] The violation categories are analyzed based on the relationship between the clothing details and the key points between the clothing details and the human body parts. The violation categories predicted by the deep neural network and the violation categories obtained based on the key point analysis are integrated to obtain the final violation situation.
[0081] A computer-readable storage medium provided in an embodiment of the present invention stores a computer program which is not limited to the method operations described above, but can also execute related operations in a method for sensing safety standardization of clothing for power marketing field operations provided in any embodiment of the present invention.
[0082] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, structures or units, which can be electrical, mechanical or other forms.
[0083] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0084] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0085] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for perceiving safety standardization of clothing for power marketing field operations, characterized in that: include: Detect the whole body clothing imaging of the operator from four angles: forward, reverse and two lateral directions; Input the whole body clothing image into a pre-trained deep neural network for standardization of work clothing safety perception to perform violation analysis and key point extraction, the deep neural network includes: a human body segmentation network for segmenting the human body, a clothing safety standard analysis network for standardization of work clothing safety perception of the human body foreground; extract the human body foreground from the whole body clothing image through the human body mask predicted by the human body segmentation network; the clothing safety standard analysis network performs regression prediction of the bounding box, category label and violation category label based on the image data in the human body foreground, the clothing safety standard analysis network performs regression prediction of the clothing mask based on the image data in the human body foreground, and the clothing safety standard analysis network performs key point prediction based on the image data in the human body foreground; The violation categories are analyzed based on the relationship between the clothing details and the key points between the clothing details and the human body parts. The violation categories predicted by the deep neural network and the violation categories obtained based on the key point analysis are integrated to obtain the final violation situation.
2. The method for perceiving safety standardization of clothing for power marketing field operations according to claim 1 is characterized in that: The human segmentation network consists of an encoder network and a corresponding decoder network; The encoder network consists of 13 convolutional layers and 5 pooling layers. In the encoder network, the pooling layers are connected after the 2nd convolutional layer, the 4th convolutional layer, the 7th convolutional layer, the 10th convolutional layer, and the 13th convolutional layer respectively. The decoder network is implemented by 13 convolutional layers and 5 upsampling. Each convolutional layer of the decoder network corresponds to a convolutional layer in the encoder network. In the decoder network, upsampling is set before the 1st convolutional layer, the 4th convolutional layer, the 7th convolutional layer, the 10th convolutional layer, and the 12th convolutional layer. The last convolution operation of the decoder network produces a single-channel feature map. The 3rd, 6th, 9th, and 11th convolutional layers of the decoder network are connected to 1×1 convolutional layers respectively. A deconvolution layer is added after the 1×1 convolutional layer to process the feature map to adapt to the original size of the input. The outputs of all deconvolutional layers and the single-channel feature map of the last convolutional layer of the encoder network are merged, and the merged results are fused through a 1×1 convolutional layer to combine features of different scales. After the features are processed by the activation function, the probability of each pixel belonging to the foreground or background of the human body is displayed, and the human body mask is obtained after the probability is binarized.
3. The method for perceiving safety standardization of clothing for power marketing field operations according to claim 1 is characterized in that: The clothing safety standard analysis network includes: a convolution layer for extracting human foreground features, the convolution layer is connected to a feature pyramid network, the feature pyramid network includes: a multi-level forward feature encoder and a reverse feature encoder based on Resnet or VGG, the forward feature encoder includes five forward feature encoding layers based on Resnet or VGG, the first four forward feature encoding layers are followed by a pooling layer for downsampling, the channel of the output feature of the forward feature encoding layer at any level is twice the channel of the output feature of the feature encoding layer at the previous level, and the output feature of the forward feature encoding layer at any level is The width and height are 1 / 2 of the width and height of the output features of the forward feature coding layer of the previous level; the reverse feature encoder includes five reverse feature coding layers based on Resnet or VGG, the first four reverse feature coding layers are followed by upsampling, the channels of the output features of the reverse feature coding layer at any level are 1 / 2 times the channels of the output features of the reverse feature coding layer of the previous level, the width and height of the output features of the reverse feature coding layer at any level are twice the width and height of the output features of the reverse feature coding layer of the previous level, and a 1×1 convolution layer is set between the corresponding forward feature encoder and the reverse feature encoder to fuse features; The features output by the feature pyramid network are spatially transformed through three Roi Align operations, and the three features after spatial transformation are respectively input into a first head module for regression prediction of bounding boxes, category labels and violation category labels, a second head module for prediction of clothing masks, and a third head module for prediction of key points. The first head module, the second head module and the third head module respectively output predicted bounding boxes, categories and violation categories, clothing masks and key points.
4. The method for perceiving safety standardization of clothing for power marketing field operations according to claim 1 is characterized in that: The clothing safety standard analysis network is trained using training data. The loss function involved in the training is a weighted sum of the cross entropy loss of clothing categories and clothing violation categories, the detection box regression loss, the clothing mask cross loss, and the key point cross entropy loss. The parameters of the deep neural network are updated by stochastic gradient descent with the goal of minimizing the loss function.
5. The method for perceiving safety standardization of clothing for power marketing field operations according to claim 4 is characterized in that: The construction of the training data includes: The full-body clothing images of the power marketing field workers are collected from four angles: forward, reverse and two lateral angles. The proportion of the characters in the full-body clothing images, the occlusions, the image size and the viewing angle of the full-body clothing images are adjusted to enrich the diversity of the full-body clothing images. Manually annotate full-body clothing images, including key point labels, human body masks, and labels of staff clothing; The labels of the staff’s clothing include: bounding box, category label, clothing mask, and violation category label.
6. The method for perceiving safety standardization of clothing for power marketing field operations according to claim 5 is characterized in that: Based on the standard specifications for safety wear in marketing field operations, a set of standards for wear violation levels and violation items are formulated:
1. Major violations include: not wearing a safety helmet, not wearing a work jacket, not wearing work pants, not wearing insulated shoes, and not wearing safety gloves; 2. General violations include: the safety helmet is not fastened with a chin strap, and the work jacket and work pants are damaged; 3. Minor violations include: the chin strap of the safety helmet is not on both sides of the ears, the chin strap buckle of the safety helmet is not standardized, and the buttons of the work clothes are not fully fastened; Construct labels for staff clothing based on the level of wearing violation and violation criteria: To determine whether the staff clothing is a major violation, the manual annotator needs to draw a bounding box for each piece of clothing of the staff and assign a category label. The bounding box indicates the clothing detected in the full-body clothing image, and the bounding box is determined by the coordinates of the two end points of the rectangular diagonal; the category label indicates the type of clothing the staff is wearing; set clothing masks for each type of clothing in the full-body clothing image, and mark each pixel mask to mark whether each element in the full-body clothing image belongs to any type of clothing; When determining that the clothing is work clothing, it is also necessary to determine whether the staff's clothing falls into general violations and minor violations. The manual annotator needs to configure a violation category label for each work clothing of the staff. The violation category label indicates the type of violation of the work clothing worn by the staff; and, configure key point labels for the ears and safety helmets, and configure key point labels for other clothing except safety helmets, so as to perform anomaly analysis based on the relationship between key point labels.
7. The method for perceiving safety standardization of clothing for power marketing field operations according to claim 5 is characterized in that: The process of setting up human body mask and clothing mask annotation includes: obtaining instance contours through contour detection, obtaining preliminary instance masks based on instance contours, and correcting instance contours through manual annotators to improve instance masks.
8. The method for perceiving safety standardization of clothing for power marketing field operations according to claim 1 is characterized in that: Violation categories based on the relationship between clothing details and key points between clothing details and human body parts include: Determine whether the chin strap of the helmet is worn on both sides of the ears based on the relationship between the key points of the ear and the key points of the chin strap of the helmet in the lateral full-body clothing imaging; Determine whether all buttons are fastened based on the relationship between the key points of the sub-buttons and the key points of the mother-buttons in the forward full-body clothing imaging; Determine whether the chin strap buckle is standardized based on the key point relationship of the chin strap buckle and the buckle of the helmet.
9. A safety standardization sensing device for clothing used in power marketing field operations, characterized in that: include: At least one processing unit, the processing unit is connected to the storage unit and the collection unit via a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the method for sensing safety standardization of clothing for power marketing field operations as described in any of claims 1-8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for sensing safety standardization of clothing for power marketing field operations as described in any of claims 1-8 is implemented.
Citation Information
Cited By
Intelligent standard dressing identification method based on target detection and human body key points
CN120472505A