An intelligent recognition method for security status based on attribute knowledge modeling
Through the method based on attribute knowledge modeling, the Transformer encoder and self-attention mechanism are used to solve the occlusion problem of safety state recognition in industrial scenarios, achieving high-precision and robust safety equipment recognition, and improving the recognition rate.
Patent Information
- Application Number
- CN202211376352.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-11-04
AI Technical Summary
The prior art is difficult to effectively identify the safety status of workers in industrial scenarios, especially the occlusion problem, and the existing methods are insufficient in accuracy and robustness, so it is impossible to accurately identify important safety equipment such as insulated gloves and insulated shoes.
Using a method based on attribute knowledge modeling, high-dimensional features are extracted through the backbone network and combined with the Transformer encoder to model the relationship between image features and attribute label vectors. The self-attention mechanism is used to learn the correlation weights between features, and the binary cross entropy loss function is trained to improve the recognition accuracy.
The accuracy and robustness of the identification of safe operation status have been significantly improved, and the recognition rate has reached more than 96% in industrial scenarios, making it highly competitive.
Smart Images

Figure CN115690682B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to an intelligent recognition method for safety states. Background Art
[0002] In recent years, with the continuous development of artificial intelligence and computer vision, traditional manual inspections and video surveillance can no longer meet the needs of the high-risk industrial production fields with frequent accidents, and people are increasingly concerned about the practical applications of innovative technologies in the recognition of safe operation states. In the Chinese patent with the publication number "CN110046557A", a method for detecting safety helmets and safety belts based on deep neural network discrimination is disclosed; in the Chinese patent with the publication number "CN114120237A", a method and system for identifying safety belts at construction sites are provided. Both provide some methods for recognizing safe operation states, but they are all based on the detection of safety helmets or safety belts, and it is difficult to solve the problem of occlusion in actual industrial scenarios that leads to undetected situations, and other safety states are not considered, such as the recognition of insulating gloves and insulating shoes necessary for electrical scenarios. At the same time, there are also problems that the accuracy is not enough to be applied to actual scenarios and the efficiency needs to be improved.
[0003] Meanwhile, with the wide deployment of security cameras, how to perform efficient pedestrian attribute recognition in surveillance scenarios has received extensive attention. Pedestrian attribute recognition is to use technologies such as computer vision to intelligently process pedestrian pictures, so as to obtain the attribute categories contained in a certain pedestrian, such as age, gender, clothing, etc. However, most of the current pedestrian attribute recognition technologies directly classify the extracted high-dimensional features, or use attention mechanisms to extract different features for different human body parts for recognition. These methods ignore the relationships existing between attributes and cannot associate attributes with more appropriate high-dimensional features. At the same time, most of these methods use convolutional neural networks for feature extraction, losing global information, and the recognition effects for some global features, such as age and gender, are not ideal. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides an intelligent recognition method for safety states based on attribute knowledge modeling. First, pictures of workers in surveillance videos are extracted and preprocessed, then high-dimensional feature extraction and generation of attribute label vectors are performed on the pictures based on a backbone network. Next, a Transformer encoder is applied to model the relationship between image features and attribute label vectors. Finally, the results of the feature and attribute outputs are processed, and the error loss and accuracy are calculated to train the network, completing the intelligent recognition of the safety states of workers. The present invention effectively improves the accuracy and robustness of attribute recognition, enabling artificial intelligence algorithms to better play a role in industrial safety.
[0005] The technical solution adopted by the present invention to solve the technical problem includes the following steps:
[0006] Step 1: Extract and preprocess worker images from surveillance videos;
[0007] Step 1-1: Capture surveillance images from the long surveillance video, use the target detection algorithm to detect workers and crop the images, obtain images containing only workers as the data set, and divide the data set into a training set and a test set according to the set ratio;
[0008] Step 1-2: Screen the images in the training set and label them with the safety operation status attribute labels;
[0009] Step 1-3: Preprocess the images of the training set and augment the data;
[0010] Step 2: Extract high-dimensional features of the image and generate attribute label tensors based on the backbone network;
[0011] Step 2-1: Use the backbone network to extract high-dimensional features;
[0012] The pictures in the training set Input to the backbone network for feature extraction and then output image feature tensor Where H and W represent the height and width of the input image, respectively, and h, w, and d represent the height, length, and number of channels of the output tensor, respectively; and an embedding layer is used to initialize the extracted image feature tensor to generate a learnable position encoding sequence P = {p1, p2, ..., p h×w},in:
[0013] p i =w0+w1q s
[0014] w1 represents the learnable parameter, w0 represents the bias, q s ∈Q;
[0015] Step 2-2: Encode all attribute labels of each image into a tensor through an embedding layer l is the number of attribute labels, and an attribute label tensor is obtained L represents the semantic information of all attribute labels included in the image;
[0016] Step 3: Use the Transformer encoder to model the relationship between the image feature tensor and the attribute label tensor;
[0017] Step 3-1: Fusion of the image feature tensor Q and the position encoding sequence P to generate a new feature tensor and
[0018] Z = Q + P
[0019] Let K = {z1, z2, …, z h×w , l1, l2, …, l l} represent the set of the feature tensor Z and the attribute label tensor L, and they are sent into the Transformer encoder together;
[0020] Step 3-2: In the Transformer encoder, learn the correlation weights between each feature of the input K through the self-attention mechanism;
[0021] Let α ij represent the correlation weight between feature k i ∈ K and k j ∈ K, and the calculation method of α ij is as follows:
[0022]
[0023] According to α ij and a non-linear layer ReLU, update the feature tensor:
[0024]
[0025] where, W Q , W K , W v represent three learnable vector matrices respectively, b1 and b2 represent bias vectors, and H = h × w + l represents the length of the input vector set;
[0026] The final output of the Transformer encoder is the feature vector K' = {z'1, z'2, …, z' h×w , l'1, l'2, …, l' l} of the relationship between attributes and features, where Z' = {z'1, z'2, …, z' h×w} represents the output of the image feature tensor, and L' = {l'1, l'2, …, l' l} represents the output of the attribute label tensor;
[0027] Step 4: Process the results of the features and attributes output in Step 3 and conduct training;
[0028] Step 4-1: For the image feature tensor Z', through dimensional transformation to Then, after activation by the average pooling layer and the fully connected layer, the final output output f is obtained as follows:
[0029] output f= σ(FC(avgpool(Z')))
[0030] Wherein, avgpool represents the average pooling layer, FC represents the fully connected layer, and σ represents the sigmoid activation function;
[0031] Step 4-2: For the attribute label tensor L', obtain the final prediction probability through an independent feed-forward network FFN; FFN includes a simple linear layer, and the calculation formula is as follows:
[0032] output l = FFN(l' i ) = σ((w i ·l' i ) + b i )
[0033] Wherein, w i represents the learnable weight, b i is a bias vector, and σ represents the sigmoid activation function;
[0034] Step 4-3: During the training process, use the binary cross-entropy loss function as the loss function for safety attribute recognition, and the formula is as follows:
[0035]
[0036] Wherein, and l respectively represent the result predicted by the model and the true label, and M represents the total number of attribute labels;
[0037] Calculate the loss Loss f of the image feature vector and the loss Loss l of the attribute label vector respectively, and obtain the final loss function:
[0038] Loss = λLoss f + Loss l
[0039] Step 4-4: During the inference process, obtain the possible probability of each attribute category through the following formula to judge the safe working state of the worker:
[0040] output = maximum(output f , output l )
[0041] Preferably, the safe working state attributes include safety helmets, safety belts, reflective vests, insulating gloves, insulating shoes, masks, and smoking attributes.
[0042] Preferably, the preprocessing and data augmentation include scale normalization, random horizontal flipping, random rotation, and random color jittering.
[0043] Preferably, the backbone network is a convolutional neural network CNN, a deep residual network Resnet, or a vision-based ViT.
[0044] The beneficial effects of the present invention are as follows:
[0045] 1. The present invention first proposes to transform the problem of identifying the safe operation state in the industrial field into a problem of identifying the semantic attributes of an image, which to a certain extent solves the disadvantages of the prior art, such as low detection accuracy and poor robustness for the safe operation state, and provides a new solution for the intelligent identification of the safe operation state in the industrial field.
[0046] 2. The present invention proposes an attribute knowledge modeling method based on a Transformer encoder, applying the self-attention mechanism to attribute recognition to fully explore the relationships between various attributes and image features and model them. Finally, five evaluation metrics, namely mean average precision (mA), accuracy, precision, recall, and F1-score (F1_Score), are used to experimentally verify the accuracy and robustness of the proposed algorithm in the actual industrial production scenario on the constructed industrial scenario dataset. Its recognition rate on the test set reaches an average precision of over 96%, showing strong competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the recognition method based on attribute knowledge modeling proposed by the present invention.
[0048] Figure 2 It is an example diagram of the industrial scenario safety state recognition dataset collected by the present invention.
[0049] Figure 3 It is a comparison diagram of the attention of certain attributes between the embodiment of the present invention and the residual convolutional network. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The present invention will be further described below in conjunction with the drawings and embodiments.
[0051] Aiming at the disadvantages of the prior art, such as low detection accuracy and poor robustness for the safe operation state in the industrial field, according to the actual scenario requirements, the present invention first proposes to transform the problem of identifying the safe operation state into a problem of identifying the semantic attributes of an image, and proposes a safe state intelligent recognition technology based on attribute knowledge modeling for the limitations existing in the attribute recognition algorithm, effectively improving the accuracy and robustness of attribute recognition, aiming to make the artificial intelligence algorithm play a better role in industrial safety.
[0052] A security state intelligent recognition method based on attribute knowledge modeling, comprising the following steps:
[0053] Step 1: Extract and preprocess the worker pictures in the surveillance video;
[0054] Step 1-1: Intercept surveillance pictures from the long surveillance video, use the object detection algorithm to detect workers and crop the pictures, obtain the pictures containing only workers as the data set, and divide the data set into a training set and a test set according to a set ratio;
[0055] Step 1-2: Screen the pictures in the training set and label the safety operation state attribute labels; the safety operation state attributes include multiple types of attributes such as safety helmets, safety belts, reflective vests, insulating gloves, insulating shoes, masks, smoking, etc.;
[0056] Step 1-3: Preprocess and augment the data of the pictures in the training set, including scale normalization, random horizontal flipping, random rotation, random color jitter, etc.;
[0057] Step 2: Based on the backbone network, perform high-dimensional feature extraction on the pictures and generate the attribute label tensors;
[0058] Step 2-1: Use the backbone network to perform high-dimensional feature extraction; the backbone network can be a convolutional neural network (CNN), such as a deep residual network (Resnet), or a vision-based ViT (Vision Transformer);
[0059] Input the pictures of the training set into the backbone network for feature extraction and then output the image feature tensor h, w, and d respectively represent the height, length, and number of channels of the output tensor; and through an embedding layer, a learnable position encoding sequence P = {p1, p2,..., p h×w} is initialized for the extracted image feature tensor, where:
[0060] p i = w0 + w1q s
[0061] w1 represents the learnable parameter, w0 represents the bias, and q s ∈Q;
[0062] Step 2-2: Encode all the attribute labels of each picture into a tensor through an embedding layer l is the number of attribute labels, and an attribute label tensor is obtained L represents the semantic information of all the attribute labels included in the picture;
[0063] Step 3: Use a Transformer encoder to model the relationship between the image feature tensor and the attribute label tensor;
[0064] Step 3-1: Fuse the image feature tensor Q with the position encoding P to generate a new feature tensor and
[0065] Z = Q + P
[0066] Let K = {z1, z2, …, z h×w , l1, l2, …, l l} represent the set of the feature tensor Z and the attribute label tensor L, and send them into the Transformer encoder together;
[0067] Step 3-2: In the Transformer encoder, learn the correlation weights between each feature of the input K through the self-attention mechanism;
[0068] Let α ij represent the correlation weight between feature k i ∈ K and k j ∈ K, and the calculation method of α ij is as follows:
[0069]
[0070] Update the feature tensor according to α ij and a non-linear layer ReLU:
[0071]
[0072] where, W Q , W K , W v represent three learnable vector matrices respectively, b1 and b2 represent bias vectors, and H = h × w + l represents the length of the input vector set;
[0073] The final output of the Transformer encoder is the feature vector K' = {z'1, z'2, …, z' h×w , l'1, l'2, …, l' l} of the relationship between attributes and features, where Z' = {z'1, z'2, …, z' h×w} represents the output of the image feature tensor, and L' = {l'1, l'2, …, l' l} represents the output of the attribute label tensor;
[0074] Step 4: Process the results of the features and attributes output in Step 3 and perform training;
[0075] Step 4-1: For the image feature tensor Z', through dimensional transformation to Then, after activation by the average pooling layer and the fully connected layer, the final output output is obtained f , as follows:
[0076] output f = σ(FC(avgpool(Z')))
[0077] where, avgpool represents the average pooling layer, FC represents the fully connected layer, and σ represents the sigmoid activation function;
[0078] Step 4-2: For the attribute label tensor L', the final predicted probability is obtained through an independent feed-forward network FFN; FFN contains a simple linear layer, and the calculation formula is as follows:
[0079] output l = FFN(l' i ) = σ((w i ·l' i ) + b i )
[0080] where, w i represents the learnable weight, b i is a bias vector, and σ represents the sigmoid activation function;
[0081] Step 4-3: During the training process, the binary cross-entropy loss function is used as the loss function for safety attribute recognition, and the formula is as follows:
[0082]
[0083] where, and l respectively represent the result predicted by the model and the true label, and M represents the total number of attribute labels;
[0084] Calculate the loss Loss f of the image feature vector and the loss Loss l of the attribute label vector respectively, and obtain the final loss function:
[0085] Loss = λLoss f + Loss l
[0086] Step 4-4: During the inference process, the possible probability of each attribute category is obtained through the following formula to discriminate the safe working state of the worker:
[0087] output = maximum(output f , output l) Specific embodiments:
[0089] 1. Image acquisition and preprocessing
[0090] In this embodiment, first, surveillance pictures are collected from the surveillance long video, and the target detection algorithm is used to detect workers and crop the pictures to obtain images containing only workers as the dataset. The dataset is divided into a training set and a test set according to 6:4; then the pictures are screened and labeled. The safety operation status attribute values include multiple types of attributes such as safety helmets, safety belts, reflective vests, insulating gloves, insulating shoes, masks, smoking, etc.; finally, the collected worker pictures are preprocessed and data augmented, including scale normalization, random horizontal flipping, random rotation, random color jitter, etc.
[0091] 2. Feature extraction and attribute vector generation
[0092] In this embodiment, the worker images after acquisition and preprocessing are input into the backbone network Resnet50 for feature extraction to obtain an output tensor where h, w, and d represent the height, length, and number of channels of the output tensor respectively. In this embodiment, Q is regarded as a vector, where where n = h×w. Therefore, an image feature tensor can be obtained after passing through the backbone network which represents the feature vector of each region of the original image. At the same time, in this example, a learnable position encoding sequence P = {p1, p2,..., p h×w} is initialized for the extracted local image feature tensor through an embedding layer, where:
[0093] p i = w0 + w1q s
[0094] w1 represents the learnable parameter, w0 represents the bias, and q s ∈Q;
[0095] For each image, in this embodiment, all the attribute labels contained in it are encoded into a tensor L = {l1, l2,..., l l} through an embedding layer, so an attribute label tensor can be obtained which represents the semantic information of all the attribute labels included in the image.
[0096] 3. Image feature and attribute knowledge modeling
[0097] A Transformer is a deep neural network mainly based on the self-attention mechanism. It was initially proposed for modeling long-sequence learning problems and is applied in the field of natural language processing. Inspired by the powerful representation ability of the Transformer, researchers proposed to extend the Transformer to computer vision tasks such as image classification and object detection. Compared with other network types (such as convolutional networks and recurrent networks), the Transformer-based model is good at capturing the relationship dependencies between different and long-distance vectors and shows competitive or even better performance on various visual benchmarks. In this embodiment, the Transformer is applied to the attribute recognition task, and through its powerful receptive field, it learns the relationship between attributes and features and models them to extract more expressive features and learn the entangled mutual relationships between attributes, greatly improving the accuracy of attribute recognition.
[0098] In this embodiment, the image feature tensor Q is first fused with the position encoding P to generate a new feature tensor Z = {z1, z2, …, z h×w}, where
[0099] Z = Q + P
[0100] Let K = {z1, z2, …, z h×w , l1, l2, …, l l} represent the set of the image feature tensor and the attribute label tensor, which are fed into the Transformer encoder together. In the encoder, the weights related to each feature embedding of the input are learned through the self-attention mechanism. Let α ij represent the correlation weight between feature k i ∈ K and k j ∈ K. The calculation method of α ij is as follows:
[0101]
[0102] Then, according to the calculated weights between k i ∈ K and k j ∈ K and a non-linear layer ReLU, the feature tensor is updated:
[0103]
[0104] Among them, W Q , W K , W vrespectively represent three learnable vector matrices, b1 and b2 represent bias vectors, and m represents the length of the input vector set, that is, h×w + l. Finally, a new feature vector K' = {z'1, z'2, …, z' h×w , l'1, l'2, …, l' l} is obtained that learns the global information and the relationship between attributes and features.
[0105] 4. Feature classification training and recognition, calculate the error loss and accuracy, and train the network;
[0106] In step 3, the output of the final Transformer encoder can be obtained as K' = {z'1, z'2, …, z' h×w , l'1, l'2, …, l' l}. In this embodiment, Z' = {z'1, z'2, …, z' h×w} is defined to represent the output of the image feature vector, and L' = {l'1, l'2, …, l' l} is defined to represent the output of the attribute label vector; for the image feature vector, after dimensionality transformation to and then activated through the average pooling layer and the fully connected layer, the output output f can be obtained, as shown below:
[0107] output f = σ(FC(avgpool(Z')))
[0108] where avgpool represents the average pooling layer, FC represents the fully connected layer, and σ represents the sigmoid activation function; for the attribute label vector, the final predicted probability is obtained through an independent feed-forward network (FFN). The FFN contains a simple linear layer, and its calculation formula is as follows:
[0109] output l = FFN(l' i ) = σ((w i ·l' i ) + b i )
[0110] where w i represents the learnable weight, b i is a bias vector, and σ represents the sigmoid activation function;
[0111] During the training process, the binary cross-entropy loss function is used as the loss function for security attribute recognition, and the formula is as follows:
[0112]
[0113] Among them, and l respectively represent the result predicted by the model and the true label, M represents the number of attribute labels, and σ represents the sigmoid activation function. Calculate the loss Loss of the image feature vector f and the loss Loss of the attribute label vector l , and the final loss function can be obtained:
[0114] Loss = λLoss f + Loss l
[0115] In this embodiment, setting λ to 0.2 can achieve the highest accuracy.
[0116] During the inference process, the possible probability of each attribute category is obtained through the following formula element maximum value to determine the safe operation state of the worker:
[0117] output = maximum(output f , output l )
[0118] Finally, the present invention uses five evaluation indicators, namely mean average precision (mA), accuracy, precision, recall, and F1-score, to experimentally verify the accuracy and robustness of the proposed method in the actual industrial production scenario on the industrial scenario dataset constructed as shown in Figure 2 . As shown in Figure 3 , it can be seen that the Transformer attribute knowledge modeling model proposed by the present invention has a more accurate attention area positioning for attributes such as safety helmets and safety belts than the residual convolutional neural network model, and its recognition rate on the test set reaches an average precision of more than 96%, which has strong competitiveness.
Claims
1. A method for intelligent identification of security status based on attribute knowledge modeling, characterized in that, The steps include: Step 1: Extract and preprocess worker images from surveillance videos; Step 1-1: Capture surveillance images from the long surveillance video, use the target detection algorithm to detect workers and crop the images, obtain images containing only workers as the data set, and divide the data set into a training set and a test set according to the set ratio; Step 1-2: Screen the images in the training set and label them with the safety operation status attribute labels; Step 1-3: Preprocess the images of the training set and augment the data; Step 2: Extract high-dimensional features of the image and generate attribute label tensors based on the backbone network; Step 2-1: Use the backbone network to extract high-dimensional features; Input the images of the training set into the backbone network for feature extraction and output the image feature tensor where H and W represent the height and width of the input image respectively, and h, w, and d represent the height, length, and number of channels of the output tensor; and through an embedding layer, a learnable position encoding sequence P = {p1, p2, …, p h×w} is initialized for the extracted image feature tensor, where: p i = w0 + w1q s w1 represents learnable parameters, w0 represents the bias, q s ∈Q; Step 2-2: Encode all the attribute labels of each image into a tensor \(L = \{l_1, l_2, \ldots, l\) l \}, where \(l\) is the number of attribute labels, obtaining an attribute label tensor \(L\) represents the semantic information of all the attribute labels included in the image; Step 3: Use the Transformer encoder to model the relationship between the image feature tensor and the attribute label tensor; Step 3-1: Fuse the image feature tensor Q with the position encoding sequence P to generate a new feature tensor Z = {z1, z2, …, z h×w}, and Z=Q+P Let \(K = \{z_1, z_2, \ldots, z h×w , l_1, l_2, \ldots, l l \}\) represent the set of the feature tensor \(Z\) and the attribute label tensor \(L\), which are fed into the Transformer encoder together; Step 3-2: In the Transformer encoder, the relevant weights between each feature of the input K are learned through the self-attention mechanism; Let α ij represent the correlation weight between feature k i ∈K and k j ∈K, and the calculation method of α ij is as follows: According to α ij and update the feature tensor with a non-linear layer ReLU: Among them, W Q , W K , W v respectively represent three learnable vector matrices, b1 and b2 represent bias vectors, and H = h × w + l represents the length of the input vector set; The output of the final Transformer encoder is the feature vector K' = {z'1, z'2, …, z' h×w , l'1, l'2, …, l' l} for the relationship between attributes and features, where Z' = {z'1, z'2, …, z' h×w} represents the output of the image feature tensor, and L' = {l'1, l'2, …, l' l} represents the output of the attribute label tensor; Step 4: Process the features and attributes output from step 3 and perform training; Step 4-1: For the image feature tensor Z', through dimensional transformation to Then, after activation by the average pooling layer and the fully connected layer, the final output output is obtained f , as follows: output f = σ(FC(avgpool(Z'))) Among them, avgpool represents the average pooling layer, FC represents the fully connected layer, and σ represents the sigmoid activation function; Step 4-2: For the attribute label tensor L', the final prediction probability is obtained through an independent feed-forward network FFN; FFN contains a simple linear layer, and the calculation formula is as follows: output l = FFN(l' i ) = σ((w i ·l' i ) + b i ) where, w i represents learnable weights, b i is a bias vector, and σ represents the sigmoid activation function; Step 4-3: During the training process, the binary cross entropy loss function is used as the loss function for security attribute identification. The formula is as follows: Among them, and l respectively represent the result predicted by the model and the true label, and M represents the total number of attribute labels; Calculate the loss of the image feature vector separately f The loss with the attribute label vector l , and obtain the final loss function: Loss=λLoss f +Loss l Step 4-4: In the reasoning process, the possible probability of each attribute category is obtained through the following formula to determine the worker's safe working status: output = maximum(output f , output l ).
2. The intelligent recognition method for security status based on attribute knowledge modeling according to claim 1, characterized in that The safety operation status attributes include safety helmets, safety belts, reflective clothing, insulating gloves, insulating shoes, masks, and smoking attributes.
3. A method for intelligent recognition of security status based on attribute knowledge modeling according to claim 1, characterized in that The preprocessing and data augmentation include scale normalization, random horizontal flipping, random rotation and random color jittering.
4. The intelligent recognition method for security status based on attribute knowledge modeling according to claim 1, characterized in that The backbone network is a convolutional neural network (CNN) or a deep residual network (Resnet) or a vision-based ViT.
Citation Information
Patent Citations
Safety helmet and safety belt detection method based on deep neural network discrimination
CN110046557A
Construction site safety belt identification method and system
CN114120237A
Picture aesthetics description modeling and description method and system based on aesthetics attribute retrieval
CN113610128A
Label-to-label-based multi-attribute prediction method and device, equipment and medium
CN114565017A