Perimeter intrusion detection and personnel attribute analysis method for safety production
By building a data fusion model and edge device deployment, the accuracy and speed of perimeter intrusion detection and personnel attribute analysis in industrial scenarios are solved, and efficient industrial production safety monitoring is achieved.
Patent Information
- Application Number
- CN202510422760.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology cannot effectively target perimeter intrusion detection and personnel attribute analysis in industrial scenarios, and there are problems with limited computing power of edge equipment and the inability to take into account both accuracy and speed.
Build a data fusion model, including a person-unmanned classification network, a person-detection network and a feature extraction network, use Tensorrt technology for quantitative compression, deploy it to edge devices for real-time detection and analysis, and optimize model performance through data augmentation and difficult sample mining.
It improves the accuracy and speed of perimeter intrusion detection and personnel attribute analysis in industrial scenarios, realizes efficient monitoring of the factory environment, and reduces the computing burden of edge equipment.
Smart Images

Figure CN120279487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent security, and in particular to a perimeter intrusion detection and personnel attribute analysis method for work safety. Background Art
[0002] As the boundary scope of production and operation units, the definition and setting of the perimeter vary according to industry scenarios. From construction sites, high-speed rails to large-scale manufacturing industries, the perimeter not only includes physical boundaries but also involves the safety scope of large-scale production workpieces. To address these challenges, perimeter intrusion detection and personnel attribute analysis systems have emerged to achieve automatic identification and tracking of intrusion events and personnel attribute analysis.
[0003] However, there are significant differences in the detection and attribute analysis of factory personnel in most scenarios: the characteristics and perspectives of factory personnel are different from common situations, and existing methods cannot effectively target industrial perimeter scenarios. Secondly, although edge computing devices play an increasingly important role in industrial production, the computing power of edge devices is limited, and the existing solutions cannot balance the issues of accuracy and speed.
[0004] Therefore, forming a detection model for factories to achieve high-performance perimeter intrusion detection and personnel attribute analysis and meet the real-time monitoring requirements of large-scale complex areas is an important way to promote work safety in industrial production and greatly reduce the risk coefficient of industrial accidents. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a perimeter intrusion detection and personnel attribute analysis method for work safety, which can timely and accurately identify and analyze the personnel entering the perimeter area and give early warnings according to the results to improve the safety of industrial production.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A perimeter intrusion detection and personnel attribute analysis method for work safety includes the following steps:
[0008] Step 1. Data fusion: Perform data fusion on the perimeter intrusion and pedestrian attribute analysis systems, construct a training set for industrial work safety, and provide it to Step 2;
[0009] Step 2. Construct a presence / absence classification network, a personnel detection network, and a feature extraction network, and use the dataset in Step 1 to train and fine-tune them to improve the network's adaptability to the specific environment of the factory.
[0010] Step 3. Model Inference: Deploy the above network model to the industrial scenario, including three modules: the human / personless classification module, the personnel detection module, and the personnel clothing recognition module. Among them, the human / personless classification module deploys a human / personless classification network, the personnel detection module deploys a personnel detection network, and the personnel clothing recognition module deploys a feature extraction and matching network. Use the corresponding tooling feature vector dataset made according to the actual scenario to achieve the corresponding pedestrian attribute analysis, and give corresponding warnings according to the actual situation.
[0011] The video frames determined to have people by the human / personless classification module are provided to the personnel detection module for pedestrian detection, and the personnel clothing recognition module determines that the matched personnel clothing belongs to the tooling category that must be worn to enter the perimeter area. When multiple consecutive frames exceed the preset threshold, it is determined that the personnel entering the perimeter belong to normal work.
[0012] Step 4. Deep Learning Inference Acceleration and Model Deployment: Use Tensorrt technology to quantize and compress each model to improve the inference speed of the model on edge devices while ensuring accuracy. Select edge devices according to the computing resource constraints for model deployment and system integration.
[0013] Most images in traditional public datasets are images in natural scenes or urban environments and are not specifically for factory environments; in addition, although such datasets contain a large number of images, annotations, and a wide range of object categories, such as people, animals, vehicles, commodities, etc., there are domain differences between factory personnel detection and most object categories; the characteristics, perspectives, etc. of factory personnel (such as postures, occlusions, and overlaps) are also different from the general personnel in the vast majority of public datasets. Therefore, simply using common public datasets cannot fully cover the special needs of factory scenarios.
[0014] Specifically, in Step 1, data fusion includes two steps: data augmentation and difficult sample screening;
[0015] For the data augmentation part, for the factory-related data collected, use a rectangular box to select the position of people as the annotation rule, roughly mark the individuals and approximate position information in the image through an automated tool, and then perform manual refinement and correction for precise positioning, and establish an xml annotation file to hierarchically store all attributes of the annotation objects;
[0016] Then, perform operations such as random direction transformation (vertical and horizontal flips and random angle rotations), size transformation (scaling), image sharpening or blurring, and brightness adjustment on the factory images in sequence to simulate different shooting conditions and changes;
[0017] Finally, it is integrated with the public MS COCO dataset and the CrowdHuman dataset for crowd detection to construct a dataset with industrial workers as the information main body and containing various environmental backgrounds of the factory.
[0018] In the difficult sample mining part, by mining and focusing on pedestrian samples that are more challenging, more difficult to classify or predict for the model, the performance, generalization ability and robustness of the model in the pedestrian classification task can be improved targeted.
[0019] First, use the initial dataset for model training, and use the trained model for prediction to obtain the object detection results, including the predicted bounding boxes and categories.
[0020] Then, divide the difficult samples according to the IOU value, and assign higher weights according to their IOU values, so as to pay more attention to these difficult examples in the next round of training.
[0021] Next, use the weighted samples (including difficult samples and ordinary samples) for the next round of model training, and perform multiple iterative trainings to further optimize the performance of the model.
[0022] Specifically, in step 2,
[0023] The person presence classification network uses the PULC ultra-lightweight image classification model (based on the PP-LCNet network, prior art);
[0024] The person detection network uses an improved YOLOv8 model;
[0025] The feature extraction network uses the PP-ShiTu model;
[0026] Use the data obtained in step 1 to train the above networks.
[0027] Specifically, in step 3,
[0028] Furthermore, the person presence classification network detects whether there are people in the video stream in the camera and determines whether to perform subsequent operations.
[0029] Specifically, image classification is achieved through the following method:
[0030] First, obtain the RTSP video stream of the network camera, and extract video frames from it as the input images of the algorithm.
[0031] Preprocess the image with PyTorch, convert it to the limited size of the input image by the PULC person presence classification model structure, convert the channel dimension and data type into the form required by the picture, and finally perform pixel normalization processing.
[0032] Use the manned and unmanned classification network for inference to obtain the inference result.
[0033] For video frames with people, provide them to the person detection network for pedestrian detection; otherwise, end this detection process.
[0034] Because in the actual application in the factory, there are mostly no people in the perimeter in most cases. Doing so can improve the real-time processing efficiency of the video and reduce the loss of edge devices.
[0035] Furthermore, the person detection network can accurately detect various types of intruding people in complex industrial scenarios.
[0036] Specifically, object detection is achieved by improving the feature pyramid network of YOLOv8 as follows:
[0037] On the one hand, optimize it from unidirectional top-down fusion to bidirectional feature flow as the feature fusion strategy to ensure the efficient flow of information between different scales: high-level feature information can be transmitted downward to improve the model's detection ability for smaller humans; low-level feature information can be transmitted upward to enhance the model's perception ability for larger humans.
[0038] On the other hand, perform weighted feature fusion to obtain the output features of each layer. Through learnable weights, let the model dynamically adjust the contribution of each scale feature by learning the weights, and at the same time avoid the information imbalance problem caused by the original simple addition in YOLOv8. When learning the weights, first take the exponential of the learnable weights and then perform Softmax normalization to ensure that the sum of the contribution degrees of all features is 1. Compared with directly using the weights, taking the exponential before normalization can ensure that all weights are positive numbers and avoid the negative weight problem.
[0039] Furthermore, prune the model to remove the parts in the neural network that contribute less to the prediction. Set thresholds for the absolute value size of each weight for pruning; set thresholds for the L1 norm of each channel to delete the channels with less importance. After pruning, fine-tune the model to restore the accuracy.
[0040] Use PyTorch to preprocess the image: adjust it to the specified size of the input image for YOLOv8, convert the channel dimension and data type into the form required by the image, and finally perform pixel normalization processing.
[0041] Use the improved YOLOv8 for inference to obtain the inference result.
[0042] Perform post-processing. Process the format of the model output and the format of the bounding boxes, and then perform non-maximum suppression to filter the overlapping bounding boxes in the model output. Finally, restore the inference result and draw both the prediction boxes and the class labels on the image.
[0043] The personnel clothing recognition module creates a corresponding work clothing feature vector dataset according to the actual scenario, uses a feature extraction and matching network based on metric learning to achieve corresponding pedestrian attribute analysis, and issues corresponding warnings according to the actual situation;
[0044] Specifically, the feature extraction and matching network is obtained through the following steps:
[0045] Step 3.1: For the clothing that may appear on factory personnel, collect template images and create a feature vector dataset for matching;
[0046] Step 3.2: For the results of pedestrian detection, use the feature extraction network obtained in Step 2 to extract features;
[0047] 1) Target image cropping: According to the person detection frame obtained by the person detection module, crop the work clothing detection area;
[0048] 2) Feature acquisition and storage: Send the cropped image into the feature extraction model for inference
[0049] Step 3.3: Match the extracted features with the feature vector dataset;
[0050] Furthermore, in Step 3.1, since in practical applications, the video mainly comes from real-time monitoring within the factory, in order to make the feature extraction framework more adaptable to the recognition task of pedestrian attribute targets for perimeter safety production, common work clothing used in the factory during actual production, safety protection clothing used in special scenarios, and various possible daily outfits are collected, and a new feature vector dataset suitable for perimeter safety production and personnel attribute analysis is constructed through methods such as screening and annotation.
[0051] Furthermore, in Step 3.2, the feature extraction part needs to complete two aspects in sequence: target image cropping, feature acquisition and storage.
[0052] First, traverse all the bounding boxes containing prediction results obtained from pedestrian detection in Step 2, extract the detection boxes with the detection category of person, and crop them from the original video frame.
[0053] Then, only intercept the head and upper body parts of the human detection box as the initial input of the feature extraction network.
[0054] Specifically, first use PyTorch to preprocess the image: convert it to the defined size of the input image by the feature extraction network, convert the channel dimension and data type into the form required by the image, and perform pixel normalization processing.
[0055] Load the feature extraction model, predict the input image, and obtain the prediction result.
[0056] Normalize the obtained 512-dimensional vector. Calculate the Euclidean norm of the feature result, and divide each element in the feature vector array by the norm value to achieve the normalization operation.
[0057] Further, in step 3.3, the feature matching part needs to match the normalized result with the feature vector dataset in step 3.1, and perform the comparison only in the most similar cluster to achieve efficient vector retrieval and tooling feature vector matching.
[0058] First, use K-means clustering to divide all vectors into k clusters and obtain the clustering center of each cluster.
[0059] Then, each database vector v i is assigned to the nearest clustering center, that is, each vector is stored only in the nearest cluster. Build an inverted file structure, and each clustering center c j maintains a list to store all vectors belonging to this cluster.
[0060] For the target vector, retrieve the 5 nearest clusters and find the nearest neighbor in their vector lists as the matching result.
[0061] Perform condition checking, requiring both conditions to be met simultaneously: matching a specific tooling (the tooling category that must be worn when entering the perimeter area), and when exceeding a certain threshold for multiple consecutive frames, determine that the person entering the perimeter belongs to normal work.
[0062] Specifically, in step 4, the deep learning inference acceleration and model deployment are obtained through the following methods:
[0063] First, perform inference acceleration. Convert the models in Pytorch and PaddlePaddle formats to onnx format models, and then perform TensorRT inference optimization on the onnx format models to convert them into engine models that can be used by TensorRT.
[0064] Then, select a suitable edge device for model deployment. From the production site to the RTSP video stream that can be used by the edge device for processing, hardware transmission configuration and deployment are required: first, the camera at the production site obtains video information from the site, then the information is transmitted to the communication base station through the switch, and then the communication base station uploads it to the server, and finally the RTSP data stream that can be processed at the edge device terminal is obtained.
[0065] Beneficial effects
[0066] Compared with the prior art, the present invention has the following advantages:
[0067] 1. The present invention proposes a perimeter intrusion detection and personnel attribute analysis method for work safety, which realizes image classification, pedestrian detection, feature extraction and matching of RTSP video streams through software and hardware deployment on edge devices, realizes personnel attribute analysis and timely warning for personnel intruding into the perimeter range, and improves work safety.
[0068] 2. The present invention performs data fusion and dataset adjustment on the perimeter intrusion and pedestrian attribute analysis system. And trains an image classification network based on the PP-LCNet network and a pedestrian detection network based on CSPNet, and improves the accuracy of personnel detection by 9.08% on the business dataset.
[0069] 3. The present invention appropriately crops the results of pedestrian detection, uses a feature extraction network based on metric learning to obtain features and then performs vector retrieval and matching, and improves the accuracy of personnel attribute analysis by 7.42% on the business dataset.
[0070] 4. The present invention improves the detection efficiency from two aspects: the system deployment price and the inference acceleration of the Tensorrt framework. For the image classification model, pedestrian detection model, and feature extraction model, the single-frame processing speed is increased by 75.6%, 88.9%, and 87.6%. On the premise of ensuring the accuracy of system detection, the detection speed is greatly improved, which is applicable to the actual production system. Description of the Drawings
[0071] Figure 1 is the overall flowchart of the perimeter intrusion detection and personnel attribute analysis method for work safety of the method of the present invention;
[0072] Figure 2 is the schematic flowchart of data augmentation of the method of the present invention;
[0073] Figure 3 is the schematic flowchart of hard sample mining of the method of the present invention;
[0074] Figure 4 is the schematic implementation framework diagram of the human / personless classification module of the method of the present invention;
[0075] Figure 5 is the schematic diagram of the inference part of the personnel detection module of the method of the present invention;
[0076] Figure 6 is the schematic implementation framework diagram of the personnel clothing recognition module of the method of the present invention;
[0077] Figure 7 is the schematic implementation framework diagram of the system deployment of the method of the present invention. Detailed Embodiments
[0078] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0079] Embodiment:
[0080] This embodiment provides a perimeter intrusion detection and personnel attribute analysis method for safe production, which can identify and analyze the personnel entering the perimeter area in a timely and accurate manner, and issue early warnings based on the results to improve the safety of industrial production. The framework schematic diagram of this method is as Figure 1 shown, and specifically includes the following steps:
[0081] Step 1. Data fusion: Perform data fusion on the perimeter intrusion and pedestrian attribute analysis system to construct a training set for industrial safe production.
[0082] Specifically:
[0083] The data fusion includes data augmentation and difficult sample mining.
[0084] For the data augmentation part (such as Figure 2 ), for the factory-related data collected, use a rectangular box to select the position of the person as the annotation rule. Through an automated tool, roughly mark the individuals and approximate position information in the image, and then perform manual refinement and correction for accurate positioning. Establish an xml annotation file to hierarchically store all the attributes of the annotation objects. Then, perform operations such as random direction transformation (vertical and horizontal flipping and random angle rotation), size transformation (scaling), image sharpening or blurring, and brightness adjustment on the factory images in sequence to simulate different shooting conditions and changes. Finally, integrate with the public MS COCO dataset and the CrowdHuman dataset for crowd detection to construct a new dataset with industrial personnel as the information main body and including various factory environmental backgrounds.
[0085] For the difficult sample mining part, specifically, for the results of data augmentation: (such as Figure 3 )
[0086] First, use the initial dataset for model training and prediction to obtain the object detection results.
[0087] Set the IOU threshold to 0.5. Samples with an IOU greater than the threshold but misclassified between the predicted bounding box and the true annotation, and samples that are correctly classified but not accurate enough in position are set as difficult negative samples; samples with an IOU less than the threshold and misclassified are set as difficult positive samples.
[0088] Select the exponential decay function for hard sample weighting. Here, w is the weight of the current sample, e is the base of the natural logarithm, and the decay rate control parameter c is determined by experimental cross-validation:
[0089] w = e (-c)×IOU
[0090] Use the weighted samples (including hard samples and normal samples) for the next round of model training, and perform multiple iterative trainings to further optimize the performance of the model.
[0091] Step 2. Model training: Fine-tune the network used to improve the network's adaptability to the specific environment of the factory. Specifically,
[0092] Use the dataset obtained in Step 1 to train the PULC ultra-lightweight image classification algorithm to obtain the human / non-human classification network used in this embodiment.
[0093] Use the dataset obtained in Step 1 to train the YOLOv8 model to obtain the person detection network used in this embodiment.
[0094] Use the dataset obtained in Step 1 to train the PP-ShiTu model to obtain the feature extraction network used in this embodiment.
[0095] Step 3. Model inference: The actual application of the model in the industrial scenario. It is mainly divided into three modules, the human / non-human classification module (such as Figure 4 ), the person detection module (such as Figure 5 ), and the person clothing recognition module (such as Figure 6 ). Specifically,
[0096] 1) Human / non-human classification module: Detect whether there are people in the video stream in the camera and determine whether to perform subsequent operations.
[0097] The specific process is as follows:
[0098] First, obtain the RTSP video stream of the network camera and extract video frames from it as the input image of the algorithm.
[0099] Preprocess the image, convert it to the defined size of the input image by the PULC human / non-human classification model structure, adapt the channel dimension and data type, and finally perform pixel normalization processing.
[0100] In this embodiment, the preprocessing process is completed using PyTorch.
[0101] Use the trained human / non-human classification network for inference to obtain the prediction result.
[0102] The predicted result is an array that stores two probabilities between 0 and 1, representing the probabilities of no person and a person in the image respectively.
[0103] For video frames with a person, they are passed to the subsequent program for pedestrian detection; otherwise, the detection process ends here.
[0104] 2) Personnel detection module: Achieve accurate detection of various intruding personnel in complex industrial scenarios.
[0105] Process the video frames determined to have a person in the person / no-person classification module. The specific process is as follows:
[0106] First, adjust the video frame: Resize it to the size (640, 640) that meets the input image requirements of YOLOv8, adapt to the channel dimension and data type, and finally perform pixel normalization.
[0107] The present invention improves the feature pyramid network of YOLOv8 as follows:
[0108] On the one hand, optimize it from unidirectional top-down fusion to bidirectional feature flow as the feature fusion strategy. Ensure efficient information flow between different scales: High-level feature information can be passed down to improve the model's detection ability for smaller humans; low-level feature information can be passed up to enhance the model's perception ability for larger humans.
[0109]
[0110] On the other hand, perform weighted feature fusion to obtain the output features of each layer. Through learnable weights, let the model dynamically adjust the contribution of each scale feature by learning the weights, while avoiding the information imbalance problem caused by the original simple addition in YOLOv8. When learning the weights, first take the exponential of the learnable weights and then perform Softmax normalization to ensure that the sum of the contribution degrees of all features is 1. Compared with directly using the weights, taking the exponential before normalization can ensure that all weights are positive numbers and avoid the negative weight problem.
[0111]
[0112] Among them, represent the feature outputs after top-down and bottom-up fusion respectively, F i top 、F i bottom represent the information from the high-resolution layer and the low-resolution layer respectively, represent the learnable high-level weights and low-level weights respectively, and ε is a minimum constant.
[0113] When learning the weights, for the learnable parameters First, take the exponent to avoid the negative weight problem, and then perform Softmax normalization to ensure that the sum of the contributions of all features is 1.
[0114]
[0115] Furthermore, prune the model to remove the parts of the neural network that contribute less to the prediction. For the absolute value size of each weight, set a threshold and prune; for the L1 norm of each channel, set a threshold and delete the channels with less importance. After pruning, fine-tune the model to restore the accuracy. Specifically as follows:
[0116] First, prune the weights, that is, remove the weights close to 0. Calculate the absolute value size of each weight S(w)=|w|, set a threshold T, if S(w)<T, prune and set w = 0.
[0117] Then, prune the channels, that is, remove the channels with less total contribution. Use the L1 norm to calculate the total contribution of the weights in the three dimensions of the c-th channel, set a pruning threshold T, and filter the channels below this threshold.
[0118]
[0119] Among them, C is the total number of channels in the network layer. For the selected channels, delete their corresponding convolutional kernels.
[0120] W pruned =W\{w c}
[0121] Among them, W is the weight tensor of the original convolutional layer, w c represents the subset of weights corresponding to the c-th channel selected to be pruned, and W pruned represents the pruned weight matrix.
[0122] Use the trained person detection network for inference to obtain the inference result.
[0123] Perform shape conversion on the prediction result, converting from (1, 84, 8400) to (8400, 85). Among them, in (1, 84, 8400), 1 represents the batch size, 84 represents the information of each prediction box, including 4 bounding box coordinates and the prediction probabilities under 80 category information, and 8400 represents the total number of prediction boxes; in (8400, 85), 8400 represents the total number of prediction boxes, and 85 represents the information of each prediction box, including the bounding box coordinates (4 values), confidence (1 value), and category information (80 values).
[0124] Convert the format of the bounding box. In YOLOv8, the bounding box coordinate representation format is the center point coordinate and width and height (x, y, w, h), and convert it to the representation form of the upper left corner and the lower right corner (x1, y1, x2, y2).
[0125] Perform non-maximum suppression to filter out overlapping bounding boxes in the model output and restore the coordinates of the inference results mapped back to the original image. Finally, draw the prediction boxes and class labels on the original image.
[0126] 3) Personnel clothing recognition module: Collect template images to make a corresponding dataset of tooling feature vectors. For the detection area of tooling, use a feature extraction and matching network based on metric learning to implement corresponding pedestrian attribute analysis and give corresponding warnings according to the actual situation.
[0127] Specifically,
[0128] First, collect common tooling used in the actual production process of the factory, safety protection clothing used in special scenarios, and various possible daily outfits. Remove images that do not meet the requirements or are redundant through screening, and use annotation tools to accurately frame and label the attributes of the clothing part. Use a feature extraction model with the PaddlePaddle structure to extract features from the obtained images and construct a dataset of tooling feature vectors suitable for perimeter safety production and personnel attribute analysis.
[0129] Perform K-Means clustering on the vectors in the tooling feature vector dataset and divide all vectors into 8 clusters.
[0130] C = {c1, c2,..., c8}, c i ∈R 512
[0131] Then assign each database vector to the nearest cluster center, that is, store each vector value in the nearest cluster. And construct an inverted file, where each cluster stores the vectors belonging to it.
[0132]
[0133] Index(c j ) = {v j1 , v j2 ..., v jm}
[0134] Where C represents the set of all cluster centers, c j represents the j-th cluster center, Assign(v i ) represents the index of the nearest cluster center to which the vector belongs, Index(c j ) represents the inverted list corresponding to the cluster center, and {v j1 , v j2 ..., v jm} represents the set of vectors stored in the inverted list.
[0135] For the results of personnel detection, extract the detection boxes with the detection category of personnel. According to the defined standing positions of the human body and the possible standing postures in the actual situation, after multiple adjustments and tests, it is stipulated to crop the upper 109 / 237 rectangular part of the personnel detection box as the tooling detection area. This part includes the head and upper body of the personnel, which reduces the dependence on the lower clothing on the one hand and ensures that the feature extraction network obtains sufficient context information to improve the effect on the other hand.
[0136]
[0137] For the tooling detection area, first use PyTorch to adjust it to the defined size (224, 224) of the input image for the feature extraction network, and convert the channel dimension and data type into the form required by the image. Finally, perform pixel normalization processing.
[0138] Load the feature extraction model to obtain the feature vector. And perform normalization processing on the obtained 512-dimensional vector: calculate the Euclidean norm of the feature vector, and divide each element in the feature vector by the norm value to achieve the normalization operation.
[0139] Retrieve the 5 nearest clusters, and perform a nearest neighbor search among all the vectors belonging to the selected cluster set C selected to select the final matching result.
[0140]
[0141] Among them, v query represents the input vector that needs to be matched currently, and c j represents the j-th clustering center, and d(v query , v i ) represents the distance between the query vector and the candidate vector.
[0142] Perform condition checking, requiring two conditions to be met simultaneously: judge that the work clothing of the matched personnel is the tooling category that must be worn to enter the perimeter area through the vector with the maximum cosine similarity; and when it exceeds the preset threshold for multiple consecutive frames, determine that the personnel entering the perimeter belong to normal work. Otherwise, it is an abnormal situation.
[0143] Step 4. Deep learning inference acceleration and model deployment: Use Tensorrt technology to quantize and compress each model to improve the inference speed of the model on edge devices while ensuring accuracy. Select edge devices according to the computing resource constraints for model deployment and system integration. (Such as Figure 7 )
[0144] Specifically, for all the above models:
[0145] First, perform inference acceleration. Convert Pytorch and PaddlePaddle into onnx models, and then perform TensorRT inference optimization on the onnx format models.
[0146] Then, select a suitable edge device for model deployment. In this example, NVIDIA Jetson Orin Nano 8GB is selected for software and hardware deployment in the environment of Linux aarch64 / Ubuntu 20.04.
[0147] For the cameras at the production site, obtain video information from the site, then transmit the information to the communication base station through a switch, and then upload it to the server by the communication base station. Finally, obtain the RTSP data stream that can be processed at the edge device terminal.
[0148] Build a simulation environment and compare the accuracy of the method of the present invention and the benchmark method in personnel detection and personnel attribute analysis. The results are shown in the following table.
[0149]
[0150] Comparison of personnel detection accuracy: The benchmark method uses the unimproved and unoptimized YOLOv8 network to test on the constructed factory business dataset. The present invention conducts data fusion and dataset adjustment for the perimeter intrusion and pedestrian attribute analysis system. And train the image classification network based on the PP-LCNet network and the pedestrian detection network based on CSPNet, and the accuracy of personnel detection is improved by 9.08% on the constructed factory business dataset.
[0151] Comparison of personnel attribute analysis accuracy: The benchmark method uses the unimproved and unoptimized PP-ShiTu network to test on the constructed factory business dataset. The present invention appropriately crops the results of pedestrian detection, uses the feature extraction network based on metric learning to obtain features and then performs retrieval and matching, and the accuracy of personnel attribute analysis is improved by 7.42% on the business dataset.
[0152] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A perimeter intrusion detection and personnel attribute analysis method for work safety, characterized in that It includes the following steps: Step 1. Data fusion: Perform data fusion on the perimeter intrusion and pedestrian attribute analysis system, construct a training set for industrial safety production, and provide it to Step 2; Step 2. Construct a person presence / absence classification network, a person detection network, and a feature extraction network, and use the dataset in Step 1 to train and fine-tune them to improve the network's adaptability to the specific factory environment; Step 3. Model inference: Deploy the above network model to the industrial scenario, including three modules: a person presence / absence classification module, a person detection module, and a person clothing recognition module; among them, the person presence / absence classification module deploys the person presence / absence classification network, the person detection module deploys the person detection network, and the person clothing recognition module deploys the feature extraction and matching network. Use the corresponding work clothing feature vector dataset made according to the actual scenario to achieve the corresponding pedestrian attribute analysis, and give corresponding warnings according to the actual situation; The video frames determined to have people by the person presence / absence classification module are provided to the person detection module for pedestrian detection, and the person clothing recognition module determines that the matched person's clothing is the work clothing category that must be worn to enter the perimeter area. If it exceeds the preset threshold for multiple consecutive frames, it is determined that the person entering the perimeter belongs to normal work; Step 4. Deep learning inference acceleration and model deployment: Use Tensorrt technology to quantize and compress each model to improve the inference speed of the model on edge devices while ensuring accuracy; select edge devices according to the computing resource constraints for model deployment and system integration.
2. The method for perimeter intrusion detection and personnel attribute analysis for work safety according to claim 1, wherein In Step 1, data fusion includes two steps: data augmentation and hard sample screening; For the data augmentation part, for the factory-related data collected, use a rectangular box to select the position of people as the annotation rule. Roughly mark the individuals and approximate position information in the image through an automated tool and perform manual refinement and correction for precise positioning. Establish an xml annotation file to hierarchically store all attributes of the annotation objects; Then, perform random direction transformation, size transformation, image sharpening or blurring processing, and brightness adjustment operations on the factory images in sequence to simulate different shooting conditions and changes; Finally, integrate with the public MS COCO dataset and the CrowdHuman dataset for crowd detection to construct a dataset with industrial personnel as the information main body and including various factory environmental backgrounds; For the hard sample mining part, first use the initial dataset for model training, and use the trained model for prediction to obtain the object detection results, including the predicted bounding boxes and categories; Then divide the hard samples according to the IOU value, and assign higher weights according to their IOU values so as to pay more attention to these hard examples in the next round of training; Next, use the weighted samples, including hard samples and ordinary samples, for the next round of model training, and perform multiple iterative trainings to further optimize the performance of the model.
3. The perimeter intrusion detection and personnel attribute analysis method for work safety according to claim 1, wherein In Step 2, the person presence / absence classification network uses the PULC ultra-lightweight image classification model; the person detection network uses the improved YOLOv8 model; the feature extraction network uses the PP-ShiTu model.
4. The perimeter intrusion detection and personnel attribute analysis method for work safety according to claim 1, characterized in that In step 3, object detection is achieved by improving the feature pyramid network of YOLOv8, specifically as follows: On the one hand, it is optimized from unidirectional top-down fusion to bidirectional feature flow as the feature fusion strategy, ensuring efficient information flow between different scales: high-level feature information can be passed down to improve the model's detection ability for smaller humans; low-level feature information can be passed up to enhance the model's perception ability for larger humans. On the other hand, weighted feature fusion is performed to obtain the output features of each layer. Through learnable weights, the model can dynamically adjust the contribution of each scale feature by learning the weights, while avoiding the information imbalance problem caused by the original simple addition in YOLOv8. When learning the weights, the learnable weights are first exponentiated and then Softmax-normalized to ensure that the sum of the contribution degrees of all features is 1. Compared with directly using weights, exponentiating before normalization can ensure that all weights are positive numbers, avoiding the problem of negative weights. Furthermore, the model is pruned to remove parts with less contribution to the prediction in the neural network. For the absolute value size of each weight, a threshold is set for pruning; for the L1 norm of each channel, a threshold is set to delete channels with less importance. After pruning, the model is fine-tuned to restore the accuracy. Use PyTorch to preprocess the image: adjust it to the specified size of the input image for YOLOv8, convert the channel dimension and data type to the form required by the image, and finally perform pixel normalization. Use the improved YOLOv8 for inference to obtain the inference result. Perform post-processing to process the format of the model output and the format of the bounding boxes, and then perform non-maximum suppression to filter out overlapping bounding boxes in the model output. Finally, restore the inference result and draw both the prediction boxes and class labels on the image.
5. The perimeter intrusion detection and personnel attribute analysis method for work safety according to claim 1, characterized in that In step 3, the personnel clothing recognition module creates a corresponding tooling feature vector dataset according to the actual scenario, uses a feature extraction and matching network based on metric learning to achieve corresponding pedestrian attribute analysis, and issues corresponding warnings according to the actual situation. Specifically, the feature extraction and matching network is obtained through the following steps: In step 3.1, for the clothing that factory personnel may wear, template images are collected to create a feature vector dataset for matching. In step 3.2, for the results of pedestrian detection, the feature extraction network obtained in step 2 is used for feature extraction. 1) Target image cropping: According to the person detection box obtained by the person detection module, crop the tooling detection area. 2) Feature acquisition and storage: Feed the cropped image into the feature extraction model for inference. In step 3.3, match the extracted features with the feature vector dataset. In step 3.1, since in practical applications, videos mainly come from real-time monitoring in the factory, in order to make the feature extraction framework more adaptable to the recognition task of pedestrian attribute targets for perimeter safety production, common work clothes used in the factory during actual production, safety protection clothing used in special scenarios, and various possible daily outfits are collected. A new feature vector dataset applicable to perimeter safety production and personnel attribute analysis is constructed through methods such as screening and annotation; In step 3.2, the feature extraction part needs to complete two aspects of content in sequence: target image cropping, feature acquisition and storage; First, traverse all the bounding boxes containing prediction results obtained from pedestrian detection in step 2, extract the detection boxes with the detection category of personnel, and crop them from the original video frames; Then, only intercept the head and upper body parts of the human detection box as the initial input of the feature extraction network; Specifically, first use PyTorch to preprocess the image: convert it to the defined size of the input image by the feature extraction network, convert the channel dimension and data type into the form required by the picture, and perform pixel normalization processing; Load the feature extraction model, make predictions on the input image, and obtain the prediction results; Perform normalization processing on the obtained 512-dimensional vector; Calculate the Euclidean norm of the feature results, and divide each element in the feature vector array by the norm value to achieve the normalization operation; In step 3.3, the feature matching part needs to match the normalized result with the feature vector dataset in step 3.1; First, use K-means clustering to divide all vectors into k clusters and obtain the clustering centers of each cluster; Then each database vector v i is assigned to the nearest cluster center, that is, each vector is stored only in the nearest cluster; an inverted file structure is constructed, and each cluster center c j maintains a list storing all vectors belonging to that cluster; For the target vector, retrieve the 5 closest clusters and find the nearest neighbor in their vector lists as the matching result; The condition check requires that two conditions be met simultaneously: matching a specific work clothing (the work clothing category that must be worn when entering the perimeter area), and when it exceeds a certain threshold for multiple consecutive frames, it is determined that the person entering the perimeter belongs to normal work.