Personnel state monitoring method and system based on YOLO algorithm
By optimizing the YOLO algorithm combined with the CBAM attention mechanism and multi-scale feature fusion technology, the problem of inefficient safety monitoring of personnel in natural gas stations is solved, high-precision real-time monitoring and abnormal behavior recognition are achieved, and the level of safety management is improved.
Patent Information
- Application Number
- CN202510170147.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art personnel safety monitoring in natural gas stations is inefficient, making it difficult to achieve high-precision all-weather monitoring, especially in complex environments, the accuracy of personnel posture recognition is insufficient.
The optimized YOLO algorithm is adopted, combined with the CBAM attention mechanism and multi-scale feature fusion technology to detect and pose estimate the human body. The key points are corrected through historical frame data, an action model is established and compared with the normal behavior library to generate an alarm signal.
Real-time and accurate monitoring of personnel status in the natural gas station has been achieved, the efficiency and safety management capabilities of abnormal behavior recognition have been improved, and it can operate stably in complex contexts.
Smart Images

Figure CN120339932A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image or video recognition, and particularly to a method and system for monitoring the status of personnel based on the YOLO algorithm. Background Art
[0002] Natural gas is one of the most important energy sources globally, featuring low carbon, green, efficient, and clean characteristics, and is widely used in multiple fields such as industry, households, and transportation. Natural gas stations, as the core nodes for the transportation, storage, and distribution of natural gas, are crucial links in the natural gas supply chain. The safety of natural gas stations is of vital importance to national energy security and the safety of people's lives and property. However, due to their complex equipment, changing environments, and high potential for sudden risks, natural gas stations often face significant challenges in safety monitoring. At natural gas stations, staff often need to operate high-pressure gases and equipment. Once a safety accident occurs, it may lead to serious casualties and property losses. Therefore, how to conduct all-weather and high-precision personnel safety monitoring through intelligent means has become an urgent need to ensure the safe operation of natural gas stations. Currently, most natural gas stations still use traditional manual inspections or video analysis based on static surveillance cameras for personnel safety monitoring. However, these traditional methods have significant drawbacks. Manual supervision is inefficient, consumes a large amount of resources, and it is difficult to conduct efficient centralized management and analysis of monitoring personnel.
[0003] In Chinese Patent Publication No.: CN112487915B, a pedestrian detection method based on the Embedded YOLO algorithm is disclosed, including the following steps: (1) Extract all pedestrian image data in the dataset, and randomly divide the extracted image data into a training set and a test set; (2) Construct an Embedded module based on a deep convolutional network; (3) Stack the Embedded modules and combine them with MobileNet, SPP, and YOLO layers to form the entire Embedded YOLO detection network model; (4) Use the training set to train the neural network of the Embedded YOLO model to obtain the optimal detection network model; (5) Detect the picture data in the test set, and evaluate the detection accuracy, speed, and lightweight of the detection results of the test set. This method has a relatively low accuracy in recognizing human postures when dealing with complex scenarios. Summary of the Invention
[0004] The present invention discloses a method and system for monitoring the status of personnel based on the YOLO algorithm, aiming to achieve automatic and real-time monitoring of the status of personnel in natural gas stations, improve the accuracy and efficiency of abnormal behavior recognition, and thus enhance the safety management ability of the station.
[0005] The technical solutions adopted by the present invention to solve the above technical problems are as follows: A method for monitoring the state of personnel based on the YOLO algorithm, comprising: Obtaining real-time image or video stream data from the cameras of the natural gas station; Determining the range of the human body in the image or video stream; Using the YOLO algorithm to obtain human key point information and performing pose estimation on the human body; Establishing an action model according to the result of the pose estimation; Judging the state of the personnel according to the action model.
[0006] Preferably, for the pose estimation of the human body, an optimized YOLO algorithm is adopted, and the accuracy of key point detection is improved by introducing the CBAM attention mechanism.
[0007] Preferably, the optimized YOLO algorithm improves the detection ability under complex backgrounds through multi-scale feature fusion technology.
[0008] Preferably, the human key point information is obtained through a predefined human skeleton model and at least includes the key joint point information of the head, facial features, torso and limbs.
[0009] Preferably, the method can also evaluate the result of the human pose estimation through a human pose estimation loss function, correct the human key points, and optimize the action model.
[0010] Preferably, the correction of the human key points is based on historical frame data and a prediction model, and the key point positions are smoothed to eliminate noise.
[0011] Preferably, the specific method for judging the state of the personnel is as follows: Comparing the action model of the current personnel with a pre-stored normal behavior sample library; Calculating the similarity index of key features and judging whether it conforms to the normal behavior pattern; Setting the thresholds for pose deviation, action speed and the degree of abnormality of key point positions; If the detected deviation of behavior features exceeds the threshold, it is considered that there is an abnormality.
[0012] Preferably, when it is judged that the state of the personnel is abnormal, an alarm signal is generated and the management terminal is notified.
[0013] Preferably, the alarm signal includes real-time alarm information sent to the management terminal and triggers an alarm device in case of major abnormalities.
[0014] A personnel state monitoring system based on the YOLO algorithm, which executes the personnel state monitoring method based on the YOLO algorithm according to any one of claims 1 to 9, comprising: A camera, which acquires images or video streams of a natural gas station yard; A data processing module, which performs recognition processing on the images or video streams; An alarm module, which generates and transmits alarm signals; A management terminal, which receives the alarm signals generated by the alarm module and performs monitoring and management operations.
[0015] Through the optimized YOLO-Pose algorithm, the present invention realizes real-time and accurate monitoring of the personnel status in a natural gas station yard, and has efficient human key point detection and abnormal behavior recognition capabilities. By introducing the CBAM attention mechanism and multi-scale feature fusion technology, the detection accuracy and robustness in complex environments are improved. The system can timely judge the abnormal status of personnel and generate alarm signals, effectively reducing safety risks and improving the safety management efficiency of the station yard. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a flow chart for monitoring the personnel status in a natural gas station yard.
[0018] Figure 2 It is a schematic diagram of human key points.
[0019] Explanation of the reference numerals in the drawings: 0 nose; 1 left eye; 2 right eye; 3 left ear; 4 right ear; 5 left shoulder; 6 right shoulder; 7 left elbow; 8 right elbow; 9 left wrist; 10 right wrist; 11 left hip; 12 right hip; 13 left knee; 14 right knee; 15 left ankle; 16 right ankle. Detailed Embodiments
[0020] Term Explanation: Anchor: Anchor point, used to locate the key points of the human body; AP (Average Precision): It refers to the area under the PR curve. The larger the value, the more accurate it is and the higher the detection accuracy; mAP: Taking the average of the AP values of all classes is mAP; IoU (Intersection Over Union): The intersection in the union. The larger the IoU value, the more consistent the predicted box and the ground truth box are. When they are exactly the same, IoU = 1; OKS (Object Keypoint Similarity): When performing human pose estimation, it is used to evaluate the similarity degree of key points of the human pose. The larger the value, the higher the similarity degree.
[0021] Embodiment This embodiment is deployed in the safety monitoring system of a natural gas station. The hardware environment includes multiple monitoring cameras and a GPU server (configured with 4 RTX3060 graphics cards and 64GB of memory). The software environment is the Ubuntu 18.04 operating system, supporting a deep learning framework for the YOLO-Pose algorithm. The system is used to monitor the pose changes and behavior trajectories of personnel in the station in real time.
[0022] The monitoring cameras in the natural gas station acquire image and video stream data in real time. The image resolution is 1920×1080, and the sampling frame rate is 30 frames per second. The collected data is transmitted through the network to the GPU server for processing.
[0023] Human detection: The image data is processed through the YOLOv5m_CBAM model. This model introduces the CBAM attention mechanism, which can improve the accuracy and robustness of human detection in complex backgrounds. The model identifies the bounding box of each human body in the image and outputs the relevant confidence level.
[0024] Human keypoint localization: The YOLO-Pose algorithm is used to estimate the keypoints of each detected human body, and the positions and confidence levels of 17 keypoints including the head, shoulders, wrists, knees, and ankles are identified. Keypoints with a confidence level lower than 0.5 are filtered to improve the reliability of pose estimation.
[0025] Pose estimation result evaluation: The IoU bounding box loss function and the OKS (Object Keypoint Similarity) loss function are used to evaluate and optimize the pose estimation results. The OKS loss function assigns weights to different keypoints to ensure that the algorithm estimates important keypoints more accurately.
[0026] Establishment of the motion trajectory model: According to the changes in keypoints in consecutive frames, a time series analysis method is used to establish a personnel motion trajectory model. This model can record the movement trajectories and pose changes of personnel in real time, providing support for subsequent behavior analysis.
[0027] Abnormal Behavior Detection: Based on the postural changes and movement trajectories of personnel, the system can automatically identify abnormal behaviors. The action model of the current personnel is compared with the pre-stored normal behavior sample library: calculate the similarity index of key features (such as key point positions, speeds, angle changes, etc.), and determine whether it conforms to the normal behavior pattern; set the abnormal determination threshold, such as the abnormal degree of posture deviation, action speed or key point position; if the detected behavior feature deviation exceeds the threshold, it is considered abnormal. For example, when it is monitored that a person maintains a falling posture for a long time or enters a prohibited area, the system will trigger an alarm and send a notification to the station yard safety management personnel.
[0028] Scenario 1: Normal Operation Monitoring: During the daily operation of the station yard, the system captures the working status of personnel in real time through surveillance cameras. The YOLO-Pose algorithm can accurately identify common postures such as standing, walking, bending, squatting, etc., ensuring that all staff are in a safe state.
[0029] Scenario 2: Emergency Event Detection: When a staff member accidentally falls and does not return to the normal posture for a long time, the system detects the abnormal behavior, immediately records the event, and issues an alarm to the station yard management center. At the same time, the algorithm analyzes the movement trajectory before and after the fall to provide data support for accident cause analysis.
[0030] Scenario 3: Dense Crowd Scenario Monitoring: In the case of station yard equipment maintenance or concentrated personnel operations, the system can accurately distinguish the key points of each person in the dense crowd, avoiding pose estimation errors caused by occlusion or key point confusion.
[0031] 5. Effect Verification: To verify the system performance, in this embodiment, the YOLO-Pose algorithm was offline trained and tested using the COCO key point dataset. The results on the test dataset are as follows: mAP(%), AP50(%), AP75(%), AP90(%) are respectively: 84.9, 85.4, 75.8, 71.0. In actual deployment, the system can operate stably in complex backgrounds, low light and dense personnel environments, meeting the requirements of natural gas station yards for personnel pose estimation and abnormal behavior detection.
[0032] Next, the present invention will be described in detail with reference to the accompanying drawings.
[0033] As Figure 1 shown, the present invention includes steps such as data input, human body detection, pose estimation, establishing a movement trajectory model, model evaluation, optimization adjustment, abnormal judgment, and issuing an alarm.
[0034] Next, the human body detection step will be introduced in detail.
[0035] Human detection is a crucial step in pose estimation algorithms, aiming to accurately identify and locate the position of the human body in the input image or video stream. In this process, to address the problem that the YOLOv5 algorithm is insensitive to target personnel in natural gas stations, human detection combines the YOLOv5m_CBAM model to improve the stability and reliability of human detection.
[0036] YOLOv5 focuses on the object detection task of 80 categories defined in the COCO dataset. For each detected Anchor, a vector containing 85 elements is predicted through the Box head part. These 85 elements not only cover the probabilities of 80 categories but also include the bounding box position of the target, the score of the target's existence, and the confidence score. In addition, 4 different shapes of Anchors are set at each grid position to adapt to targets of different sizes and shapes.
[0037] YOLOv5m_CBAM is an object detection model that combines a convolutional neural network (CNN) and an attention mechanism. It inherits the efficiency and real-time performance of the YOLO series and, at the same time, introduces the CBAM (Convolutional Block Attention Module) attention module to further improve the detection accuracy of the model. Through a large amount of data training, the YOLOv5m_CBAM model can learn the characteristics of the human body and thus accurately identify the bounding box of the human body in a complex background.
[0038] In the safety monitoring of natural gas stations, due to the complex and changeable environment, human detection faces many challenges. However, the excellent performance of the YOLOv5m_CBAM model enables it to operate stably in this environment and output the position information of the human body in real time. This allows us to track and monitor the personnel in the natural gas station in real time, and timely discover and handle potential safety risks.
[0039] Next, the steps of human key point localization will be introduced in detail.
[0040] In this task, 17 key points of the human body need to be identified and located, as Figure 2 shown. The position and confidence of each key point need to be identified, which means that for each Anchor, the key point detection head needs to predict 51 elements to describe these key points.
[0041] At the same time, to completely describe a detected human body instance, the Box head also needs to predict 6 additional elements, covering information such as the bounding box and confidence of the human body. Therefore, for each Anchor containing n key points, the overall prediction vector will be defined by these components, ensuring that the model can accurately detect and estimate the pose of the human body. Among them, the overall prediction vector is defined as:
[0042] When performing human keypoint detection, accurate prediction of keypoint confidence is crucial as it directly relates to the reliability of subsequent processing and analysis. Confidence reflects the model's level of confidence in the accuracy of the keypoint positions. During training, this confidence is determined based on the visibility flags of the keypoints.
[0043] Specifically, if a keypoint is visible in the image or can be inferred even when partially occluded, the ground truth confidence of that keypoint is marked as 1, indicating high confidence. Conversely, if a keypoint is outside the field of view of the image, i.e., completely invisible, its confidence is set to zero, indicating that there is no reliable keypoint at that position. During the inference phase of the model, maintaining a keypoint confidence higher than 0.5 becomes an important criterion.
[0044] This means that only those keypoints predicted by the model with high confidence are considered valid and used for subsequent processing and analysis, while all keypoint predictions with confidence below this threshold are regarded as unreliable and accordingly ignored. This step is necessary as it helps filter out inaccurate keypoints that may lead to incorrect pose estimation, especially important when dealing with complex backgrounds or partial occlusions. During model evaluation, the predicted keypoint confidence is not directly used. However, the prediction of confidence is still an indispensable part of the entire keypoint detection framework as it enables the model to distinguish which keypoints in the image are credible during prediction. This contrasts with existing bottom-up methods based on heatmaps, which usually do not need to filter out keypoints outside the field of view as these keypoints are not included in the detection results from the beginning. For a given image, the anchor points matched to a person store their entire 2D pose as well as the bounding boxes. The box coordinates are transformed according to the anchor point center, and the box size is normalized according to the height and width of the anchor point. Similarly, the keypoint positions are also transformed to the anchor point center. However, the keypoints are not normalized using the anchor point height and width. Both the keypoints and the boxes are predicted to the center of the anchor point. Since the improvement of YOLO-Pose is independent of the width and height of the anchor, YOLO-Pose can be easily extended to anchor-free object detection methods.
[0045] Next, a detailed introduction to the human pose estimation loss function will be given.
[0046] (1) IoU-based bounding box loss function Most object detectors optimize advanced variants of the IoU loss, such as GIoU, DIoU, or CIoU loss, rather than distance-based losses for box detection, because these losses are scale-invariant and directly optimize the evaluation metric itself. YOLO-Pose uses CIoU Loss for bounding box supervision. For the ground truth bounding box matched by the k-th anchor at position (i, j) and scale s, the loss is defined as: where, is the predicted box of the k-th anchor at position (i, j) and scale s. In YOLO-Pose, there are three anchors at each position, and predictions are made at four scales.
[0047] (2) Human Pose Loss Function Formula In the field of keypoint detection, Object Keypoint Similarity (OKS) is a metric commonly used to evaluate model performance. Traditional bottom-up methods usually rely on heatmaps to detect keypoints and use the L1 loss function for optimization. However, the L1 loss does not always lead to the best OKS performance because it does not take into account the scale of the object or the type of keypoint. Since a heatmap is a probability distribution map, it is not possible to directly use OKS as a loss function in a heatmap-only method. OKS can only be used as a loss function when regressing the keypoint positions. In their research, Geng et al. adopted a scale-normalized L1 loss for keypoint regression. A key advantage of this method is that it allows the algorithm to directly regress keypoints towards the anchor center, thus optimizing the evaluation metric itself rather than a surrogate loss function. This idea is similar to the concept of the Intersection over Union (IOU) loss, but extends it from bounding boxes to keypoints. In the case of keypoints, Object Keypoint Similarity (OKS) can be regarded as the equivalent of IOU. The OKS loss is essentially scale-invariant, which means it can maintain consistent performance for objects of different scales. [1] In addition, the OKS loss can also distinguish different keypoints and assign higher weights to some keypoints, which reflects their relative importance in the overall performance evaluation. By introducing the scale-normalized L1 loss and directly optimizing OKS, this method provides a more refined and effective optimization strategy for keypoint detection, enabling the model to better handle different types of keypoints while maintaining scale invariance.
[0048] If the ground truth bounding box matches the anchor at position (i, j) and scale s, the keypoints will be predicted relative to the center of the anchor. The OKS is calculated separately for each keypoint and then summed to give the final OKS loss or keypoint IOU loss.
[0049]
[0050] where d n is the Euclidean distance between the predicted position and the ground truth position of the n-th keypoint; k n is the specific weight of the keypoint; s is the scale of the object; δ(v n ) is the visibility flag for each keypoint. Corresponding to each keypoint, a confidence parameter is learned, which indicates whether the person has the keypoint.
[0051] where is the predicted confidence of the n-th keypoint. Generally, such a loss function will include classification loss, bounding box regression loss, and keypoint-related loss. If the ground truth bounding box matches the anchor, the loss at position (i, j) is valid for the anchor of scale s.
[0052] Finally, the total losses for all scales, anchors, and positions are summed: where the hyperparameters: λ cls = 0.5, λ box = 0.05, λ kpts = 0.1 and λ kpts_conf = 0.5, and these hyperparameters are mainly used to balance losses of different scales.
[0053] (1) represents the classification loss, and this part of the loss function is responsible for optimizing the model's ability to recognize object categories. In the YOLO series of algorithms, this is usually achieved through Binary Cross-Entropy Loss, ensuring that the model can accurately classify different objects; (2) represents the bounding box regression loss, and this part of the loss function is used to optimize the difference between the predicted bounding box and the true bounding box by the model. In the YOLO algorithm, this usually involves calculating the distance from the center point, such as metrics like 1-CIoU (Complete Intersection over Union), to reduce the localization error; (3) Represents the key point position loss, which is the part specifically for pose estimation in YOLO-Pose. To more accurately predict the key point positions, YOLO-Pose adopts a loss function based on the L2 distance, which helps the model learn the exact positions of the key points; (4) Represents the key point confidence loss. YOLO-Pose optimizes the prediction accuracy of the model on the presence or absence of key points through binary cross-entropy loss.
[0054] Since YOLO-Pose is an algorithm that integrates human detection and pose estimation, its goal is to optimize the performance of the model through the combination of these losses, enabling the model to accurately perform human pose estimation while maintaining a relatively fast speed. YOLO-Pose also takes into account the scale-normalized L1 loss, which is a step towards the OKS loss. Since the heatmap is a probability map, it is impossible to directly use OKS as the loss function for pure heatmap methods. Only when regressing the key point positions can OKS be used as the loss function. In this way, YOLO-Pose can better handle key points of different scales and types. Generally speaking, YOLO-Pose realizes the comprehensive optimization of the performance of the algorithm model through the combination of these loss functions, enabling it to accurately perform human pose estimation while maintaining a relatively fast speed.
[0055] The YOLO-Pose algorithm can achieve an end-to-end training mode, which is different from traditional two-stage pose estimation methods based on Heatmap. In traditional methods, object detection and pose estimation are carried out separately and often rely on an indirect L1 loss, which not only increases the training complexity but also affects the direct optimization of the key point similarity (OKS) of the model. YOLO-Pose can jointly perform bounding box detection and pose estimation through a single forward pass, greatly improving the training efficiency and model performance. Different from methods that rely on alternative loss functions, YOLO-Pose optimizes the OKS metric itself, which means that it directly focuses on improving the accuracy of pose estimation during model training. In this way, YOLO-Pose can more accurately identify and evaluate human poses, making it superior to traditional methods in terms of accuracy and efficiency.
[0056] The specific experimental configuration is as follows: The algorithm model is deployed on an Ubuntu 18.04 server for training. The basic configuration is 4 RTX3060 GPU graphics cards and 64G of memory. During the training process, 300 iteration cycles are set, the initial learning rate is 0.01, and the learning rate is processed using the step decay method. When the iteration cycle reaches 100, the learning rate drops to 1 / 10 of the current value; when the iteration cycle reaches 200, the learning rate drops to 1 / 10 of the current value again. The backpropagation process uses the gradient descent method, where the momentum parameter is set to 0.9 and the batchsize is set to 128.
[0057] Next, a detailed introduction to the ablation experiment of human key points will be given.
[0058] To evaluate the impact of different components or features in the model on the overall performance, a human pose estimation ablation experiment is introduced. By systematically removing (or "ablating") certain parts of the model and observing how this change affects the performance of the model.
[0059] The proposed pose estimation model is experimented with the bottom-up HigherHRNet-w32, HRNet, LightweightOpenPose, OpenPose, and DEKR-W32 models on the COCO keypoint dataset. The experiments are all carried out in the same experimental environment and configuration. The algorithm model evaluates the model on the COCO dataset, which consists of more than 200,000 images, 250,000 person instances, and 17 key points. It includes three sets: train2017, val2017, and test-dev2017. The train2017 set contains 57K images, while the val2017 and test-dev2017 sets contain 5K and 20K images respectively. The model is trained on the train2017 set and the results are tested on the val2017 set and the test-dev2017 set. The results are shown in Table 1: Table 1
[0060] Because there is a distance error between the predicted value of the human key point detection algorithm and the true key point, and the size of the human image and the offset of the manually marked position will also affect the prediction result, the key point detection evaluation index OKS is used to calculate the similarity between the predicted value and the true key point. Based on the characteristics of key point detection, using P OKS The metric average precision AP that measures the entire algorithm is calculated as follows: given a constant t, if the current P OKS >t, it indicates that the key point is correctly detected; if P KS <t, it indicates that it is not successfully detected or there are missed detections or false detections. Therefore, for all P OKSCount the number greater than t and calculate its ratio to all P OKS The ratio is the AP value, and its formula is Using P OKS 0.50, 0.75, and 0.90 (50AP, 75AP, and 90AP) represent the matching accuracies of different requirements for key points, and the average mAP is used to represent the average accuracy of all key points. To measure the scale of the algorithm, the frame rate is used to calculate the number of pictures detected per second. The experimental results of 5 models on the COCO dataset are shown in Table 1. Among them, the mAP of models such as HigherHRNet-w32, HRNet, Lightweight OpenPose, OpenPose, and DEKR-W32 are 81.1%, 79.8%, 82.4%, 80.5%, and 80.7% respectively, and the mAP of the algorithm model is 84.9%.
[0061] When testing the video clips of personnel in the natural gas station using the YOLO-Pose algorithm, in order to ensure that the algorithm can accurately identify and track the postures of personnel in various complex environments, the following key scenarios are mainly concerned: the detection of independent key points of personnel, the discrimination of overlapping key points of personnel, the recognition of multiple postures of personnel, and the posture estimation of key points of personnel at different distances.
[0062] Next, the test results of the human pose algorithm will be explained in detail.
[0063] 1. Detection of independent key points of personnel: In this scenario, the main concern is whether the algorithm can accurately detect each key point of a single person, such as the head, shoulders, arms, waist, thighs, and ankles. The staff in the natural gas station are far apart from each other and there is no occlusion. The YOLO-Pose algorithm successfully identified the size of the target box for each person and the human key points of the personnel. These key points are clearly marked and displayed, indicating that the algorithm can accurately locate and track the key parts of a single person in this environment and maintain good performance without occlusion. The YOLO-Pose algorithm can achieve high-precision detection of independent personnel key points in various environments by combining deep learning and pose estimation techniques.
[0064] 2. Identification of Key Points with Overlapping Persons: In complex scenarios, especially when there is close contact or partial occlusion among persons, the performance of human pose estimation algorithms faces a severe test. In such cases, the algorithm needs to have a high level of discrimination ability to distinguish and track the key points of overlapping human bodies. YOLO-Pose can still accurately identify the skeletal key points of human poses when facing various complex occlusion patterns such as severe occlusion, partial occlusion, close contact, and staggered occlusion. This ability not only demonstrates the strong performance of the algorithm in dealing with occlusion problems but also has important significance for practical application scenarios such as personnel activity monitoring and safety management in dense crowd environments like natural gas stations. The effectiveness of this algorithm lies in its ability to identify and distinguish individuals in complex interactions and maintain a high level of accuracy even when the line of sight is blocked or the space is limited. The YOLO-Pose algorithm shows significant advantages in this regard, and it can effectively identify and track individual key points in cases of close contact or partial occlusion.
[0065] 3. Estimation of Multiple Human Poses: In the natural gas station environment, the actions and postures of workers are diverse, covering various postures such as standing, walking, bending, and squatting. These different body postures not only reflect the diversity of daily work but also may indicate potential safety risks or hidden dangers in operations. The YOLO-Pose algorithm can accurately detect the skeletal key points of workers in various complex backgrounds with bending, standing, and different walking postures. This ability to recognize multiple postures is crucial for the safety management of natural gas stations because it can monitor the behavior of workers in real-time and promptly detect possible unsafe actions or potential risks.
[0066] 4. Pose Estimation of Key Points for People at Different Distances: In the natural gas station environment, people have a wide range of activities, and the distance from the camera varies accordingly, which poses a challenge to the scale invariance of the human pose estimation algorithm. Scale invariance refers to the ability of the algorithm to accurately identify and estimate human key points in images of different scales, which is crucial for ensuring the comprehensiveness and effectiveness of the monitoring system. The YOLO-Pose algorithm has shown excellent scale adaptability in actual tests. The research results indicate that regardless of whether people are at a far or near position from the camera, YOLO-Pose can accurately estimate the positions of their key points. In various complex environments at long distances, medium distances, relatively short distances, and short distances, the algorithm can effectively detect the skeletal key points, demonstrating its stability and accuracy at different scales. This scale invariance is of great significance for the monitoring of large natural gas stations. Since the staff may be active at any position within the camera's field of view, the algorithm must be able to remain efficient at these different scales. This feature of YOLO-Pose ensures that the monitoring system can quickly respond to any potential safety risks within the station, take timely measures, and thus guarantee the safety of the staff and the stable operation of the station.
Claims
1. A method for monitoring personnel status based on the YOLO algorithm, characterized in that, Comprising: Obtaining real-time image or video stream data from a camera at a natural gas station; Determining the range of the human body in the image or video stream; Using the YOLO algorithm to obtain human key point information and perform pose estimation on the human body; Establishing an action model based on the result of the pose estimation; Judging the personnel status according to the action model.
2. The method for monitoring personnel status based on the YOLO algorithm according to claim 1, wherein, When performing pose estimation on the human body, the optimized YOLO algorithm is adopted, and the accuracy of key point detection is improved by introducing the CBAM attention mechanism.
3. The personnel status monitoring method based on the YOLO algorithm according to claim 2, wherein The optimized YOLO algorithm improves the detection ability under complex backgrounds through multi-scale feature fusion technology.
4. The method for monitoring personnel status based on the YOLO algorithm according to claim 1, wherein The human key point information is obtained through a predefined human skeleton model and at least includes key joint point information of the head, facial features, torso and limbs.
5. The method for monitoring personnel status based on the YOLO algorithm according to claim 1, characterized in that It also includes evaluating the result of the human pose estimation through a human pose estimation loss function, correcting the human key points, and optimizing the action model.
6. The method for monitoring the personnel state based on the YOLO algorithm according to claim 5, wherein, When correcting the human key points, based on historical frame data and a prediction model, the positions of the key points are smoothed to eliminate noise.
7. The method for monitoring the personnel status based on the YOLO algorithm according to claim 1, wherein The specific method for judging the personnel status is as follows: Comparing the action model of the current personnel with a pre-stored normal behavior sample library; Calculating the similarity index of key features and judging whether it conforms to the normal behavior pattern; Setting thresholds for pose deviation, action speed, and the degree of abnormality of key point positions; If the detected behavior feature deviation exceeds the threshold, it is considered that an abnormality exists.
8. The method for monitoring personnel status based on the YOLO algorithm according to claim 1 or 7, characterized in that, It also includes generating an alarm signal and notifying the management terminal when the personnel status is judged to be abnormal.
9. The method for monitoring the state of personnel based on the YOLO algorithm according to claim 8, wherein, The alarm signal includes real-time alarm information sent to the management terminal and triggers an alarm device in the event of a major abnormality.
10. A personnel status monitoring system based on the YOLO algorithm, which executes the personnel status monitoring method based on the YOLO algorithm described in any one of claims 1 to 9, characterized in that, Comprising: A camera for obtaining images or video streams of a natural gas station; A data processing module for performing recognition processing on the images or video streams; An alarm module for generating and transmitting alarm signals; A management terminal for receiving the alarm signals generated by the alarm module and performing monitoring and management operations.
Citation Information
Patent Citations
A pedestrian detection method based on Embedded YOLO algorithm
CN112487915B
Cited By
Video monitoring early warning method based on AI analysis
CN121170684A
Production line personnel behavior identification method and system
CN121305679A