Mining area personnel dangerous behavior early warning method and system

By using a combination of multi-task deep learning model and YOLOv5 learning model in the mine operation environment, the problem of identifying hazardous behaviors in the mine operation environment is solved, and the rapid and accurate identification and early warning of dangerous behaviors of people in the mine area is achieved, and the overall safety of mining area production is improved.

CN120014529APending Publication Date: 2025-05-16SICHUAN GUXU COAL DEV CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411889754.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In mine operation environments, it is difficult for the prior art to effectively identify and early warning of dangerous behaviors of staff, especially in low light conditions, where the accuracy and speed of image recognition are limited.

Method used

The multi-task deep learning model is used as the backbone network, combined with the YOLOv5 learning model to extract the advanced features of video frames, and multiple task branches are added on it, including OpenPose, Mask R-CNN, Faster R-CNN and ResNet50-LSTM models, which are used to identify personnel off posts, not wearing hard helmets, belt deviations and 'three violations' statistical analysis respectively.

Benefits of technology

It has achieved rapid and accurate identification and early warning of dangerous behaviors of people in mining areas, improved the overall safety of production in mining areas, and eliminated the occurrence of dangerous behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014529A_ABST
    Figure CN120014529A_ABST
Patent Text Reader

Abstract

The invention discloses a mining area personnel dangerous behavior early warning method and system, and relates to the technical field of intelligent video monitoring, and the method comprises the steps: carrying out the region division according to production functions; acquiring corresponding video data in different areas; based on the multiple pieces of behavior video data, a multi-task deep learning model is adopted as a backbone network, and advanced features of video frames are extracted; on the basis of a backbone network, a plurality of task branches are added, and each branch is responsible for different identification tasks; a corresponding alarm is given out according to the recognition result; according to the invention, early warning can be comprehensively and accurately carried out on dangerous behaviors of workers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent video monitoring, and in particular to a method and system for early warning of dangerous behaviors of personnel in mining areas. Background Art

[0002] my country is a country rich in mineral resources. However, since mineral mining operations involve many large machines and are generally located in mountainous areas or underground environments, mineral mining work is often accompanied by more risks.

[0003] The underground environment of a mine is generally complex, damp and cold, with muddy tunnels and dim lighting, which makes it easy for operators to commit dangerous and illegal acts. In addition, for operating areas with different functions, the types of large machines in the area are also different, so the requirements for standardized production are also different.

[0004] For example, in the ore conveying area, if the workers are too close to the conveyor belt, they may be knocked down by the conveyor belt, causing the ore on the belt surface to fall and injure the workers, resulting in a safety accident; or in the dangerous mine tunnel area where there is a risk of landslide, if the workers approach the area, then once a landslide occurs, the workers' safety factor will be greatly reduced.

[0005] At present, in traditional scenarios (non-underground operation scenarios), there are already a large number of existing technologies to identify operators' falls, climbing over fixed areas, approaching dangerous equipment, etc., but this fall behavior model cannot be directly applied to underground operation scenarios. Because the underground operation scene is dimly lit, the images collected by the video stream are low-light images, which need to be processed. However, new noise is easily introduced during the low-light image data processing, which brings greater uncertainty to the subsequent image recognition. In addition, the underground operation environment is complex and there are many obstructions. These dangerous and illegal behaviors are quite different from daily falls, climbing over, approaching and other dangers or violations, which brings difficulties to the identification work.

[0006] The Chinese patent with the publication number CN113111840A discloses a method for early warning of illegal and dangerous behaviors of workers in coal mine fully mechanized working faces. It adopts a new non-blind denoising method to pre-process image data, uses a pedestrian border detection model to identify workers to obtain borders, and uses the pre-defined dangerous equipment ROI area to determine whether there are dangerous behaviors such as falling, climbing over into the scraper, and approaching the cutting head. It uses a method combining skeleton key point recognition and sequence model to identify the dangerous behavior of falling of workers around the equipment, and optimizes the algorithm of the judgment process, thereby improving the judgment accuracy and ensuring that the judgment results are accurate and timely.

[0007] However, this method is only used to judge the actions of workers and does not take into account the overall safety of mining production, which has certain limitations.

[0008] Therefore, we propose a system that can comprehensively and accurately warn workers of dangerous behaviors. Summary of the invention

[0009] The purpose of the present invention is to provide a method and system for early warning of dangerous behaviors of personnel in mining areas, which can comprehensively and accurately warn of dangerous behaviors of personnel.

[0010] The present invention is achieved through the following technical solutions:

[0011] A method for early warning of dangerous behaviors of personnel in mining areas, comprising:

[0012] Regional division according to production function;

[0013] Obtain corresponding video data in different areas;

[0014] Based on multiple behavioral video data, a multi-task deep learning model is used as the backbone network to extract high-level features of video frames;

[0015] On the basis of the backbone network, multiple task branches are added, each branch is responsible for different recognition tasks;

[0016] According to the recognition results, corresponding alarms are issued.

[0017] Furthermore, the backbone network adopts the YOLOv5 learning model.

[0018] Furthermore, the process of extracting high-level features of video frames using the YOLOv5 learning model is as follows:

[0019] Preprocess the video data;

[0020] Input the preprocessed video frames into the YOLOv5 learning model;

[0021] Through the convolutional layer, residual block and downsampling layer in sequence, the low-level features and high-level features of the video frame are gradually extracted.

[0022] Furthermore, the pre-processing process is:

[0023] Crop and scale video frames to a fixed size;

[0024] Normalize pixel values ​​to between 0 and 1;

[0025] Use data augmentation techniques to process normalized video data.

[0026] Furthermore, the multiple task branches include an OpenPose model, a Mask R-CNN model, a Faster R-CNN model, and a ResNet50-LSTM model, wherein the OpenPose model is used to complete the identification of personnel off-duty;

[0027] The Mask R-CNN model is used to identify people who are not wearing helmets;

[0028] The Faster R-CNN model is used to complete belt deviation identification;

[0029] The ResNet50-LSTM model is used to complete the "three violations" statistical analysis.

[0030] Furthermore, the OpenPose model extracts key points of the human body from high-level features through convolutional layers and upsampling layers, thereby determining whether a person is out of position.

[0031] Furthermore, the processing process of the Mask R-CNN model is:

[0032] Generate a mask for each object from high-level features through convolutional layers and RoI pooling layers;

[0033] Based on the generated object masks, identify the masks of all personnel and the masks of all helmets;

[0034] Extract bounding boxes from masks of person and helmet;

[0035] Calculate the intersection and union ratio of each person's bounding box with all helmet bounding boxes;

[0036] If the intersection-over-union ratio of a person's bounding box and a helmet's bounding box is greater than a certain threshold, the person is considered to be wearing a helmet.

[0037] Furthermore, the processing process of the ResNet50-LSTM model is:

[0038] Obtain video data containing "three violations" and other normal behaviors;

[0039] Decomposing video data into a continuous sequence of frames;

[0040] Label each video frame to mark whether it contains the "three violations" behavior;

[0041] Use ResNet50 to extract features from video frames;

[0042] The extracted feature sequence and the decomposed continuous frame sequence are used as the input of the LSTM model;

[0043] The LSTM model outputs the prediction results for each time step to determine whether there are any "three violations" behaviors.

[0044] A mine personnel dangerous behavior early warning system, including a camera, an intelligent image recognition device and a video intelligent recognition and analysis server;

[0045] The cameras are distributed in various functional areas within the mining area and are used to obtain corresponding video data in different areas;

[0046] The intelligent image recognition device is provided in each functional area and is used to perform specific recognition tasks on the video data of each functional area;

[0047] The video intelligent recognition and analysis server is used to receive the output results of multiple intelligent image recognition devices, manage and preview videos, identify potential safety hazards, and manage them according to type, and has real-time alarm, event backtracking, fault self-diagnosis and record query functions.

[0048] Furthermore, the camera also includes a fill light for improving the clarity of the video captured by the camera.

[0049] The technical solution of the present invention has at least the following advantages and beneficial effects:

[0050] The invention discloses a method for early warning of dangerous behaviors of personnel in mining areas. By adopting a model of adding multiple task branches on the basis of a backbone network, the speed of identifying and analyzing the behaviors of personnel in various areas can be effectively improved, thereby quickly issuing early warnings and preventing the occurrence of dangerous behaviors of personnel.

[0051] In addition, there is a corresponding task branch model for each functional area, and specific identification and analysis can be carried out to comprehensively obtain various personnel behaviors in the mining area. In addition, due to its strong targeting, the identification and analysis results of personnel are also more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 The present invention is a schematic flow chart of a method. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0054] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0056] In the description of the present invention, it should be noted that if the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the accompanying drawings, or is the orientation or position relationship in which the product of the application is usually placed when used. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0057] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0058] Example 1

[0059] As attached Figure 1 A method for early warning of dangerous behavior of personnel in a mining area is shown, comprising:

[0060] Regional division according to production function;

[0061] Including areas where personnel are prohibited from entering, such as adit and inclined tunnels, tunnel production areas, and belt feeding areas;

[0062] Obtain corresponding video data in different areas;

[0063] Based on the above areas, the video data obtained are: the images in the prohibited area, the images in the lane, and the images in the belt feeding area;

[0064] Based on multiple behavioral video data, a multi-task deep learning model is used as the backbone network to extract high-level features of video frames;

[0065] On the basis of the backbone network, multiple task branches are added, each branch is responsible for different recognition tasks;

[0066] That is, according to different mining areas, different task branches correspond to different identification tasks. For example, for areas where personnel are prohibited from entering, the corresponding identification task is to determine whether there are workers entering the area; for tunnel production areas, the corresponding identification task is to determine whether there are workers who are off duty or sleeping in the area; for belt feeding areas, the corresponding identification task is to determine whether the belt in the area is deviated, and whether there are workers nearby when the belt deviates; and in all areas, the corresponding branches are used to determine whether the workers are wearing safety helmets and whether the workers have committed "three violations" in construction;

[0067] According to the recognition results, a corresponding alarm is issued;

[0068] That is, if a staff member enters an area where entry is prohibited, a staff member is off duty or sleeping in the tunnel production area, or a belt deviates in the belt feeding area and a staff member is nearby when the belt deviates, or in all areas, a staff member is not wearing a safety helmet, in the above situations, a warning will be issued through a voice broadcast, and if a staff member performs "three violations" construction, it will be recorded.

[0069] Example 2

[0070] The backbone network adopts the YOLOv5 learning model because the model has a good balance between real-time and accuracy.

[0071] In addition, the process of extracting high-level features of video frames using the YOLOv5 learning model is as follows:

[0072] Preprocess the video data;

[0073] The pre-processing process is as follows:

[0074] Crop and scale video frames to a fixed size, e.g., 416x416;

[0075] Normalize pixel values ​​to between 0 and 1;

[0076] Use data augmentation techniques to process normalized video data, where the data augmentation techniques include random cropping, flipping, and brightness adjustment methods to increase data diversity;

[0077] Input the preprocessed video frames into the YOLOv5 learning model;

[0078] Through the convolution layer, residual block and downsampling layer in sequence, the low-level features and high-level features of the video frame are gradually extracted, among which the high-level features include visual features, dynamic features and multimodal features; visual features are extracted through convolutional neural networks, dynamic features are extracted through optical flow and LSTM / GRU, and features of different modalities are combined through multimodal fusion technology. These high-level features can effectively represent objects, actions and scenes in the video, providing rich information for multi-task deep learning models.

[0079] Example 3

[0080] The multiple task branches include an OpenPose model, a Mask R-CNN model, a Faster R-CNN model, and a ResNet50-LSTM model, wherein the OpenPose model is used to complete the off-duty identification of personnel;

[0081] The Mask R-CNN model is used to identify people who are not wearing helmets;

[0082] The Faster R-CNN model is used to complete belt deviation identification;

[0083] The ResNet50-LSTM model is used to complete the "three violations" statistical analysis.

[0084] The OpenPose model extracts key points of the human body from high-level features through convolutional layers and upsampling layers to determine whether a person is out of his or her post.

[0085] The processing process of the Mask R-CNN model is:

[0086] Generate a mask for each object from high-level features through convolutional layers and RoI pooling layers;

[0087] Based on the generated object masks, identify the masks of all personnel and the masks of all helmets;

[0088] Extract bounding boxes from masks of person and helmet;

[0089] Calculate the intersection and union ratio of each person's bounding box with all helmet bounding boxes;

[0090] If the intersection-over-union ratio of a person's bounding box and a helmet's bounding box is greater than a certain threshold, the person is considered to be wearing a helmet.

[0091] The processing of the Faster R-CNN model is similar to that of the Mask R-CNN model. The Faster R-CNN model is used to identify the masks of the belt and conveyor rollers to extract the bounding boxes and perform subsequent calculations.

[0092] The processing process of the ResNet50-LSTM model is:

[0093] Obtain video data containing "three violations" and other normal behaviors, where the video data is external input data for comparison of subsequent LSTM models;

[0094] Decomposing video data into a continuous sequence of frames;

[0095] Label each video frame to mark whether it contains the "three violations" behavior;

[0096] Use ResNet50 to extract features from video frames;

[0097] The extracted feature sequence and the decomposed continuous frame sequence are used as the input of the LSTM model. The LSTM model compares the real-time video captured by the camera with the input video to see if the "three violations" are consistent.

[0098] The LSTM model outputs the prediction results for each time step to determine whether there are any "three violations" behaviors.

[0099] Example 4

[0100] A mine personnel dangerous behavior early warning system, including a camera, an intelligent image recognition device and a video intelligent recognition and analysis server;

[0101] The cameras are distributed in various functional areas within the mining area and are used to obtain corresponding video data in different areas;

[0102] The intelligent image recognition device is provided in each functional area, and is used to perform specific recognition tasks on the video data of each functional area; that is, the intelligent image recognition device can directly perform the recognition task of the area to improve the recognition efficiency, and the intelligent image recognition device includes a backbone network and corresponding branch tasks. For example, the intelligent image recognition device located in the tunnel production area includes a backbone network, an OpenPose model, a Mask R-CNN model, and a ResNet50-LSTM model, which are used to identify whether there are workers who are off duty or sleeping in the tunnel, whether the workers are wearing safety helmets, and whether the construction actions of the workers are "three violations";

[0103] The video intelligent recognition and analysis server is used to receive the output results of multiple intelligent image recognition devices, manage and preview videos, identify potential safety hazards, and manage them according to type, and has real-time alarm, event backtracking, fault self-diagnosis and record query functions.

[0104] In addition, based on this system, the following operations can be performed on the behaviors of the alarms identified by the task:

[0105] 1) Area intrusion and vehicle-free identification

[0106] According to the designated prohibited areas, the illegal behavior of people entering these dangerous areas can be identified and voice alarms can be issued on the spot. After identifying the violation, pictures and videos can be captured to form alarm records, and the violation alarm information can be pushed to the client; the violators can be remotely called and intercomed; the violation alarm information can be screened and confirmed, and the penalty results can be generated.

[0107] In addition, if the system identifies a person who has entered a prohibited area such as a horizontal tunnel or an inclined tunnel, a voice alarm will be issued on the spot to prevent the person from entering and causing a safety accident. At the same time, if someone is identified in the tunnel, the hardware safety circuit of the winch can be disconnected, the winch can be locked and the vehicle can be started; after the winch is running, if a person enters the tunnel, a voice alarm will be issued in time, and the vehicle will be stopped by linkage control. After identifying a violation, pictures and videos can be captured to form an alarm record, and the violation alarm information can be pushed to the client; the violator can be remotely called and intercomed; the violation alarm information can be screened and confirmed, and the penalty result can be generated.

[0108] 2) Identification of personnel off duty

[0109] Identify the violation of personnel leaving the work area for more than the specified time (the off-duty determination time can be set). After identifying the violation, you can capture pictures and videos to form an alarm record, push the violation alarm information to the client, filter and confirm the violation alarm information, and generate the penalty result.

[0110] 3) Identification of personnel sleeping on duty

[0111] Identify the violation of personnel sleeping on duty and issue a voice alarm on the spot. After identifying the violation, you can capture pictures and videos to form an alarm record, push the violation alarm information to the client, filter and confirm the violation alarm information, and generate the penalty result.

[0112] 4) Identification of not wearing a helmet

[0113] Identify the violation of personnel not wearing a helmet and issue a voice alarm on the spot. After identifying the violation, you can capture pictures and videos to form an alarm record, push the violation alarm information to the client, filter and confirm the violation alarm information, and generate the penalty result.

[0114] 5) Belt deviation identification

[0115] Identify hidden dangers of belt deviation. After identifying the hidden danger, you can capture pictures and videos to form an alarm record, push the alarm information to the client, and output linkage control information to the on-site control system.

[0116] In particular, the camera also includes a fill light for improving the clarity of the video captured by the camera.

[0117] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for early warning of dangerous behaviors of personnel in mining areas, characterized by: include: Regional division according to production function; Obtain corresponding video data in different areas; Based on multiple behavioral video data, a multi-task deep learning model is used as the backbone network to extract high-level features of video frames; On the basis of the backbone network, multiple task branches are added, each branch is responsible for different recognition tasks; According to the recognition results, corresponding alarms are issued.

2. The method for early warning of dangerous behaviors of personnel in mining areas according to claim 1, characterized in that: The backbone network adopts the YOLOv5 learning model.

3. The method for early warning of dangerous behaviors of personnel in mining areas according to claim 2, characterized in that: The process of extracting high-level features of video frames using the YOLOv5 learning model is as follows: Preprocess the video data; Input the preprocessed video frames into the YOLOv5 learning model; Through the convolutional layer, residual block and downsampling layer in sequence, the low-level features and high-level features of the video frame are gradually extracted.

4. The method for early warning of dangerous behaviors of personnel in mining areas as claimed in claim 3, characterized in that: The pre-processing process is as follows: Crop and scale video frames to a fixed size; Normalize pixel values ​​to between 0 and 1; Use data augmentation techniques to process normalized video data.

5. The method for early warning of dangerous behaviors of personnel in mining areas according to claim 1, characterized in that: The multiple task branches include an OpenPose model, a Mask R-CNN model, a Faster R-CNN model, and a ResNet50-LSTM model, wherein the OpenPose model is used to complete the off-duty identification of personnel; The Mask R-CNN model is used to identify people who are not wearing helmets; The Faster R-CNN model is used to complete belt deviation identification; The ResNet50-LSTM model is used to complete the "three violations" statistical analysis.

6. The method for early warning of dangerous behaviors of personnel in mining areas according to claim 5, characterized in that: The OpenPose model extracts key points of the human body from high-level features through convolutional layers and upsampling layers, thereby determining whether a person is out of position.

7. The method for early warning of dangerous behaviors of personnel in mining areas according to claim 5, characterized in that: The processing process of the Mask R-CNN model is: Generate a mask for each object from high-level features through convolutional layers and RoI pooling layers; Based on the generated object masks, identify the masks of all personnel and the masks of all helmets; Extract bounding boxes from masks of person and helmet; Calculate the intersection and union ratio of each person's bounding box with all helmet bounding boxes; If the intersection-over-union ratio of a person's bounding box and a helmet's bounding box is greater than a certain threshold, the person is considered to be wearing a helmet.

8. The method for early warning of dangerous behaviors of personnel in mining areas according to claim 5, characterized in that: The processing process of the ResNet50-LSTM model is: Obtain video data containing "three violations" and other normal behaviors; Decomposing video data into a continuous sequence of frames; Label each video frame to mark whether it contains "three violations"; Use ResNet50 to extract features from video frames; The extracted feature sequence and the decomposed continuous frame sequence are used as the input of the LSTM model; The LSTM model outputs the prediction results for each time step to determine whether there are "three violations" behaviors.

9. A warning system for dangerous behaviors of personnel in mining areas, including a camera, an intelligent image recognition device and a video intelligent recognition and analysis server; The cameras are distributed in various functional areas within the mining area and are used to obtain corresponding video data in different areas; The intelligent image recognition device is provided in each functional area and is used to perform specific recognition tasks on the video data of each functional area; The video intelligent recognition and analysis server is used to receive the output results of multiple intelligent image recognition devices, manage and preview videos, identify potential safety hazards, and manage them according to type, and has real-time alarm, event backtracking, fault self-diagnosis and record query functions.

10. The mine personnel dangerous behavior early warning system as claimed in claim 9, characterized in that: The camera also includes a fill light for improving the clarity of the video captured by the camera.

Citation Information

Patent Citations

  • A coal mine fully mechanized coal mining face operator violation and dangerous behavior early warning method

    CN113111840A