Method and device for identifying intrusion of underground coal mine personnel

By using monocular depth estimation method and bone key point technology underground in coal mines, combined with Kalman filtering and Hungarian matching algorithm, the problem of low pedestrian detection accuracy in coal mines is solved, the false alarm rate is reduced, and more accurate pedestrian intrusion recognition is achieved.

CN119296133BActive Publication Date: 2025-07-29CHINA COAL RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411303222.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-07-29
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

The existing pedestrian detection methods have low accuracy in dim and fuzzy environments under coal mines, resulting in a high false alarm rate for electronic fence intrusion detection.

Method used

The depth image of the image is determined by monocular depth estimation method, and the two-dimensional and three-dimensional coordinates of the detection frame and bone key points of the pedestrian target are detected, combined with Kalman filtering and Hungarian matching algorithm, it is determined whether the pedestrian invades the electronic fence area of the target space.

Benefits of technology

It improves the accuracy of pedestrian detection in coal mines, reduces the false alarm rate of electronic fence intrusion detection, and achieves more accurate pedestrian location tracking and intrusion identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119296133B_ABST
    Figure CN119296133B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for identifying the intrusion of personnel underground in a coal mine. The method includes: acquiring a current frame image and determining a depth image corresponding to the current frame image by using a monocular depth estimation method; detecting pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and extracting skeletal key points based on the current frame image in the regions where the detection frames are located to obtain the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame; determining the three-dimensional coordinates of the skeletal key points of each pedestrian target according to the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame, the depth image, and the internal parameter matrix of the image acquisition device; and determining whether each pedestrian target intrudes into the target space electronic fence area according to the three-dimensional coordinates of each detection frame and / or the skeletal key points of each pedestrian target. Thus, the accuracy of pedestrian detection underground in a coal mine can be improved and the false alarm rate can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for identifying the intrusion of personnel underground in coal mines. Background Art

[0002] In dangerous environments such as underground coal mines, ensuring the safety of workers is of utmost importance. The underground environment is relatively complex, with many unsafe factors. Preventing underground workers from entering areas with unsafe factors is an important measure to ensure the safety of underground workers.

[0003] Visible light cameras are usually installed underground in coal mines to monitor unsafe factors in the underground environment. However, manually monitoring camera images is usually inefficient, with low detection rates, and it is difficult to concentrate on monitoring camera images for a long time. Nowadays, artificial intelligence technology has been widely applied in many fields, and pedestrian detection based on visible light images has also been applied in many fields. However, most of the existing pedestrian detection methods are aimed at general ground scenes, and the accuracy is not high when applied to the dim and blurred environment underground in coal mines, resulting in many false alarms in the intrusion detection results of the electronic fence. Summary of the Invention

[0004] The present invention provides a method and device for identifying the intrusion of personnel underground in coal mines to at least solve one of the technical problems in the related art to a certain extent. The technical solution of the present invention is as follows:

[0005] According to the first aspect of the embodiments of the present invention, a method for identifying the intrusion of personnel underground in coal mines is provided, including: obtaining a current frame image collected by an image acquisition device arranged underground in a coal mine, and determining a depth image corresponding to the current frame image by using a monocular depth estimation method; detecting pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and extracting skeletal key points based on the current frame image in the area where each detection frame is located to obtain two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame; wherein, the detection frame is a matrix frame that minimally encloses the corresponding pedestrian target; determining three-dimensional coordinates of the skeletal key points of each pedestrian target according to the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame, the depth image, and the internal parameter matrix of the image acquisition device; determining whether each pedestrian target intrudes into the target space electronic fence area according to the three-dimensional coordinates of each detection frame and / or the skeletal key points of each pedestrian target; wherein, the target space electronic fence is determined based on the depth image and a preset area to be determined.

[0006] According to a second aspect of the embodiments of the present invention, there is provided a device for identifying intrusion of personnel underground in a coal mine, including: a first processing module, configured to obtain a current frame image collected by an image acquisition device arranged underground in a coal mine, and determine a depth image corresponding to the current frame image by using a monocular depth estimation method; a second processing module, configured to detect pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and perform skeleton key point extraction on the current frame image in the regions where the detection frames are located to obtain two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each of the detection frames; wherein, the detection frame is a matrix frame that minimally encloses the corresponding pedestrian target; a first determination module, configured to determine three-dimensional coordinates of the skeleton key points of each pedestrian target according to the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection frame, the depth image, and the internal parameter matrix of the image acquisition device; a second determination module, configured to determine whether each pedestrian target intrudes into a target space electronic fence area according to each detection frame and / or the three-dimensional coordinates of the skeleton key points of each pedestrian target; wherein, the target space electronic fence is determined based on the depth image and a preset area to be determined.

[0007] According to a third aspect of the embodiments of the present invention, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the method for identifying intrusion of personnel underground in a coal mine as described in the embodiments of the first aspect of the present invention.

[0008] According to a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method for identifying intrusion of personnel underground in a coal mine as described in the embodiments of the first aspect of the present invention.

[0009] According to a fifth aspect of the embodiments of the present invention, there is provided a computer program product, including: a computer program, which when executed by a processor, implements the method for identifying intrusion of personnel underground in a coal mine as described in the embodiments of the first aspect of the present invention.

[0010] The technical solutions provided by the embodiments of the present invention at least bring the following beneficial effects:

[0011] In this technical solution, the current frame image collected by the image acquisition device arranged underground in the coal mine is obtained, and the monocular depth estimation method is used to determine the depth image corresponding to the current frame image; the pedestrian targets in the current frame image are detected to obtain the detection frames corresponding to at least one pedestrian target in the current frame image, and the two-dimensional coordinates of the skeletal key points of the pedestrian targets corresponding to the detection frames are extracted based on the current frame image in the regions where the detection frames are located; wherein, the detection frame is a matrix frame that minimally encloses the corresponding pedestrian target; according to the two-dimensional coordinates of the skeletal key points of the pedestrian targets corresponding to the detection frames, the depth image, and the internal parameter matrix of the image acquisition device, the three-dimensional coordinates of the skeletal key points of the pedestrian targets are determined; according to the three-dimensional coordinates of the detection frames and / or the skeletal key points of the pedestrian targets, it is determined whether each pedestrian target invades the target space electronic fence area; wherein, the target space electronic fence is determined based on the depth image and the preset area to be determined. Thus, aiming at the problem that there are a large number of false alarms when the existing pedestrian detection methods are applied in the dim and fuzzy environment underground in the coal mine, resulting in a large number of false alarms in the electronic fence intrusion detection, the present invention combines monocular depth estimation and human bones to perform three-dimensional position tracking on the pedestrian targets in the current frame image collected by the image acquisition device arranged underground in the coal mine, which can improve the accuracy of pedestrian detection underground in the coal mine and reduce the false alarm rate of the electronic fence intrusion detection.

[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention, and do not constitute an improper limitation of the present invention.

[0014] Figure 1 is a schematic flow chart of the method for identifying the intrusion of personnel underground in the coal mine shown in the first embodiment of the present invention;

[0015] Figure 2 is a schematic flow chart of the method for identifying the intrusion of personnel underground in the coal mine shown in the second embodiment of the present invention;

[0016] Figure 3 is a schematic flow chart of the method for identifying the intrusion of personnel underground in the coal mine shown in the third embodiment of the present invention;

[0017] Figure 4 is a schematic flow chart of the method for identifying the intrusion of personnel underground in the coal mine shown in the fourth embodiment of the present invention;

[0018] Figure 5 is a schematic flow chart of the method for identifying the intrusion of personnel underground in the coal mine shown in the fifth embodiment of the present invention;

[0019] Figure 6 Schematic diagram of the principle for identifying the intrusion of underground coal mine personnel provided by the sixth embodiment of the present invention;

[0020] Figure 7 Schematic diagram of the structure of the device for identifying the intrusion of underground coal mine personnel shown in the seventh embodiment of the present invention;

[0021] Figure 8 Schematic diagram of the structure of the electronic device 800 shown in an exemplary embodiment of the present invention. Detailed implementation manners

[0022] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0023] It should be noted that in the description and claims of the present invention and the above-mentioned accompanying drawings, the terms "first", "second", etc. are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0024] It should be noted that in the technical solutions of the present invention, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing are all carried out on the premise of obtaining the user's consent, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0025] Different from open spaces, the underground coal mine space is closed, narrow, and dim, with many obstructions, which poses great challenges to the detection of underground coal mine workers. Most of the existing pedestrian detection methods are for general ground scenarios. When applied in the dim and blurred environment of underground coal mines, the accuracy is not high, often resulting in many false alarms, bringing quite a lot of trouble to subsequent further processing such as pedestrian tracking, and also causing a relatively high false alarm rate for the intrusion detection of the electronic fence.

[0026] In view of the above problems, the present invention proposes a method for identifying the intrusion of underground coal mine personnel. By combining monocular depth estimation and human skeletons, more robust target tracking methods, motion detection, multi-point detection and other means can be realized, which can, while ensuring the accuracy of alarms, monitor and track the positions of personnel in restricted areas in real time, issue alarms in a timely manner and take necessary measures, thereby minimizing the possibility of accidents to the greatest extent.

[0027] The method and device for identifying the intrusion of underground coal mine personnel according to the embodiments of the present invention will be described below with reference to the accompanying drawings.

[0028] Figure 1 It is a schematic flowchart of the method for identifying the intrusion of underground coal mine personnel shown in the first embodiment of the present invention.

[0029] In the embodiments of the present invention, the method for identifying the intrusion of underground coal mine personnel is taken as an example of being configured in a device for identifying the intrusion of underground coal mine personnel. The device for identifying the intrusion of underground coal mine personnel can be applied to any electronic device so that the electronic device can perform the function of identifying the intrusion of underground coal mine personnel.

[0030] Among them, the electronic device can be any device with computing ability, such as a personal computer, a mobile terminal, a server (or cloud), etc. The mobile terminal can be a hardware device with various operating systems, a touch screen and / or a display screen, such as a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc.

[0031] As Figure 1 shown, the method for identifying the intrusion of underground coal mine personnel includes the following steps:

[0032] Step 101: Obtain the current frame image collected by the image acquisition device arranged underground in the coal mine, and determine the depth image corresponding to the current frame image by using the monocular depth estimation method.

[0033] Among them, the image acquisition device can be any device used to capture image data, such as a camera, a video camera, a camera, other devices with a photographing function (such as a mobile phone, a tablet computer, etc.), and so on.

[0034] In the embodiments of the present invention, the device for identifying the intrusion of underground coal mine personnel can obtain the current frame image collected by the image acquisition device arranged underground in the coal mine through various public, legal, and compliant means. For example, during the process of the image acquisition device collecting images, the images collected by the image acquisition device can be obtained in real time through network transmission.

[0035] Among them, the depth image (Depth Images), also known as the range image (Range Image), is a special image representation method. Its core lies in recording the distance (depth) value from the image acquisition device to each point in the scene as the pixel value. This kind of image directly reflects the geometric shape of the visible surface of the scene, is the projection of three-dimensional space information on a two-dimensional plane, and is of great significance for understanding the spatial structure of objects and performing three-dimensional reconstruction.

[0036] In the present invention, in order to better track the position of pedestrians, a monocular depth estimation method is adopted to estimate the depth of the current frame image. As a possible implementation, the current frame image can be input into a monocular depth estimation network to obtain the depth image corresponding to the current frame image, from which the depth information d of each pixel point p can be obtained. p .

[0037] Step 102: Detect pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and perform skeleton key point extraction based on the current frame image in the regions where the detection frames are located to obtain the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection frame.

[0038] In the embodiment of the present invention, the coal mine underground personnel intrusion recognition device can detect pedestrian targets in the current frame image, so as to obtain detection frames corresponding to at least one pedestrian target in the current frame image. Among them, the detection frame is a matrix frame that minimally encloses the corresponding pedestrian target. Optionally, the detection frame can be output in the form of a quadruple, and the four numbers included respectively represent the positions of the left, top, right, and bottom sides of the matrix frame that minimally encloses the corresponding pedestrian target, represented by l, t, r, and b respectively. Then, the upper left corner point of the rectangular frame can be represented as p (lt) = [l, t] T , and the lower right corner point can be represented as p (rb) = [r, b] T , and the center point can be represented as p (center) = [(l + r) / 2, (t + b) / 2] T .

[0039] As a possible implementation, a coal mine underground pedestrian dataset can be used to train a deep neural network to obtain a pedestrian detection network optimized using the coal mine underground pedestrian dataset, and then this pedestrian detection network is used to detect pedestrian targets in the current frame image. Specifically, the current frame image can be input into this pedestrian detection network to obtain the detection frames corresponding to at least one pedestrian target output by this pedestrian detection network.

[0040] In the present invention, in order to better track the position of pedestrians, after obtaining the detection frames corresponding to each pedestrian target, skeleton key point extraction is performed based on the current frame image in the regions where the detection frames are located to obtain the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection frame. Among them, the skeleton key points can include 17 human body skeleton key points such as nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle.

[0041] As a possible implementation, for any detection box, the current frame image of the area where the detection box is located can be cropped and input into the skeleton key point extraction network to obtain the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian target in the detection box.

[0042] Step 103: According to the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection box, the depth image, and the internal parameter matrix of the image acquisition device, determine the three-dimensional coordinates of the skeleton key points of each pedestrian target.

[0043] As a possible implementation, the depth information d of each pixel point p can be obtained from the depth image corresponding to the current frame image p , and then using the internal parameter matrix K of the image acquisition device, the following calculation formula is used to determine the three-dimensional spatial coordinates of any pixel point p in the current frame image:

[0044] x = d p K -1 p

[0045] where the coordinates of the pixel point p are given in homogeneous coordinate form.

[0046] In summary, since the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection box are pixel points p one by one, the above calculation formula can be used to determine the three-dimensional coordinates of the skeleton key points of each pedestrian target.

[0047] Step 104: Determine whether each pedestrian target invades the target space electronic fence area according to the three-dimensional coordinates of each detection box and / or the skeleton key points of each pedestrian target.

[0048] Among them, the target space electronic fence is determined based on the depth image and a preset area to be determined.

[0049] Among them, the area to be determined can be any region of interest (ROI) preset by the user in the current frame image.

[0050] Among them, the ROI is a closed figure composed of an ordered two-dimensional point sequence, and the area on the right hand side along the point sequence order is the internal area of the closed figure. Let the ordered two-dimensional point sequence be p (ROI),i , i = 1, 2, …, L (ROI) The length of is L (ROI) , for any pixel point p in the current frame image, if the pixel point p is adjacent to the points in any ordered two-dimensional point sequence, that is, i < L (ROI) , j = i + 1 and i = L (ROI) , j = 1, there is: the pixel point p is at the pixel point p (ROI),i and the pixel point p (ROI),jIf the pixel point p is on the right side of the formed directed line segment, it can be determined that the pixel point p is inside the ROI.

[0051] It should be noted that the ROI preset by the user in the current frame image is a planar region in the current frame image. To achieve the tracking effect in space, the present invention will utilize monocular depth estimation. According to the preset ROI, the corresponding region in space (target space electronic region) of the ROI is extracted to perform intrusion detection in the target space electronic region.

[0052] As a possible implementation manner, the determination process of the target space electronic fence can be as follows: According to the depth image and the two-dimensional coordinates of each point in the two-dimensional point sequence constituting the region to be determined, the three-dimensional coordinates of each point in the two-dimensional point sequence are determined; the Hough transform is used to extract multiple planes within the region to be determined, and the number of inliers of each plane is used as the area occupied by the corresponding plane in the region to be determined; the plane with the largest area is selected from each plane as the ground plane, and the normal vector of the ground plane and the three-dimensional coordinates of the target intersection point are obtained; wherein, the target intersection point is the intersection point of the straight line along the normal vector direction of the ground plane from the coordinate origin and the ground plane; according to the normal vector of the ground plane and the three-dimensional coordinates of the target intersection point, the three-dimensional coordinates of each point in the two-dimensional point sequence are projected onto the ground plane to obtain the target space electronic fence.

[0053] For example, first, according to the depth image corresponding to the current frame image, the depth information d of each pixel point p within the ROI is extracted p , and the three-dimensional coordinates of each pixel point p within the ROI are determined; then the Hough transform is used to extract all planes and their normal vectors, and the number of inliers of each plane is calculated as the area occupied by the plane in the ROI; then the plane with the largest area is selected as the ground plane P, and its normal vector n is obtained P , as well as the intersection point o of the straight line along the normal vector direction of the ground plane from the coordinate origin and the ground plane; finally, the following calculation formula is used to project the three-dimensional coordinates x s,(ROI),i of the ordered two-dimensional point sequence constituting the ROI onto the ground plane:

[0054] x g,(ROI),i = x s,(ROI),i - (x g,(ROI),i - o)n P

[0055] The ROI on the ground plane is obtained, which is used as the basis for pedestrian electronic fence intrusion detection, that is, the target space electronic fence.

[0056] In the embodiment of the present invention, after obtaining each detection box and the three-dimensional coordinates of the skeletal key points of each pedestrian target, it is possible to determine whether each pedestrian target invades the target space electronic fence area based on each detection box and / or the three-dimensional coordinates of the skeletal key points of each pedestrian target.

[0057] As a possible implementation, the skeletal key points at least include the left and right ankle skeletal key points. For any pedestrian target, when the confidence levels of the left and right ankle skeletal key points of the pedestrian target are both greater than the confidence level threshold, the three-dimensional coordinates of the ankle skeletal key point with a higher confidence level among the left and right ankle skeletal key points of the pedestrian target are used as the foot position of the pedestrian target; when the confidence levels of the left and right ankle skeletal key points of the pedestrian target are not both greater than the confidence level threshold, the three-dimensional coordinates of the midpoint of the lower edge of the detection box corresponding to the pedestrian target are used as the foot position of the pedestrian target; wherein, the three-dimensional coordinates of the midpoint of the lower edge of the detection box corresponding to the pedestrian target are determined based on the two-dimensional coordinates of the midpoint of the lower edge of the detection box corresponding to the pedestrian target, the depth image, and the internal parameter matrix of the image acquisition device; project the foot position of the pedestrian target onto the ground plane to determine whether the foot position of the pedestrian target is within the target space electronic fence area, and when the foot position of the pedestrian target is within the target space electronic fence area, determine that the pedestrian target invades the target space electronic fence area; wherein, the ground plane is the ground plane determined during the process of determining the target space electronic fence area.

[0058] That is to say, in the present invention, the foot position of the pedestrian target is used as a reference to determine whether the pedestrian target invades the target space electronic fence area. For the case where the confidence levels of the left and right ankle skeletal key points are higher than the confidence level threshold, the present invention uses the ankle skeletal key point with a higher confidence level as the foot position of the pedestrian target, otherwise it is assumed that the midpoint of the lower edge of the detection box corresponding to the pedestrian target is the foot position. Then, project the three-dimensional coordinates of the determined foot position onto the ground plane, and further determine whether the pedestrian target is within the target space electronic fence area. If it is within the target space electronic fence area, it is considered that there is an intrusion situation.

[0059] The method for identifying the intrusion of personnel underground in a coal mine according to an embodiment of the present invention obtains the current frame image collected by an image acquisition device arranged underground in the coal mine, and determines the depth image corresponding to the current frame image by using a monocular depth estimation method; detects pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and extracts skeletal key points based on the current frame image in the area where each detection frame is located to obtain the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame; wherein, the detection frame is a matrix frame that minimally encloses the corresponding pedestrian target; determines the three-dimensional coordinates of the skeletal key points of each pedestrian target according to the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame, the depth image, and the internal parameter matrix of the image acquisition device; determines whether each pedestrian target intrudes into the target space electronic fence area according to the three-dimensional coordinates of each detection frame and / or the skeletal key points of each pedestrian target; wherein, the target space electronic fence is determined based on the depth image and a preset area to be determined. Therefore, for the existing pedestrian detection method, when applied in the dim and fuzzy environment underground in the coal mine, there are a large number of false alarms, resulting in a large number of false alarms in the electronic fence intrusion detection. The present invention combines monocular depth estimation and human bones to perform three-dimensional position tracking on pedestrian targets in the current frame image collected by an image acquisition device arranged underground in the coal mine, which can improve the accuracy of pedestrian detection underground in the coal mine and reduce the false alarm rate of electronic fence intrusion detection.

[0060] It should be noted that in the present invention, before determining whether each pedestrian target intrudes into the target space electronic fence area according to the three-dimensional coordinates of each detection frame and / or the skeletal key points of each pedestrian target, pedestrian tracking can also be performed to remove non-pedestrian targets in each pedestrian target. The following combines Figure 2 to illustrate this process.

[0061] Figure 2 is a schematic flowchart of the method for identifying the intrusion of personnel underground in a coal mine shown in the second embodiment of the present invention.

[0062] As Figure 2 shown, the method for identifying the intrusion of personnel underground in a coal mine includes the following steps:

[0063] Step 201, obtain the current frame image collected by an image acquisition device arranged underground in the coal mine, and determine the depth image corresponding to the current frame image by using a monocular depth estimation method.

[0064] Step 202, detect pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and extract skeletal key points based on the current frame image in the area where each detection frame is located to obtain the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame.

[0065] Step 203: Determine the three-dimensional coordinates of the skeletal key points of each pedestrian target based on the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian target in each detection box, the depth image, and the internal parameter matrix of the image acquisition device.

[0066] Step 204: Track the detection boxes of the pedestrian targets using the Kalman filter based on multiple frames of images before the current frame image to obtain the predicted detection boxes corresponding to at least one predicted pedestrian target in the current frame image, and / or track the skeletal key points of the pedestrian targets using the Kalman filter to obtain the predicted values of the three-dimensional coordinates of the skeletal key points of at least one predicted pedestrian target in the current frame image.

[0067] To achieve the tracking and filtering of the pedestrian positions, the present invention will use the Kalman filter method to track the detection boxes of the pedestrian targets and / or the skeletal key points of the pedestrian targets.

[0068] As a possible implementation, the detection box is composed of the two-dimensional coordinates of the left side, upper side, right side, and lower side of the matrix box that minimally encloses the corresponding pedestrian target. Thus, the two-dimensional coordinates of the upper left corner point and the lower right corner point of the detection box corresponding to at least one pedestrian target in multiple frames of images before the current frame image can be used as the state vector, the second-order difference of the first point to be tracked is used as the first preset vector, and the first-order difference remaining unchanged is taken as the target to construct the first motion equation:

[0069]

[0070] where, x i represents the two-dimensional coordinates of the first point to be tracked in the current frame image, x i-1 represents the two-dimensional coordinates of the first point to be tracked in the previous frame image before the current frame image, x i-2 represents the two-dimensional coordinates of the first point to be tracked in the frame image before the previous frame image before the current frame image,

[0071]

[0072] A x is the first state transition matrix, indicating that the second-order difference of the first point to be tracked is the first preset vector and the first-order difference remains unchanged, where I is the identity matrix; where the first point to be tracked is the upper left corner point or the lower right corner point of the detection box corresponding to any pedestrian target in multiple frames of images before the current frame image, the two-dimensional coordinates of the upper left corner point of any detection box are determined based on the two-dimensional coordinates of the left side and the upper side of the detection box, and the two-dimensional coordinates of the lower right corner point of any detection box are determined based on the two-dimensional coordinates of the right side and the lower side of the detection box;

[0073] Based on the first motion equation, determine the predicted detection boxes corresponding to at least one predicted pedestrian target in the current frame image; and / or,

[0074] The three-dimensional coordinates of the skeletal key points of at least one pedestrian target in multiple frames of images before the current frame image can be used as the state vector, the second-order difference of the second point to be tracked is used as the second preset vector, and the first-order difference remains unchanged as the target to construct the second motion equation:

[0075]

[0076] Among them, x′ i represents the three-dimensional coordinates of the second point to be tracked in the current frame image, x i-1 ′ represents the three-dimensional coordinates of the second point to be tracked in the previous frame image before the current frame image, x i-2 ′ represents the three-dimensional coordinates of the second point to be tracked in the frame image before the previous frame image before the current frame image,

[0077]

[0078] A x ′ is the second state transition matrix, indicating that the second-order difference of the second point to be tracked is the second preset vector and the first-order difference remains unchanged, where I is the identity matrix; among them, the second point to be tracked is any skeletal key point of any pedestrian target in multiple frames of images before the current frame image;

[0079] According to the second motion equation, determine the predicted values of the three-dimensional coordinates of the skeletal key points of at least one predicted pedestrian target in the current frame image.

[0080] Step 205, according to at least one of each predicted detection box, the predicted values of the three-dimensional coordinates of the skeletal key points of each predicted pedestrian target, each detection box, and the three-dimensional coordinates of the skeletal key points of each pedestrian target, use the Hungarian matching algorithm to match each pedestrian target and each predicted pedestrian target to remove non-pedestrian targets from each pedestrian target.

[0081] The present invention uses the Hungarian matching algorithm to determine the matching degree between each pedestrian target in the previous frame image and each pedestrian target in the current frame image, and matches the pedestrian targets between the two frames.

[0082] As a possible implementation, for any pair of pedestrian targets to be matched and predicted pedestrian targets, determine the number of skeletal key points with confidence higher than the confidence threshold in the set of skeletal key points of the pedestrian target and the predicted pedestrian target; when the number is not less than the number threshold, use the Euclidean distance of each pair of skeletal key points of the pedestrian target and the predicted pedestrian target as the judgment criterion, and adopt the Hungarian matching algorithm to match each pedestrian target and each predicted pedestrian target to remove non-pedestrian targets in each pedestrian target; when the number is less than the number threshold, use the intersection over union (IoU) of the detection box corresponding to the pedestrian target and the predicted detection box corresponding to the predicted pedestrian target as the judgment criterion, and adopt the Hungarian matching algorithm to match each pedestrian target and each predicted pedestrian target to remove non-pedestrian targets in each pedestrian target.

[0083] That is, during the pedestrian target matching process, the Euclidean distance of the skeletal key points and the IoU of the minimum bounding rectangle of the target are used as the measurement criteria.

[0084] For example, after obtaining the predicted detection box corresponding to at least one predicted pedestrian target in the current frame image and the predicted three-dimensional coordinate values of the skeletal key points of at least one predicted pedestrian target in the current frame image using Kalman filtering, use the Euclidean distance of the skeletal key points as the judgment criterion and match using the Hungarian matching algorithm. Specifically, use the average Euclidean distance l1 of the points of the 9 pairs of human skeletal key points such as nose, left eye (leye), right eye (reye), left ear (lear), right ear (rear), left shoulder (lshoulder), right shoulder (rshoulder), left hip (lhip), and right hip (rhip) in the two targets (pedestrian target and predicted pedestrian target) to be judged for matching, where the confidence c i is higher than the confidence threshold T c to determine whether to match:

[0085]

[0086] where w ∈ (nose, leye, reye, lear, rear, lshoulder, rshoulder, lhip, rhip).

[0087] If the number of points where the confidence c i is higher than the confidence threshold T c is less than 2 (the number threshold), then use the IoU of the minimum bounding rectangles of the two targets to be judged for matching to determine whether to match.

[0088] Moreover, after obtaining the matched detection targets, the state vector and covariance matrix of the Kalman filter will be updated.

[0089] In addition, when tracking, tracking and exit tracking thresholds are also set. The tracking is confirmed only when the target is tracked for multiple consecutive frames, and the loss of target tracking is confirmed only when the target is not tracked for multiple consecutive frames, so as to prevent the fluctuations of the target detection results within a short period of time from affecting the tracking of the target.

[0090] Step 206: Determine whether each pedestrian target invades the target space electronic fence area according to the three-dimensional coordinates of each detection box and / or the skeletal key points of each pedestrian target.

[0091] It should be noted that the execution processes of steps 201 to 203 and step 206 can be implemented in any one of the embodiments of the present invention respectively. The embodiments of the present invention do not make any limitations on this, nor will they be elaborated further.

[0092] The method for identifying the intrusion of underground coal mine personnel according to the embodiments of the present invention tracks the detection boxes of pedestrian targets by using Kalman filtering based on multiple frames of images before the current frame image to obtain the predicted detection boxes corresponding to at least one predicted pedestrian target in the current frame image, and / or tracks the skeletal key points of pedestrian targets by using Kalman filtering to obtain the predicted values of the three-dimensional coordinates of the skeletal key points of at least one predicted pedestrian target in the current frame image; according to at least one of each predicted detection box, the predicted values of the three-dimensional coordinates of the skeletal key points of each predicted pedestrian target, each detection box and the three-dimensional coordinates of the skeletal key points of each pedestrian target, the Hungarian matching algorithm is used to match each pedestrian target and each predicted pedestrian target to remove the non-pedestrian targets among each pedestrian target. Thus, the tracking and filtering of the pedestrian positions are realized by using Kalman filtering, and on the basis of Hungarian matching, the corner points of the minimum bounding rectangle and the bone points are used as the matching judgment basis of the Hungarian matching algorithm to realize target tracking, and the non-pedestrian targets among each pedestrian target can be accurately removed, improving the accuracy of pedestrian detection.

[0093] It should be noted that in the present invention, before determining whether each pedestrian target invades the target space electronic fence area according to the three-dimensional coordinates of each detection box and / or the skeletal key points of each pedestrian target, in addition to pedestrian tracking, target motion confirmation can also be performed to further remove the non-pedestrian targets among each pedestrian target. The following combines Figure 3 , to illustrate this process.

[0094] Figure 3 is a schematic flowchart of the method for identifying the intrusion of underground coal mine personnel shown in the third embodiment of the present invention.

[0095] As Figure 3 shown, the method for identifying the intrusion of underground coal mine personnel includes the following steps:

[0096] Step 301: Obtain the current frame image collected by the image acquisition device arranged underground in the coal mine, and use the monocular depth estimation method to determine the depth image corresponding to the current frame image.

[0097] Step 302: Detect pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and perform bone key point extraction based on the current frame image in the area where each detection frame is located to obtain the two-dimensional coordinates of the bone key points of the corresponding pedestrian targets in each detection frame.

[0098] Step 303: Determine the three-dimensional coordinates of the bone key points of each pedestrian target according to the two-dimensional coordinates of the bone key points of the corresponding pedestrian targets in each detection frame, the depth image, and the internal parameter matrix of the image acquisition device.

[0099] Step 304: For any pedestrian target, determine the target motion confirmation index corresponding to the pedestrian target according to the detection frame corresponding to the pedestrian target in at least one frame image before the current frame image and the detection frame corresponding to the pedestrian target in the current frame image.

[0100] The results of pedestrian detection may have false detections, detecting some areas that seem to be pedestrians as pedestrians. In the underground coal mine environment, these false detection results are often stationary objects in the environment. Therefore, in order to remove these false detection results, the present invention also performs target motion confirmation after pedestrian tracking.

[0101] As a possible implementation, the detection frame is composed of the two-dimensional coordinates of the left side, upper side, right side, and lower side of the matrix frame that minimally encloses the corresponding pedestrian target. Thus, the variances of each left side, the variances of each upper side, the variances of each right side, and the variances of each lower side can be determined according to the two-dimensional coordinates of the left side, upper side, right side, and lower side of the detection frame corresponding to the candidate pedestrian target in at least one frame image before the current frame image and the two-dimensional coordinates of the left side, upper side, right side, and lower side of the detection frame corresponding to the candidate pedestrian target in the current frame image; determine the first target variance with the largest variance from the variances of each left side and the variances of each upper side, and determine the second target variance with the largest variance from the variances of each right side and the variances of each lower side; take the minimum value of the first target variance and the second target variance as the target motion confirmation index.

[0102] For example, let the quadruple of all detection results arranged in chronological order from when the pedestrian tracking target is detected to the latest tracking result be (l i , t i , r i , b i ), which correspond to the left side, upper side, right side, and lower side of this rectangular frame respectively, where i represents the i-th detection result. Use the target motion confirmation index m to confirm the motion amplitude of the detection result:

[0103] m = min(max(var i (l i ), var i (t i ))), max(var i (r i ), var i (b i )))

[0104] Among them, max is to take the maximum value, min is to take the minimum value, and var is to calculate the variance.

[0105] Step 305, when the target motion confirmation index is not greater than the target motion confirmation threshold, determine that the pedestrian target is a non - pedestrian target, and remove the pedestrian target from all pedestrian targets.

[0106] In the embodiment of the present invention, when the target motion confirmation index corresponding to the pedestrian target is not greater than the target motion confirmation threshold, it is considered that the pedestrian target is a non - pedestrian target. At this time, it is necessary to remove the pedestrian target from all pedestrian targets. On the contrary, when the target motion confirmation index corresponding to the pedestrian target is greater than the target motion confirmation threshold, it is considered that the pedestrian target is a pedestrian target. At this time, there is no need to perform a removal operation on the pedestrian target.

[0107] Step 306, determine whether each pedestrian target invades the target space electronic fence area according to the three - dimensional coordinates of the detection frames and / or the skeletal key points of each pedestrian target.

[0108] It should be noted that the execution processes of steps 301 to 303 and step 306 can be implemented in any one of the embodiments of the present invention respectively. The embodiments of the present invention do not make any limitations in this regard and will not be elaborated further.

[0109] The method for identifying the intrusion of underground coal mine personnel in the embodiment of the present invention determines the target motion confirmation index corresponding to the pedestrian target by aiming at any pedestrian target according to the detection frame corresponding to the pedestrian target in at least one frame of image before the current frame image and the detection frame corresponding to the pedestrian target in the current frame image; when the target motion confirmation index is not greater than the target motion confirmation threshold, determine that the pedestrian target is a non - pedestrian target, and remove the pedestrian target from all pedestrian targets. Thus, by using the four corner points of the minimum bounding rectangle to confirm the moving target, some mis - detected targets can be eliminated, the accuracy of pedestrian detection in underground coal mines can be improved, and the false alarm in the challenging environment of underground coal mines can be reduced.

[0110] It should be noted that in the present invention, an alarm can also be given when any pedestrian target invades the target space electronic fence. The following combines Figure 4 , to illustrate this process.

[0111] Figure 4 It is a schematic flowchart of the method for identifying the intrusion of underground coal mine personnel shown in the fourth embodiment of the present invention.

[0112] As Figure 4 shown, the method for identifying the intrusion of underground coal mine personnel includes the following steps:

[0113] Step 401: Obtain the current frame image collected by the image acquisition device arranged underground in the coal mine, and use the monocular depth estimation method to determine the depth image corresponding to the current frame image.

[0114] Step 402: Detect the pedestrian targets in the current frame image to obtain the detection frames corresponding to at least one pedestrian target in the current frame image, and extract the skeletal key points based on the current frame image in the area where each detection frame is located to obtain the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame.

[0115] Step 403: Determine the three-dimensional coordinates of the skeletal key points of each pedestrian target according to the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame, the depth image, and the internal parameter matrix of the image acquisition device.

[0116] Step 404: Determine whether each pedestrian target invades the target space electronic fence area according to the three-dimensional coordinates of each detection frame and / or the skeletal key points of each pedestrian target.

[0117] Step 405: Alarm when any pedestrian target invades the target space electronic fence.

[0118] In the embodiment of the present invention, an alarm can be given when any pedestrian target invades the target space electronic fence.

[0119] It should be noted that false alarms are inevitable in visual detection, and jitter in detection results is also inevitable. To make the alarm smoother, the present invention filters the intrusion results before giving an alarm.

[0120] As a possible implementation, exponential moving average can be used to determine the alarm signal at the current moment according to the alarm signal at the previous moment and the intrusion signal at the current moment; wherein, the alarm signal is used to indicate whether to give an alarm, and the intrusion signal is used to indicate whether someone invades the target space electronic fence.

[0121] For example, let a i be the alarm signal at the i-th moment, indicating whether an alarm is needed, be the intrusion signal at the i-th moment, indicating whether someone invades the electronic fence area, then

[0122]

[0123] Among them, η ranges from 0 < η < 1 and is the coefficient of exponential moving average.

[0124] It should be noted that the execution processes of steps 401 to 404 can be implemented in any one of the embodiments of the present invention. The embodiments of the present invention do not make any limitations in this regard and will not be elaborated further.

[0125] The method for identifying the intrusion of underground coal mine personnel according to the embodiment of the present invention alarms when any pedestrian target intrudes into the target space electronic fence, and before the alarm, exponential moving average is used to filter the intrusion result to avoid false alarms and reduce the false alarm rate.

[0126] To clearly illustrate the above embodiments, examples are given below for illustration.

[0127] Figure 5 It is a schematic flowchart of the method for identifying the intrusion of underground coal mine personnel shown in the fifth embodiment of the present invention.

[0128] Figure 6 It is a schematic diagram of the principle of the method for identifying the intrusion of underground coal mine personnel provided by the sixth embodiment of the present invention.

[0129] Such as Figure 5 and Figure 6 As shown, the present invention designs a simple and effective method for identifying the intrusion of underground coal mine personnel for the closed, narrow, dim, and fuzzy environment in underground coal mines. The present invention uses a pedestrian detection network optimized by an underground coal mine dataset to detect pedestrians in underground coal mine images, and uses monocular depth estimation and human skeletons to track the three-dimensional positions of the detected targets, then eliminates false targets through motion detection, confirms intrusion through multiple points, and alarms for the intrusion target.

[0130] Based on any one of the embodiments of the present invention, as Figure 5 and Figure 6 shown, the method for identifying the intrusion of underground coal mine personnel according to the embodiment of the present invention can also be implemented based on the following steps:

[0131] Step 1 Monocular depth estimation

[0132] To better track the position of pedestrians, the present invention will use the monocular depth estimation method to estimate the depth of the current frame image, so as to obtain the approximate depth of the current frame. Input the current frame into the monocular depth estimation network to obtain the depth image of the current frame, from which the depth information d of each pixel point p can be obtained p .

[0133] Using the internal parameter matrix K of the image acquisition device, any pixel point p can be restored to the three-dimensional space coordinates

[0134] x = dp K -1 p

[0135] Among them, the coordinates of the pixel point p are given in the form of homogeneous coordinates.

[0136] Step 2 Spatial electronic fence area extraction

[0137] The present invention requires an ROI to be input as the area to be determined. The ROI is a closed figure composed of an ordered two-dimensional point sequence, and the area on the right hand side along the order of the point sequence is the inner area of the closed figure. Let the ordered two-dimensional point sequence be p (ROI),i , i = 1, 2,..., L (ROI) The length of which is L (ROI) , for any pixel point p in the current frame image, if the pixel point p is adjacent to the points in any ordered two-dimensional point sequence, that is, i < L (ROI) , j = i + 1 and i = L (ROI) , j = 1, there is: the pixel point p is on the right hand side of the directed line segment formed by the pixel point p (ROI),i and the pixel point p (ROI),j , then it can be determined that the pixel point p is located inside the ROI.

[0138] The ROI preset by the user in the current frame image is a planar area in the current frame image. In order to achieve the tracking effect in space, the present invention will utilize monocular depth estimation, extract the corresponding area in space according to the input ROI area, and perform electronic fence intrusion detection on the area in space.

[0139] Specifically, first, the depth of the pixels inside the input ROI in the image is extracted and restored to three-dimensional space coordinates; then, the Hough transform is used to extract all planes and their normal vectors, and the number of inliers of each plane is calculated as the area occupied by the plane in the ROI area; then, the plane with the largest area is selected as the ground plane P, and its normal vector n P is obtained, as well as the intersection point o of the straight line along the normal vector direction of the ground plane from the coordinate origin and the ground plane; finally, the three-dimensional coordinates x s,(ROI),i of the ordered two-dimensional point sequence constituting the ROI are projected onto the ground plane:

[0140] x g,(ROI),i = x s,(ROI),i - (x g,(ROI),i - o)n P

[0141] The ROI on the ground plane is obtained, which is used as the basis for pedestrian electronic fence intrusion detection, that is, the target spatial electronic fence.

[0142] Step 3 Pedestrian detection

[0143] Accurate pedestrian detection is a prerequisite for pedestrian tracking and intrusion detection. The present invention uses a coal mine underground pedestrian dataset to train a deep neural network, and then infers the input image to achieve the detection of pedestrian targets in the visible light images of coal mine underground. The image collected by the visible light camera in the coal mine underground is input into the pedestrian detection network to obtain the position of the pedestrian target. The target position is output in the form of multiple quadruples. The four numbers included represent the positions of the left, upper, right, and lower sides of the minimum bounding rectangle of the target, denoted as l, t, r, and b respectively. Then the upper left corner point of the rectangle can be expressed as p (lt) =[l,t] T , and the lower right corner point can be expressed as p (rb) =[r,b] T , and the center point can be expressed as p (center) =[(l + r) / 2, (t + b) / 2] T .

[0144] Step 4 Skeleton key point extraction

[0145] Skeleton key point extraction is usually used for pedestrian pose analysis. To achieve better pedestrian tracking, the present invention will extract skeleton key points for pedestrians. The image of the area of the pedestrian bounding rectangle obtained through pedestrian detection is cropped out and input into the skeleton key point extraction network, and the skeleton key points of the pedestrian can be obtained, obtaining the positions of 17 human skeleton key points including the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle, which are p (nose) , p (leye) , p (reye) , p (lear) , p (rear) , p (lshoulder) , p (rshoulder) , p (lelbow) , p (relbow) , p (lwrist) , p (rwrist) , p (lhip) , p (rhip) , p (lknee) , p (rknee) , p (lankle) , p (rankle) .

[0146] Step 5 Kalman filtering

[0147] To achieve the tracking and filtering of the pedestrian position, the present invention will use the Kalman filtering method to track and filter the pedestrian skeleton key points. The present invention uses the pedestrian skeleton key points as the state vector, assuming that the second-order difference of the point to be tracked is 0 and maintaining the first-order difference. Let x i be the coordinates of the three-dimensional point to be tracked at the current time i, then its second-order motion equation is

[0148]

[0149] Among them, x i-1 is the three-dimensional coordinate at the previous moment, i.e., the (i - 1)-th moment, and x i-2 is the three-dimensional coordinate at the (i - 2)-th moment.

[0150]

[0151] is the state transition matrix, indicating that the second-order difference of the point to be tracked is 0, maintaining the first-order difference, and I is the identity matrix.

[0152] In addition to tracking and filtering the key points of the human body skeleton, the present invention also uses Kalman filtering to track the minimum bounding rectangle obtained by pedestrian detection. Specifically, the upper left corner point and the lower right corner point of the minimum bounding rectangle for tracking are used to form a state vector with their two-dimensional coordinates, and the assumption of its motion equation is the same as that of the motion equation of the key points of the skeleton, that is, the second-order difference of the point to be tracked is 0 for pedestrian tracking.

[0153] Step 6 Pedestrian tracking

[0154] After obtaining the detection box using the detection network, the present invention uses the Hungarian matching algorithm to determine the matching degree between each target in the previous frame and each target in the current frame, and matches the pedestrian targets between the two frames. When matching targets, the present invention uses the Euclidean distance of the key points of the skeleton and the intersection over union of the minimum bounding rectangle of the target as the measurement criteria.

[0155] First, use the second-order motion equation to predict the minimum bounding rectangle and the key points of the skeleton of the target in the previous frame, predict their positions in the current frame, and obtain the minimum bounding rectangle P i-1 of the target in the previous frame and the predicted value of the key points of the skeleton x i-1 in the current frame. .

[0156] Then, using the Euclidean distance of the key points of the skeleton as the determination criterion, use the Hungarian matching algorithm for matching. Specifically, use the average Euclidean distance l1 of the points with confidence c i higher than the confidence threshold T c among the 9 pairs of key points of the human body skeleton such as the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left hip, and right hip in the two targets to be judged for matching to determine whether to match:

[0157]

[0158] Among them, w ∈ (nose, leye, reye, lear, rear, lshoulder, rshoulder, lhip, rhip).

[0159] If the confidence level c i is higher than the confidence level threshold T c and the number of points is less than 2, then the intersection-over-union of the minimum bounding rectangles of the two targets to be judged for matching is used to measure the matching degree to determine whether they match.

[0160] After obtaining the matched detection target, the state vector and covariance matrix of the Kalman filter will be updated.

[0161] During tracking, tracking and exit tracking thresholds are also set. The target is confirmed to be tracked only when it is tracked in multiple consecutive frames, and the target tracking is confirmed to be lost only when the target is lost in multiple consecutive frames, preventing the fluctuations in the target detection results within a short period of time from affecting the tracking of the target.

[0162] Step 7 Target motion confirmation

[0163] The results of pedestrian detection may have false detections, detecting some areas that seem to be pedestrians as pedestrians. In the underground environment, these incorrect detection results are often stationary objects in the environment. Therefore, in order to remove these incorrect detection results, after pedestrian tracking, the present invention also performs target motion confirmation, using the changes of the four numbers in the detection results during the tracking process as the confirmation index.

[0164] Let the quadruple of all the detection results arranged in chronological order between the pedestrian tracking target from being detected to the latest tracking result be (l i , t i , r i , b i ), corresponding to the left side, upper side, right side, and lower side of the rectangular box respectively, where i represents the i-th detection result. The motion amplitude of the detection result is confirmed using the target motion confirmation index m:

[0165] m = min(max(var i (l i ), var i (t i )), max(var i (r i ), var i (b i )))

[0166] where max is to take the maximum value, min is to take the minimum value, and var is to calculate the variance. When m is greater than the target motion confirmation threshold, the tracked target is confirmed as a pedestrian target.

[0167] Step 8 Intrusion confirmation

[0168] After obtaining the detection box of the pedestrian target, it is necessary to determine whether the target has invaded. The present invention uses the foot position as a reference to determine whether a pedestrian has invaded the electronic fence area. For the confidence levels c of the left and right ankle bone key points i higher than the confidence threshold T c in the case, the present invention will use the ankle bone key point with a higher confidence level as the foot position, otherwise assume that the midpoint of the lower edge of the pedestrian target detection box is the foot position.

[0169] The present invention will project the three-dimensional coordinates of the foot onto the ground plane to further determine whether the pedestrian's foot is within the electronic fence area. If it is within the electronic fence, it is considered that there is an invasion situation.

[0170] Step 9 Alarm filtering

[0171] Visual detection will inevitably have false alarm situations and detection result jitter situations. In order to make the alarm smoother, the present invention will also filter the invasion results before the alarm. The present invention uses the Exponential Moving Average (EMA) to filter the invasion results.

[0172] Specifically, let a i be the alarm signal at time i, indicating whether an alarm is needed, and a i be the invasion signal at time i, indicating whether someone has invaded the electronic fence area, then

[0173]

[0174] where η ranges from 0 < η < 1 and is the coefficient of the exponential moving average.

[0175] In summary, the position of the electronic fence area in space is extracted to achieve more accurate electronic fence intrusion detection; the three-dimensional coordinates of the bone key points are restored using monocular depth estimation, and then the position of the pedestrian is tracked and filtered using Kalman filtering; on the basis of Hungarian matching, the corner points of the minimum bounding rectangle and the bone points are used as the matching judgment basis for the Hungarian matching algorithm to achieve target tracking; the four corner points of the minimum bounding rectangle are used to confirm the moving target and eliminate some misdetected targets. Thus, for the problem that the pedestrian detection algorithm has a large number of false alarms in the narrow and dim environment of the coal mine underground, resulting in more false alarms in the electronic fence intrusion detection, the present invention optimizes the pre-trained pedestrian detection model using the coal mine underground dataset, then restores the pedestrian position to three-dimensional coordinates using monocular depth estimation and performs tracking, and then applies moving target confirmation and filters the alarm results to determine the target invading the electronic fence area of the space, improving the accuracy of pedestrian detection in the coal mine underground and reducing the false alarms in the challenging environment of the coal mine underground.

[0176] Corresponding to the method for identifying the intrusion of underground coal mine personnel provided in the above embodiment, the present invention also provides a device for identifying the intrusion of underground coal mine personnel. Since the device for identifying the intrusion of underground coal mine personnel provided in the embodiment of the present invention corresponds to the method for identifying the intrusion of underground coal mine personnel provided in the above embodiment, the implementation manner of the method for identifying the intrusion of underground coal mine personnel is also applicable to the device for identifying the intrusion of underground coal mine personnel provided in the embodiment of the present invention, and will not be described in detail in the embodiment of the present invention.

[0177] Figure 7 It is a schematic structural diagram of the device for identifying the intrusion of underground coal mine personnel shown in the seventh embodiment of the present invention.

[0178] As Figure 7 shown, the device 700 for identifying the intrusion of underground coal mine personnel includes: a first processing module 710, a second processing module 720, a first determination module 730, and a second determination module 740.

[0179] Among them, the first processing module 710 is used to obtain the current frame image collected by the image acquisition device arranged underground in the coal mine, and determine the depth image corresponding to the current frame image by using the monocular depth estimation method; the second processing module 720 is used to detect the pedestrian targets in the current frame image to obtain at least one detection box corresponding to the pedestrian targets in the current frame image, and perform skeleton key point extraction based on the current frame image in the area where each detection box is located to obtain the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection box; wherein, the detection box is a matrix box that minimally encloses the corresponding pedestrian target; the first determination module 730 is used to determine the three-dimensional coordinates of the skeleton key points of each pedestrian target according to the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection box, the depth image, and the internal parameter matrix of the image acquisition device; the second determination module 740 is used to determine whether each pedestrian target intrudes into the target space electronic fence area according to each detection box and / or the three-dimensional coordinates of the skeleton key points of each pedestrian target; wherein, the target space electronic fence is determined based on the depth image and the preset area to be determined.

[0180] As a possible implementation of an embodiment of the present invention, the skeletal key points at least include the left and right ankle skeletal key points; the second determination module 740 is further configured to: for any pedestrian target, when the confidence levels of the left and right ankle skeletal key points of the pedestrian target are both greater than the confidence threshold, use the three-dimensional coordinates of the ankle skeletal key point with a higher confidence level among the left and right ankle skeletal key points of the pedestrian target as the foot position of the pedestrian target; when the confidence levels of the left and right ankle skeletal key points of the pedestrian target are not both greater than the confidence threshold, use the three-dimensional coordinates of the midpoint of the lower side of the detection frame corresponding to the pedestrian target as the foot position of the pedestrian target; wherein, the three-dimensional coordinates of the midpoint of the lower side of the detection frame corresponding to the pedestrian target are determined based on the two-dimensional coordinates of the midpoint of the lower side of the detection frame corresponding to the pedestrian target, the depth image, and the internal parameter matrix of the image acquisition device; project the foot position of the pedestrian target onto the ground plane to determine whether the foot position of the pedestrian target is within the target space electronic fence area, and when the foot position of the pedestrian target is within the target space electronic fence area, determine that the pedestrian target invades the target space electronic fence area; wherein, the ground plane is the ground plane determined during the process of determining the target space electronic fence area.

[0181] As a possible implementation of an embodiment of the present invention, the above device further includes: a tracking module, configured to perform tracking on the detection frame of the pedestrian target by using Kalman filtering according to multiple frames of images before the current frame image to obtain at least one predicted detection frame corresponding to a predicted pedestrian target in the current frame image, and / or perform tracking on the skeletal key points of the pedestrian target by using Kalman filtering to obtain the predicted value of the three-dimensional coordinates of the skeletal key points of at least one predicted pedestrian target in the current frame image; a matching module, configured to perform matching on each pedestrian target and each predicted pedestrian target by using the Hungarian matching algorithm according to at least one of each predicted detection frame, the predicted value of the three-dimensional coordinates of the skeletal key points of each predicted pedestrian target, each detection frame, and the three-dimensional coordinates of the skeletal key points of each pedestrian target, so as to remove non-pedestrian targets among each pedestrian target.

[0182] As a possible implementation of an embodiment of the present invention, the detection frame is composed of the two-dimensional coordinates of the left side, the upper side, the right side, and the lower side of the matrix frame that minimally encloses the corresponding pedestrian target; the tracking module is further configured to: use the two-dimensional coordinates of the upper left corner point and the lower right corner point of the detection frame corresponding to at least one pedestrian target in multiple frames of images before the current frame image as the state vector, use the second-order difference of the first point to be tracked as the first preset vector, and keep the first-order difference unchanged as the target to construct the first motion equation:

[0183]

[0184] wherein, x i represents the two-dimensional coordinates of the first point to be tracked in the current frame image, x i-1Denote the two-dimensional coordinates of the first point to be tracked in the previous frame image before the current frame image, x i-2 Denote the two-dimensional coordinates of the first point to be tracked in the frame image before the previous frame image before the current frame image

[0185]

[0186] A x is the first state transition matrix, indicating that the second-order difference of the first point to be tracked is the first preset vector and the first-order difference remains unchanged, where I is the identity matrix; among them, the first point to be tracked is the upper left corner point or the lower right corner point of the detection box corresponding to any pedestrian target in multiple frame images before the current frame image, and the two-dimensional coordinates of the upper left corner point of any detection box are determined based on the two-dimensional coordinates of the left side and the upper side of the detection box, and the two-dimensional coordinates of the lower right corner point of any detection box are determined based on the two-dimensional coordinates of the right side and the lower side of the detection box; according to the first motion equation, determine the predicted detection box corresponding to at least one predicted pedestrian target in the current frame image; and / or, use the three-dimensional coordinates of the skeleton key points of at least one pedestrian target in multiple frame images before the current frame image as the state vector, and use the second-order difference of the second point to be tracked as the second preset vector and the first-order difference remains unchanged as the target to construct the second motion equation:

[0187]

[0188] Among them, x′ i Denote the three-dimensional coordinates of the second point to be tracked in the current frame image, x i-1 ′ denotes the three-dimensional coordinates of the second point to be tracked in the previous frame image before the current frame image, x i-2 ′ denotes the three-dimensional coordinates of the second point to be tracked in the frame image before the previous frame image before the current frame image

[0189]

[0190] A x ′ is the second state transition matrix, indicating that the second-order difference of the second point to be tracked is the second preset vector and the first-order difference remains unchanged, where I is the identity matrix; among them, the second point to be tracked is any skeleton key point of any pedestrian target in multiple frame images before the current frame image; according to the second motion equation, determine the predicted value of the three-dimensional coordinates of the skeleton key points of at least one predicted pedestrian target in the current frame image

[0191] As a possible implementation of an embodiment of the present invention, the matching module is further configured to: for any pair of pedestrian targets to be matched and predicted pedestrian targets, determine the number of skeletal key points with a confidence level higher than the confidence level threshold among the set of skeletal key points of the pedestrian target and the predicted pedestrian target; in the case where the number is not less than the number threshold, use the Euclidean distance between each pair of skeletal key points of the pedestrian target and the predicted pedestrian target as the determination criterion, and use the Hungarian matching algorithm to match each pedestrian target and each predicted pedestrian target to remove non-pedestrian targets from each pedestrian target; in the case where the number is less than the number threshold, use the intersection over union of the detection box corresponding to the pedestrian target and the predicted detection box corresponding to the predicted pedestrian target as the determination criterion, and use the Hungarian matching algorithm to match each pedestrian target and each predicted pedestrian target to remove non-pedestrian targets from each pedestrian target.

[0192] As a possible implementation of an embodiment of the present invention, the above device further includes: a third determination module, configured to, for any pedestrian target, determine a target motion confirmation index corresponding to the pedestrian target according to the detection box corresponding to the pedestrian target in at least one frame of image before the current frame image and the detection box corresponding to the pedestrian target in the current frame image; a third processing module, configured to, in the case where the target motion confirmation index corresponding to the pedestrian target is not greater than the target motion confirmation threshold, determine that the pedestrian target is a non-pedestrian target and remove the pedestrian target from each pedestrian target.

[0193] As a possible implementation of an embodiment of the present invention, the detection box is composed of the two-dimensional coordinates of the left side, upper side, right side, and lower side of the matrix box that minimally encloses the corresponding pedestrian target; the third determination module is further configured to: determine the variances of each left side, the variances of each upper side, the variances of each right side, and the variances of each lower side according to the two-dimensional coordinates of the left side, upper side, right side, and lower side of the detection box corresponding to the candidate pedestrian target in at least one frame of image before the current frame image and the two-dimensional coordinates of the left side, upper side, right side, and lower side of the detection box corresponding to the candidate pedestrian target in the current frame image; determine the first target variance with the largest variance from the variances of each left side and the variances of each upper side, and determine the second target variance with the largest variance from the variances of each right side and the variances of each lower side; use the minimum value of the first target variance and the second target variance as the target motion confirmation index.

[0194] As a possible implementation manner of an embodiment of the present invention, the above device further includes: a fourth determination module, configured to determine the three-dimensional coordinates of each point in the two-dimensional point sequence according to the depth image and the two-dimensional coordinates of each point in the two-dimensional point sequence constituting the area to be determined; a fourth processing module, configured to use the Hough transform to extract multiple planes within the area to be determined, and use the number of inliers of each plane as the area occupied by the corresponding plane in the area to be determined; a fifth processing module, configured to select the plane with the largest area from each plane as the ground plane, and obtain the normal vector of the ground plane and the three-dimensional coordinates of the target intersection point; wherein, the target intersection point is the intersection point of the straight line along the normal vector direction of the ground plane from the coordinate origin and the ground plane; a projection module, configured to project the three-dimensional coordinates of each point in the two-dimensional point sequence onto the ground plane according to the normal vector of the ground plane and the three-dimensional coordinates of the target intersection point, so as to obtain the target space electronic fence.

[0195] As a possible implementation manner of an embodiment of the present invention, the above device further includes: an alarm module, configured to give an alarm when any pedestrian target invades the target space electronic fence.

[0196] As a possible implementation manner of an embodiment of the present invention, the above device further includes: a fifth determination module, configured to use exponential moving average to determine the alarm signal at the current moment according to the alarm signal at the previous moment and the intrusion signal at the current moment; wherein, the alarm signal is used to indicate whether to give an alarm, and the intrusion signal is used to indicate whether someone invades the target space electronic fence.

[0197] The personnel intrusion recognition device for coal mines in the embodiments of the present invention obtains the current frame image collected by an image acquisition device arranged in the coal mine, and determines the depth image corresponding to the current frame image by using the monocular depth estimation method; detects pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and extracts skeletal key points based on the current frame image in the area where each detection frame is located to obtain the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame; wherein, the detection frame is a matrix frame that minimally encloses the corresponding pedestrian target; determines the three-dimensional coordinates of the skeletal key points of each pedestrian target according to the two-dimensional coordinates of the skeletal key points of the corresponding pedestrian targets in each detection frame, the depth image, and the internal parameter matrix of the image acquisition device; determines whether each pedestrian target invades the target space electronic fence area according to the three-dimensional coordinates of each detection frame and / or the skeletal key points of each pedestrian target; wherein, the target space electronic fence is determined based on the depth image and a preset area to be determined. Thus, for the existing pedestrian detection methods, when applied in the dim and blurred environment of coal mines, there are a large number of false alarms, resulting in a large number of false alarms in the electronic fence intrusion detection. The present invention combines monocular depth estimation and human bones to perform three-dimensional position tracking on pedestrian targets in the current frame image collected by the image acquisition device arranged in the coal mine, which can improve the accuracy of pedestrian detection in coal mines and reduce the false alarm rate of electronic fence intrusion detection.

[0198] In an exemplary embodiment, the present invention also provides an electronic device.

[0199] Wherein, the electronic device includes:

[0200] A processor;

[0201] A memory for storing instructions executable by the processor;

[0202] Wherein, the processor is configured to execute instructions to implement the method for recognizing personnel intrusion in coal mines proposed in any of the foregoing embodiments.

[0203] As an example, Figure 8 is a schematic structural diagram of an electronic device 800 shown in an exemplary embodiment of the present invention. As Figure 8 shown, the above-mentioned electronic device 800 may further include:

[0204] A memory 810 and a processor 820, a bus 830 connecting different components (including the memory 810 and the processor 820), and the memory 810 stores a computer program, and when the processor 820 executes the program, it implements the method for recognizing personnel intrusion in coal mines in the embodiments of the present invention.

[0205] The bus 830 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor bus, or a local bus using any of the various bus architectures. By way of example, such architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0206] The electronic device 800 typically includes a variety of computer-readable media. These media can be any available media that can be accessed by the electronic device 800, including both volatile and nonvolatile media, removable and non-removable media.

[0207] The memory 810 may also include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 840 and / or cache memory 850. The server 800 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, the storage system 860 can be used for reading from and writing to non-removable, nonvolatile magnetic media ( Figure 8 not shown and typically called a "hard disk drive enclosure"). Although Figure 8 not shown in the figure, a magnetic disk drive enclosure for reading from and writing to a removable nonvolatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive enclosure for reading from and writing to a removable nonvolatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) can be provided. In these cases, each drive enclosure can be connected to the bus 830 via one or more data media interfaces. The memory 810 may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of the embodiments of the present invention.

[0208] A program / utility 880 having a set (at least one) of program modules 870 can be stored, for example, in the memory 810, and such program modules 870 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a network environment. The program modules 870 typically carry out the functions and / or methods of the embodiments described herein.

[0209] The electronic device 800 can also communicate with one or more external devices 890 (such as a keyboard, a pointing device, a display 891, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 892. Moreover, the electronic device 800 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 893. As shown in the figure, the network adapter 893 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 800, including but not limited to: microcode, device driver packers, redundant processing units, external disk drive pack arrays, RAID systems, tape drive packers, and data backup storage systems, etc.

[0210] The processor 820 executes various functional applications and data processing by running the programs stored in the memory 810.

[0211] It should be noted that for the implementation process and technical principle of the electronic device in this embodiment, refer to the foregoing explanation of the method for identifying the intrusion of underground coal mine personnel in the embodiments of the present invention, and details are not described herein again.

[0212] In an exemplary embodiment, the present invention also provides a computer-readable storage medium including instructions, such as a memory including instructions. The above instructions can be executed by the processor of the electronic device to complete the method for identifying the intrusion of underground coal mine personnel proposed in any of the foregoing embodiments. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0213] In an exemplary embodiment, the present invention also provides a computer program product including a computer program / instructions. When the above computer program / instructions are executed by a processor, the method for identifying the intrusion of underground coal mine personnel proposed in any of the foregoing embodiments is implemented.

[0214] Those skilled in the art will readily think of other implementation schemes of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not disclosed in the present invention. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0215] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for identifying the intrusion of personnel underground in coal mines, characterized in that, Including: Obtain the current frame image collected by an image acquisition device arranged underground in a coal mine, and use a monocular depth estimation method to determine the depth image corresponding to the current frame image; Detect pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and perform skeleton key point extraction based on the current frame image in the area where each detection frame is located to obtain the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection frame; wherein, the detection frame is a matrix frame that minimally encloses the corresponding pedestrian target; Determine the three-dimensional coordinates of the skeleton key points of each pedestrian target according to the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection frame, the depth image, and the internal parameter matrix of the image acquisition device; According to multiple frames of images before the current frame image, use Kalman filtering to track the detection frames of pedestrian targets to obtain prediction detection frames corresponding to at least one predicted pedestrian target in the current frame image, and / or use Kalman filtering to track the skeleton key points of pedestrian targets to obtain three-dimensional coordinate prediction values of the skeleton key points of at least one predicted pedestrian target in the current frame image; According to at least one of each prediction detection frame, the three-dimensional coordinate prediction values of the skeleton key points of each predicted pedestrian target, each detection frame, and the three-dimensional coordinates of the skeleton key points of each pedestrian target, use the Hungarian matching algorithm to match each pedestrian target and each predicted pedestrian target to remove non-pedestrian targets from each pedestrian target; Determine whether each pedestrian target invades the target space electronic fence area according to each detection frame and / or the three-dimensional coordinates of the skeleton key points of each pedestrian target; wherein, the target space electronic fence is determined based on the depth image and a preset area to be determined, and the area to be determined is any region of interest preset in the current frame image; Wherein, the detection frame is composed of the two-dimensional coordinates of the left side, upper side, right side, and lower side of the matrix frame that minimally encloses the corresponding pedestrian target.

2. The method according to claim 1, wherein The skeleton key points at least include the left and right ankle skeleton key points; the determining whether each pedestrian target invades the target space electronic fence area according to each detection frame and / or the three-dimensional coordinates of the skeleton key points of each pedestrian target includes: For any one of the pedestrian targets, when the confidence levels of the left and right ankle skeleton key points of the pedestrian target are both greater than the confidence level threshold, use the three-dimensional coordinate of the ankle skeleton key point with a higher confidence level among the left and right ankle skeleton key points of the pedestrian target as the foot position of the pedestrian target; When the confidence levels of the left and right ankle skeleton key points of the pedestrian target are not both greater than the confidence level threshold, use the three-dimensional coordinate of the midpoint of the lower side of the detection frame corresponding to the pedestrian target as the foot position of the pedestrian target; wherein, the three-dimensional coordinate of the midpoint of the lower side of the detection frame corresponding to the pedestrian target is determined based on the two-dimensional coordinate of the midpoint of the lower side of the detection frame corresponding to the pedestrian target, the depth image, and the internal parameter matrix of the image acquisition device; Project the foot position of the pedestrian target onto the ground plane to determine whether the foot position of the pedestrian target is within the target space electronic fence area, and determine that the pedestrian target invades the target space electronic fence area when the foot position of the pedestrian target is within the target space electronic fence area; wherein, the ground plane is the ground plane determined during the determination of the target space electronic fence area.

3. The method according to claim 1, wherein Tracking the detection box of the pedestrian target by using Kalman filtering based on multiple frames of images before the current frame image to obtain the predicted detection box corresponding to at least one predicted pedestrian target in the current frame image, and / or tracking the skeletal key points of the pedestrian target by using Kalman filtering to obtain the three-dimensional coordinate prediction values of the skeletal key points of at least one predicted pedestrian target in the current frame image, including: Taking the two-dimensional coordinates of the upper left corner point and the lower right corner point of the detection box corresponding to at least one pedestrian target in multiple frames of images before the current frame image as the state vector, taking the second-order difference of the first point to be tracked as the first preset vector, and keeping the first-order difference unchanged as the target, to construct the first motion equation: where x i represents the two-dimensional coordinates of the first point to be tracked in the current frame image, and x i-1 represents the two-dimensional coordinates of the first point to be tracked in the previous frame image before the current frame image, and x i-2 represents the two-dimensional coordinates of the first point to be tracked in the frame image before the previous frame image before the current frame image A x is the first state transition matrix, indicating that the second-order difference of the first point to be tracked is the first preset vector and the first-order difference remains unchanged, where I is the identity matrix; wherein, the first point to be tracked is the upper left corner point or the lower right corner point of the detection frame corresponding to any pedestrian target in multiple frames of images before the current frame image, and the two-dimensional coordinates of the upper left corner point of any detection frame are determined based on the two-dimensional coordinates of the left side and the upper side of the detection frame, and the two-dimensional coordinates of the lower right corner point of any detection frame are determined based on the two-dimensional coordinates of the right side and the lower side of the detection frame; Determining the predicted detection box corresponding to at least one predicted pedestrian target in the current frame image according to the first motion equation; and / or, Taking the three-dimensional coordinates of the skeletal key points of at least one pedestrian target in multiple frames of images before the current frame image as the state vector, taking the second-order difference of the second point to be tracked as the second preset vector, and keeping the first-order difference unchanged as the target, to construct the second motion equation: where x i ′ represents the three-dimensional coordinates of the second point to be tracked in the current frame image, x i-1 ′ represents the three-dimensional coordinates of the second point to be tracked in the previous frame image before the current frame image, x i-2 ′ represents the three-dimensional coordinates of the second point to be tracked in the frame image before the previous frame image before the current frame image A x ′ is the second state transition matrix, indicating that the second-order difference of the second point to be tracked is the second preset vector and the first-order difference remains unchanged, where I is the identity matrix; wherein, the second point to be tracked is any bone key point of any pedestrian target in multiple frames of images before the current frame image. Determining the three-dimensional coordinate prediction values of the skeletal key points of at least one predicted pedestrian target in the current frame image according to the second motion equation.

4. The method according to claim 1, characterized in that, Matching each pedestrian target and each predicted pedestrian target by using the Hungarian matching algorithm according to at least one of each predicted detection box, the three-dimensional coordinate prediction values of the skeletal key points of each predicted pedestrian target, each detection box, and the three-dimensional coordinates of the skeletal key points of each pedestrian target, so as to remove non-pedestrian targets among each pedestrian target, including: For any pair of the pedestrian target and the predicted pedestrian target to be matched, determining the number of skeletal key points with a confidence level higher than the confidence level threshold among the set pairs of skeletal key points of the pedestrian target and the predicted pedestrian target; When the number is not less than the number threshold, using the Euclidean distance of each pair of skeletal key points of the pedestrian target and the predicted pedestrian target as the determination criterion, and using the Hungarian matching algorithm to match each pedestrian target and each predicted pedestrian target, so as to remove non-pedestrian targets among each pedestrian target; When the number is less than the number threshold, using the intersection over union of the detection box corresponding to the pedestrian target and the predicted detection box corresponding to the predicted pedestrian target as the determination criterion, and using the Hungarian matching algorithm to match each pedestrian target and each predicted pedestrian target, so as to remove non-pedestrian targets among each pedestrian target.

5. The method according to claim 1, characterized in that Before determining whether each pedestrian target invades the target space electronic fence area according to the three-dimensional coordinates of each detection box and / or the skeletal key points of each pedestrian target, the method further includes: For any one of the pedestrian targets, determining a target motion confirmation index corresponding to the pedestrian target according to the detection box corresponding to the pedestrian target in at least one frame of image before the current frame image and the detection box corresponding to the pedestrian target in the current frame image; When the target motion confirmation index corresponding to the pedestrian target is not greater than the target motion confirmation threshold, determining that the pedestrian target is a non-pedestrian target and removing the pedestrian target from each of the pedestrian targets.

6. The method according to claim 5, wherein The detection box is composed of the two-dimensional coordinates of the left side, the upper side, the right side, and the lower side of the matrix box that minimally encloses the corresponding pedestrian target; determining the target motion confirmation index corresponding to the pedestrian target according to the detection box corresponding to the pedestrian target in at least one frame of image before the current frame image and the detection box corresponding to the pedestrian target in the current frame image includes: Determining the variances of each of the left sides, the variances of each of the upper sides, the variances of each of the right sides, and the variances of each of the lower sides according to the two-dimensional coordinates of the left side, the upper side, the right side, and the lower side of the detection box corresponding to the candidate pedestrian target in at least one frame of image before the current frame image and the two-dimensional coordinates of the left side, the upper side, the right side, and the lower side of the detection box corresponding to the candidate pedestrian target in the current frame image; Determining a first target variance with the largest variance from the variances of each of the left sides and the variances of each of the upper sides, and determining a second target variance with the largest variance from the variances of each of the right sides and the variances of each of the lower sides; Taking the minimum value of the first target variance and the second target variance as the target motion confirmation index.

7. The method according to claim 1, characterized in that The determination process of the target space electronic fence includes: Determining the three-dimensional coordinates of each point in the two-dimensional point sequence according to the depth image and the two-dimensional coordinates of each point in the two-dimensional point sequence constituting the area to be determined; Using the Hough transform to extract multiple planes within the area to be determined, and taking the number of inliers of each plane as the area occupied by the corresponding plane in the area to be determined; Selecting the plane with the largest area from each of the planes as the ground plane, and obtaining the normal vector of the ground plane and the three-dimensional coordinates of the target intersection point; wherein, the target intersection point is the intersection point of the straight line along the normal vector direction of the ground plane from the coordinate origin and the ground plane; Projecting the three-dimensional coordinates of each point in the two-dimensional point sequence onto the ground plane according to the normal vector of the ground plane and the three-dimensional coordinates of the target intersection point to obtain the target space electronic fence.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: When any one of the pedestrian targets invades the target space electronic fence, an alarm is issued.

9. The method according to claim 8, wherein Before issuing the alarm, the method further includes: Using exponential moving average to determine the alarm signal at the current moment according to the alarm signal at the previous moment and the intrusion signal at the current moment; wherein, the alarm signal is used to indicate whether to issue an alarm, and the intrusion signal is used to indicate whether someone invades the target space electronic fence.

10. An underground coal mine personnel intrusion recognition device, characterized in that, including: The first processing module is used to obtain the current frame image collected by the image acquisition device arranged underground in the coal mine, and determine the depth image corresponding to the current frame image by using the monocular depth estimation method; The second processing module is used to detect pedestrian targets in the current frame image to obtain detection frames corresponding to at least one pedestrian target in the current frame image, and perform skeleton key point extraction based on the current frame image in the area where each detection frame is located to obtain the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection frame; wherein, the detection frame is a matrix frame that minimally encloses the corresponding pedestrian target; The first determination module is used to determine the three-dimensional coordinates of the skeleton key points of each pedestrian target according to the two-dimensional coordinates of the skeleton key points of the corresponding pedestrian targets in each detection frame, the depth image, and the internal parameter matrix of the image acquisition device; The tracking module is used to track the detection frames of pedestrian targets by using Kalman filtering according to multiple frames of images before the current frame image to obtain prediction detection frames corresponding to at least one predicted pedestrian target in the current frame image, and / or track the skeleton key points of pedestrian targets by using Kalman filtering to obtain three-dimensional coordinate prediction values of the skeleton key points of at least one predicted pedestrian target in the current frame image; according to at least one of each prediction detection frame, the three-dimensional coordinate prediction values of the skeleton key points of each predicted pedestrian target, each detection frame, and the three-dimensional coordinates of the skeleton key points of each pedestrian target, use the Hungarian matching algorithm to match each pedestrian target and each predicted pedestrian target to remove non-pedestrian targets among each pedestrian target; The second determination module is used to determine whether each pedestrian target invades the target space electronic fence area according to the three-dimensional coordinates of each detection frame and / or the skeleton key points of each pedestrian target; wherein, the target space electronic fence is determined based on the depth image and a preset area to be determined, and the area to be determined is any region of interest preset in the current frame image; Wherein, the detection frame is composed of the two-dimensional coordinates of the left side, upper side, right side, and lower side of the matrix frame that minimally encloses the corresponding pedestrian target.

Citation Information

Patent Citations

  • Human body action recognition method and device, storage medium and vehicle

    CN115578720A

  • Method for evaluating risk of worker invading construction dangerous area based on computer vision

    CN117079185A