An automatic identification system for people and vehicles underground
Image data collected by cameras combined with K-means++ and YOLO algorithms, combined with wind speed, air pressure, gas concentration and temperature data, to identify the behavior of people and vehicles in the mine, solving the accurate identification problem of automatic damper system in the mine, and reducing the risk of safety accidents.
Patent Information
- Application Number
- CN202311066498.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-23
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2043-08-23
AI Technical Summary
It is difficult to accurately identify people and vehicles in complex environments with low visibility, less visible light and lots of dust, resulting in safety accidents and resource waste risks.
Image data is collected by cameras, combined with K-means++ algorithm and YOLO algorithm for target detection, and using wind speed, air pressure, gas concentration and temperature data, the human and vehicle behavior in the mine is identified and analyzed through the data processing center, and the start and closing of the damper is controlled.
It realizes accurate identification and detection of people and vehicles in underground environments, timely judges abnormal dangers, reduces the risk of safety accidents, and improves identification efficiency and accuracy.
Smart Images

Figure CN117037070B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underground human and vehicle identification, and in particular to an underground human and vehicle automatic identification system. Background Art
[0002] The automatic dampers currently used in various mines usually use sensors in the damper control system to detect the activities of people and vehicles to automatically open or close the dampers.
[0003] However, using sensors to identify people and vehicles in mines has many shortcomings. For example, it is difficult to accurately identify targets in the complex environment of low visibility, little visible light, and a lot of dust underground, which can easily lead to safety accidents.
[0004] Moreover, in the highly dangerous environment of a mine, the damper control system's control of the damper will constantly affect the changes in gas and air pressure in the mine. If the damper fails to recognize a vehicle or a person, resulting in a failure to open the damper, it will waste resources and create the risk of dangerous accidents. For this reason, an automatic identification system for vehicles and people underground is proposed. Summary of the Invention
[0005] In view of the deficiencies in the prior art, the present invention aims to provide an automatic identification system for people and vehicles underground to solve the problems raised in the above background technology.
[0006] According to one aspect of the present application, an automatic identification system for underground vehicles and people is provided, wherein the identification method of the automatic identification system comprises the following steps:
[0007] S1. The camera collects an image dataset of vehicle and pedestrian features. Computer vision technology uses image processing and image detection technology to process the image dataset, obtain the parameters of the initial detection frame, and send it to the data processing center. The sensor on the damper collects data on wind speed, air pressure, gas concentration, and temperature around the damper, and sends it to the data processing center.
[0008] S2. The data processing center uses the K-means++ algorithm to cluster the processed image dataset to obtain the parameters of the initial clustering frame. Based on the K-means clustering algorithm, the initial detection frame parameters and the initial clustering frame parameters are clustered to obtain the size of the prior frame. Then, based on the YOLO algorithm, the size of the prior frame obtained after clustering is used for model training to obtain the target detection model data.
[0009] In step S2, a sample is randomly selected from the data set as the initial cluster center. The selection of the initial cluster center in the K-means++ algorithm is as follows:
[0010] A1. Randomly extract a sample from the image dataset as the initial cluster center C i ;
[0011] A2. First, calculate the shortest distance D(x) between each sample and the current cluster center;
[0012] Calculate the probability of each sample being the next cluster center Finally, the next cluster center is selected according to the roulette method;
[0013] A3. Repeat step 2 until K cluster centers are selected.
[0014] In step S2, the data processing center uses the target detection model data based on the YOLO algorithm to perform target detection on the image dataset that collects vehicle and pedestrian features, obtains vehicle and pedestrian information data, and uses the MSRCR algorithm to enhance the image dataset;
[0015] The MSRCR algorithm is as follows:
[0016]
[0017] Where, log is the logarithm of 10, C i is the color restoration factor of the i-th channel, which is used to adjust the color ratio of the three channels in the RGB color channel. i (x, y) is the image of the i-th channel, f[*] is the mapping function of the color space, β is the gain constant, α is the controlled nonlinear intensity, R MSRi (x, y) is the output image of the MSR processing of the i-th channel, R MSRCRi (x, y) is the output image of the MSRCR processing of the i-th channel;
[0018] S3. The data processing center processes the collected data on wind speed, air pressure, gas concentration, and temperature around the damper, and combines it with the acquired target monitoring model data to identify and detect the behavior of vehicles and pedestrians in the mine, ultimately identifying the vehicle or behavior.
[0019] S4. The data processing center is based on a time synchronization protocol and analyzes the behavior of vehicles and pedestrians to determine whether there are abnormal dangers for vehicles and pedestrians, and controls the opening and closing of the dampers.
[0020] Preferably, the iterative update of cluster centers in the K-means algorithm is as follows:
[0021] C i is the cluster center, X i for each sample in the image dataset.
[0022] Preferably, the K-means clustering algorithm predicts objects at three different scales of the prior frame size, wherein three target bounding boxes are predicted for each size.
[0023] Preferably, the sizes of the a priori frame include 13×13, 26×26, and 52×52.
[0024] Preferably, the K-means++ algorithm uses the intersection-over-union (IOU) as the distance formula, which is as follows:
[0025] box tran is the prior target box, box truth is the true target box.
[0026] Preferably, the target detection model data includes a loss function, which consists of three parts: coordinate error of the bounding box, confidence error of the bounding box, and classification error.
[0027] Preferably, the algorithm of the loss function is as follows:
[0028] loss=coordError+iouError+classError
[0029]
[0030] Among them, coordError is the coordinate error, iouError is the confidence error, classError is the classification error, λ coord is the penalty weight coefficient of the coordinate error, set to 5, K×K is the number of grids divided into images, and B is the number of bounding boxes corresponding to each grid. is the horizontal coordinate of the center point of the predicted i-th bounding box, is the ordinate of the center point of the predicted i-th bounding box, is the width of the predicted i-th bounding box, is the height of the predicted i-th bounding box, x i is the horizontal coordinate of the center point of the true i-th bounding box, y i is the true vertical coordinate of the center point of the i-th bounding box, ω i is the actual width of the i-th bounding box, h i is the height of the true i-th bounding box, λ noobj is the penalty weight coefficient of confidence error, set to 0.5, is the confidence of the predicted target in the i-th grid, C i is the confidence that the target exists in the true i-th grid, is the confidence of the predicted i-th grid target category, Di is the true confidence of the i-th grid target category, is the predicted category probability of the i-th grid, p i (c) is the true category probability of the i-th grid.
[0031] Python-based image processing of object detection model data.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] The present invention uses a camera to collect image data and obtain target monitoring model data to achieve the purpose of identifying people and vehicles. At the same time, by combining data on wind speed, air pressure, gas concentration and ambient temperature, it can comprehensively analyze the behavior of people and vehicles in the mine, achieve accurate identification and detection, and be able to promptly determine whether there is an abnormal danger, and ensure the safety of the mine environment by controlling the start and close of the damper. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Flowchart of a calculation method of an underground person-vehicle automatic identification system according to an embodiment of the present invention;
[0035] Figure 2 A network structure diagram of Darknet-53 of a calculation method for an underground person-vehicle automatic identification system according to an embodiment of the present invention;
[0036] Figure 3 2. A structural diagram of a residual component of a calculation method of an underground person-vehicle automatic identification system according to an embodiment of the present invention;
[0037] Figure 4 is a graph showing an average intersection-over-union ratio curve of a calculation method of an underground vehicle and person automatic identification system according to an embodiment of the present invention;
[0038] Figure 5 3. It is a curve diagram of the change of the average loss function of the calculation method of the underground person-vehicle automatic identification system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0040] See also Figure 1 - Figure 5As shown, this embodiment provides an automatic identification system for underground people and vehicles, and the identification method of the automatic identification system includes the following steps:
[0041] S1. The camera collects an image dataset of vehicle and pedestrian features. Computer vision technology uses image processing and image detection technology to process the image dataset, obtain the parameters of the initial detection frame, and send it to the data processing center. The sensor on the damper collects data on wind speed, air pressure, gas concentration, and temperature around the damper, and sends it to the data processing center.
[0042] S2. The data processing center uses the K-means++ algorithm to cluster the processed image dataset to obtain the parameters of the initial clustering frame. Based on the K-means clustering algorithm, the initial detection frame parameters and the initial clustering frame parameters are clustered to obtain the size of the prior frame. Then, based on the YOLO algorithm, the size of the prior frame obtained after clustering is used for model training to obtain the target detection model data.
[0043] S3. The data processing center processes the collected data on wind speed, air pressure, gas concentration, and temperature around the damper, and combines it with the acquired target monitoring model data to identify and detect the behavior of vehicles and pedestrians in the mine, ultimately identifying the vehicle or behavior.
[0044] S4. The data processing center is based on a time synchronization protocol and analyzes the behavior of vehicles and pedestrians to determine whether there are abnormal dangers for vehicles and pedestrians, and controls the opening and closing of the dampers.
[0045] In the above technical solution, machine learning is achieved by model training through the K-means++ algorithm in S2 and the K-means clustering algorithm in S3, as well as by utilizing the size of the prior frame obtained after clustering. Machine learning is used to process large image data sets, cluster image data sets and obtain the size of the prior frame, thereby training a target detection model and further achieving target detection and recognition of vehicles and pedestrians.
[0046] In the above technical solution, the YOLO algorithm further adopts the Darknet-53 network structure as the backbone network. In the S3 step, the Darknet-53 network model enhances the expression ability and learning ability of the model by adding residual components, thereby improving the accuracy and efficiency of target detection, and setting shortcut connections in the network to speed up information transmission and training. The output of the Darknet-53 network structure is passed to a feature extraction layer and then to a multi-layer convolutional neural network. The YOLO algorithm can capture the features of image data sets of different scales and levels, thereby improving the accuracy and robustness of target detection.
[0047] In the above technical solution, the YOLO algorithm further introduces the MSRCR algorithm to enhance the image data, and adds a color restoration factor C to the MSRCR algorithm to solve the problem of color distortion caused by contrast enhancement in local areas of the image.
[0048] According to the above technical solution, in step S2, the data processing center uses the target detection model data based on the YOLO algorithm to perform target detection on the image dataset that collects vehicle and pedestrian features, and obtains vehicle and pedestrian information data, and uses the MSRCR algorithm to enhance the image dataset.
[0049] The MSRCR algorithm is as follows:
[0050]
[0051] Where, log is the logarithm of 10, C i is the color restoration factor of the i-th channel, which is used to adjust the color ratio of the three channels in the RGB color channel. i (x, y) is the image of the i-th channel, f[*] is the mapping function of the color space, β is the gain constant, α is the controlled nonlinear intensity, R MSRi (x, y) is the output image of the MSR processing of the i-th channel, R MSRCRi (x,y) is the output image of the MSRCR processing of the i-th channel.
[0052] Specifically, such as Figure 4 、 Figure 5 As shown in Figure 2, the network structure of Darknet-53 and the structure of the residual component include 53 convolutional layers and a large number of 3×3, 1×1 sampling, each with a sampling step of 2.
[0053] Among them, the YOLO algorithm uses the K-means clustering method to obtain the size of the prior frame, and predicts objects at three different scales: 13×13, 26×26, and 52×52. Each scale predicts three target bounding boxes. When detecting an image, the initial network is divided into K×K, and C target categories need to be predicted. The tensor obtained at each scale is K·K·[3×(4+1+C)], where 4 is the offset coordinate of the target bounding box and 1 is the confidence score. Using multi-scale features for object detection can effectively detect small-sized objects.
[0054] Furthermore, the YOLO algorithm uses the output of logistic to predict object categories. Since objects appearing in a bounding box may belong to different categories, logistic can be used for multi-label classification, so that an object can have multiple labels, which can increase the flexibility and accuracy of the YOLO algorithm.
[0055] Among them, when detecting vehicles and people, the size of the candidate bounding boxes of vehicles and people directly affects the accuracy and speed of training detection. In order to adapt to the characteristics of vehicle and human recognition in a mine environment and achieve the best training effect, the YOLO algorithm uses the K-means++ algorithm to cluster the data set to obtain the parameters of the initial detection box and the initial clustering box. The K-means++ algorithm is used to improve the traditional K-means algorithm and reduce the randomness of the traditional K-means algorithm. The specific steps of the improved algorithm are as follows:
[0056] A1. Randomly extract a sample from the image dataset as the initial cluster center C i ;
[0057] A2. First, calculate the shortest distance D(x) between each sample and the current cluster center;
[0058] Calculate the probability of each sample being the next cluster center Finally, the next cluster center is selected according to the roulette method;
[0059] A3. Repeat step 2 until K cluster centers are selected.
[0060] A4. For each sample X in the image dataset i , calculate its distance to the K cluster centers, and divide it into the class corresponding to the cluster center with the smallest distance;
[0061] A5. For each category C i , recalculate its cluster center C i ;
[0062] A6. Repeat steps 4 and 5 above until the location of the cluster center no longer changes;
[0063] The algorithm is:
[0064] Furthermore, the K-means++ algorithm uses the intersection-over-union (IOU) as the distance formula, which reduces the error caused by large borders. tran is the prior target box, box truth is the real target frame, and the formula is as follows:
[0065]
[0066] Furthermore, the target detection model data includes a loss function, which consists of three parts: the coordinate error of the bounding box, the confidence error of the bounding box, and the classification error. The algorithm of the loss function is as follows:
[0067] loss=coordError+iouError+classError
[0068]
[0069]
[0070] Among them, coordError is the coordinate error, iouError is the confidence error, classError is the classification error, λ coord is the penalty weight coefficient of the coordinate error, set to 5, K×K is the number of grids divided into images, and B is the number of bounding boxes corresponding to each grid. is the horizontal coordinate of the center point of the predicted i-th bounding box, is the ordinate of the center point of the predicted i-th bounding box, is the width of the predicted i-th bounding box, is the height of the predicted i-th bounding box, x i is the horizontal coordinate of the center point of the true i-th bounding box, y i is the true vertical coordinate of the center point of the i-th bounding box, ω i is the actual width of the i-th bounding box, h i is the height of the true i-th bounding box, λ noobj is the penalty weight coefficient of confidence error, set to 0.5, is the confidence of the predicted target in the i-th grid, C i is the confidence that the target exists in the true i-th grid, is the confidence of the predicted i-th grid target category, D i is the true confidence of the i-th grid target category, is the predicted category probability of the i-th grid, p i (c) is the true category probability of the i-th grid.
[0071] It should also be noted that by using Python to process the images of the target detection model data, it is convenient to preprocess the image data set, improve the accuracy of target detection, and facilitate the integration of the target detection model with the deep learning framework, using its powerful computing power to accelerate the training and reasoning of the model, thereby improving the accuracy and robustness of underground human and vehicle identification, thereby reducing the risk of accidents and ensuring production safety.
[0072] In summary: By combining MSRCR image enhancement processing and YOLO algorithm, and using K-means++ clustering optimization algorithm to determine the size of the prior frame, and adopting a reasonable loss function, high-precision and high-speed detection of people and vehicles is achieved. Through MSRCR image enhancement processing, the contrast and brightness of the image can be enhanced, making the target clearer and more obvious, which helps to improve the detection accuracy. The K-means++ clustering optimization algorithm can effectively optimize the size of the prior frame and improve the detection speed. At the same time, the improved YOLO algorithm has the characteristics of high efficiency and accuracy, can quickly identify targets, and ensure the real-time detection, thereby realizing effective application in the underground air door control system, improving recognition efficiency and accuracy, and by combining the data of wind speed, air pressure, gas concentration and ambient temperature collected by the air door, the human and vehicle behavior in the mine can be comprehensively analyzed, accurate identification and detection can be achieved, and it can be judged in time whether there is abnormal danger.
[0073] Parts not described in the present invention are the same as those in the prior art or can be implemented using the prior art. Although the embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations may be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An automatic identification system for people and vehicles underground, characterized in that: The identification method of the automatic identification system comprises the following steps: S1. The camera collects an image dataset of vehicle and pedestrian features. Computer vision technology uses image processing and image detection technology to process the image dataset, obtain the parameters of the initial detection frame, and send it to the data processing center. The sensor on the damper collects data on wind speed, air pressure, gas concentration, and temperature around the damper, and sends it to the data processing center. S2. The data processing center uses the K-means++ algorithm to cluster the processed image dataset to obtain the parameters of the initial clustering frame. Based on the K-means clustering algorithm, the initial detection frame parameters and the initial clustering frame parameters are clustered to obtain the size of the prior frame. Then, based on the YOLO algorithm, the size of the prior frame obtained after clustering is used for model training to obtain the target detection model data. In step S2, a sample is randomly selected from the data set as the initial cluster center. The selection of the initial cluster center in the K-means++ algorithm is as follows: A1. Randomly extract a sample from the image dataset as the initial cluster center C i ; A2. First, calculate the shortest distance D(x) between each sample and the current cluster center; Calculate the probability of each sample being the next cluster center Finally, the next cluster center is selected according to the roulette method; A3. Repeat step 2 until K cluster centers are selected. In step S2, the data processing center uses the target detection model data based on the YOLO algorithm to perform target detection on the image dataset that collects vehicle and pedestrian features, obtains vehicle and pedestrian information data, and uses the MSRCR algorithm to enhance the image dataset; The MSRCR algorithm is as follows: Where, log is the logarithm of 10, C i is the color restoration factor of the i-th channel, which is used to adjust the color ratio of the three channels in the RGB color channel. i (x, y) is the image of the i-th channel, f[*] is the mapping function of the color space, β is the gain constant, α is the controlled nonlinear intensity, R MSRi (x, y) is the output image of the MSR processing of the i-th channel, R MSRCRi (x, y) is the output image of the MSRCR processing of the i-th channel; S3. The data processing center processes the collected data on wind speed, air pressure, gas concentration, and temperature around the damper, and combines it with the acquired target monitoring model data to identify and detect the behavior of vehicles and pedestrians in the mine, ultimately identifying the vehicle or behavior. S4. The data processing center is based on a time synchronization protocol and analyzes the behavior of vehicles and pedestrians to determine whether there are abnormal dangers for vehicles and pedestrians, and controls the opening and closing of the dampers.
2. The automatic identification system for underground vehicles and people according to claim 1 is characterized in that: The iterative update of cluster centers in the K-means algorithm is as follows: C i is the cluster center, X i for each sample in the image dataset.
3. The automatic identification system for underground people and vehicles according to claim 1, characterized in that: The K-means clustering algorithm predicts objects at three different scales of the prior frame size, where three target bounding boxes are predicted for each size.
4. The automatic identification system for underground people and vehicles according to claim 3, characterized in that: in, The sizes of the prior frames include 13×13, 26×26, and 52×52.
5. The automatic identification system for underground people and vehicles according to claim 1, characterized in that: The K-means++ algorithm uses the intersection-over-union (IOU) as the distance formula. The formula is as follows: box tran is the prior target box, box truth is the true target box.
6. The automatic identification system for underground vehicles and people according to claim 1, characterized in that: The object detection model data includes a loss function, which consists of three parts: the coordinate error of the bounding box, the confidence error of the bounding box, and the classification error.
7. The automatic identification system for underground vehicles and people according to claim 6, characterized in that: The algorithm of the loss function is as follows: loss=coordError+iouError+classError Among them, coordError is the coordinate error, iouError is the confidence error, classError is the classification error, λ coord is the penalty weight coefficient of the coordinate error, set to 5, K×K is the number of grids divided into images, and B is the number of bounding boxes corresponding to each grid. is the horizontal coordinate of the center point of the predicted i-th bounding box, is the ordinate of the center point of the predicted i-th bounding box, is the width of the predicted i-th bounding box, is the height of the predicted i-th bounding box, x i is the horizontal coordinate of the center point of the true i-th bounding box, y i is the true vertical coordinate of the center point of the i-th bounding box, ω i is the actual width of the i-th bounding box, h i is the height of the true i-th bounding box, λ noobj is the penalty weight coefficient of confidence error, set to 0.5, is the confidence of the predicted target in the i-th grid, C i is the confidence that the target exists in the true i-th grid, is the confidence of the predicted i-th grid target category, D i is the true confidence of the i-th grid target category, is the predicted category probability of the i-th grid, p i (c) is the true category probability of the i-th grid.
Citation Information
Patent Citations
Deep learning dam crack detection method suitable for complex underwater environment
CN114299060A
Obstacle detection method and device in complex weather
CN115376108A