A method, system, device and storage medium for detecting violations of mine workers
By embedding the EMA attention module and self-attention mechanism detection head DyHead in the YOLOv8 model, the problem of poor target detection effect of YOLOv8 in complex mine environments is solved, and higher detection accuracy and performance are achieved, effectively preventing mine operation accidents.
Patent Information
- Application Number
- CN202411511282.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-10-28
AI Technical Summary
YOLOv8 has poor target detection effect in complex mine environments, especially in the case of insufficient light and severe interference, with poor detection accuracy and performance.
The EMA attention module is embedded after the three C2f modules of the original YOLOv8 model backbone network, and the self-attention mechanism dynamic object detection head DyHead replaces the detection head of the original network to form an improved YOLOv8 model. This model enhances feature extraction ability through EMA attention mechanism and improves target positioning accuracy through DyHead detection head.
The improved YOLOv8 model significantly improves the accuracy and performance of target detection in complex mine environments, can better adapt to changes in the underground target scale, and avoid losing characteristic information of small targets, effectively preventing mine operation accidents.
Smart Images

Figure CN119479064B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition and processing, and particularly relates to a method, system, device and storage medium for detecting illegal targets of mine workers. Background Art
[0002] In recent years, the mine industry in China has achieved remarkable development in scale. The scale of mines is not only reflected in the number and production capacity of mines, but also in the scale and enhanced production capacity of individual mines. With the progress of technology and the improvement of management, many mines have achieved large-scale and intensive production, further improving production capacity and efficiency.
[0003] Mine accidents are often caused by the combined action of multiple factors, including natural factors such as geological structures and gas outbursts, and also including human factors such as operation errors and improper management. Among them, the illegal production of mine workers is the most easily preventable.
[0004] The realization of safety detection for mine workers is mainly based on object detection algorithms for video images. Object detection algorithms are divided into two categories: traditional machine learning and deep neural networks. Traditional machine learning algorithms have low pertinence, high time complexity, and window redundancy; moreover, the manually designed features have poor robustness and weak generalization ability, resulting in the gradual replacement of traditional machine learning algorithms by deep learning algorithms. After introducing the convolutional neural network (CNN) method into object detection by the region-based deep convolutional neural network R-CNN, the detection performance of the model has been improved. Subsequently, the proposed Fast R-CNN and Faster R-CNN have greatly improved the detection efficiency. Fast R-CNN improved R-CNN by sending the entire image into the convolutional network to extract features, then extracting the feature representations of candidate regions through the RoI pooling layer on the feature map, and finally performing classification and bounding box regression through the fully connected layer. As two-stage detection algorithms, these two algorithms usually perform well in terms of accuracy, but due to the need for two independent stages, they are slower. Subsequently, single-stage detection algorithms represented by YOLO and SSD were proposed, which solved the problem of the slow speed of two-stage detection algorithms and enabled them to be applied to scenarios with high real-time requirements. However, the detection accuracy and performance are still poor in scenarios with poor lighting and severe interference. Therefore, the above various algorithm models still have drawbacks when detecting in the complex environment of mines. Summary of the Invention
[0005] In order to solve the problem of poor object detection effect of YOLOv8 in the complex environment of mines, the present invention provides a method, system, device and storage medium for detecting illegal targets of mine workers.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A method for detecting illegal mining workers, comprising the following steps:
[0008] The EMA attention modules are respectively embedded after the three C2f modules of the original YOLOv8 model backbone network, and the dynamic target detection head DyHead with self-attention mechanism is introduced to replace the detection head of the original network to form an improved YOLOv8 model; the improved YOLOv8 model is trained to obtain a target detection model for identifying illegal behaviors of mine workers;
[0009] Get pictures of mine workers working in the mine;
[0010] The acquired image is input into the target detection model. After extracting the features of the image from the bottom layer to the high layer through the backbone network, the attention mechanism focuses on the important features. The image is convolved and downsampled to extract feature maps containing semantic information at different levels. The feature maps are sent to the DyHead detection head after feature fusion through the PAN-FPN structure, and the target violation detection results are generated through the DyHead detection head.
[0011] Preferably, the method further comprises obtaining a data set by adopting a self-built underground personnel detection data set and a target detection data set PASCAL VOC2012, wherein the self-built underground personnel detection data set is specifically: obtaining a plurality of mine operation pictures by a camera in a constructed mine operation simulation scene; the target detection data set PASCAL VOC2012 contains a plurality of object images, which contain annotations of the category, position and size information of the objects; and selecting data in the data set as a training set in proportion;
[0012] The improved YOLOv8 model is trained using the data of the training set.
[0013] Preferably, the target violation detection result determines whether the mine worker's behavior is in violation through a bounding box and a category label, specifically: marking the location information of the violation in the image through the bounding box; and outputting the worker's violation type through the category label.
[0014] Preferably, the processing process of the DyHead detection head specifically includes the following steps:
[0015] The input vector obtained from Neck is sent to three perception-enhanced attention modules. The three perception-enhanced modules obtain enhanced outputs through three attentions applied at different positions, specifically:
[0016] W(F)=π C (π S (π L (F)·F)·F)·F;
[0017] Among them, F represents the input feature, and π C , π S , π L represent the channel dimension, the spatial dimension, and the scale dimension respectively; the operation of dynamically fusing features of different scales according to the semantic importance of different scales is performed on the input L dimension, specifically:
[0018]
[0019] Among them, f is a linear function of a 1×1 convolutional layer, and the activation function is hard–sigmoid; represents the global average pooling operation on the feature map F. By summing the spatial dimension and the channel dimension of the feature map, the global average information of the feature map is obtained; S represents the spatial dimension, that is, the height and width of the feature map, and C represents the number of channels;
[0020] Use deformable convolution sparse attention learning to aggregate elements of different levels at the same spatial position, specifically:
[0021]
[0022] Among them, L represents the number of positions in the spatial dimension, that is, the number of pixel points in the feature map, represents the normalization processing of the spatial dimension; l represents a specific spatial position index in the feature map; W l,k is the spatial weighting term, indicating the weight occupied by the feature value at a specific spatial position l and offset k; K is the number of sparse sampling positions; p k +Δp k is the position moved by the self-learned spatial offset, and Δp k is learned by deformable convolution and focuses on the ambiguous area; Δm k represents the self-learned importance at the position Δp k ;
[0023] Automatically select to open or close the channel through the switch control hyperparameter whether to learn the activation threshold to adapt to different tasks, specifically:
[0024] π C (F)·F = max(α 1 (F)·F c +β 1 (F), α 2 (F)·F c +β 2 (F));
[0025] Among them, max() is a hyperparameter function for learning the activation threshold of channels; F c is the feature split at the c-th channel; α is a weighting coefficient related to the input feature F; β is an offset term, a learnable parameter related to the input feature F.
[0026] Preferably, the evaluation metrics for the improved YOLOv8 model include detection accuracy and model complexity. The detection accuracy is measured by the accuracy P, recall R, and mean average precision mAP of the model; the model complexity is measured by the number of model parameters and the model computational volume.
[0027] The present invention also provides a system for detecting illegal behaviors of mine workers, specifically including:
[0028] A model construction module, which is used to respectively embed EMA attention modules after three C2f modules of the backbone network of the original YOLOv8 model, and introduce a self-attention mechanism dynamic object detection head DyHead to replace the detection head of the original network, thus constructing an improved YOLOv8 model; train the improved YOLOv8 model to obtain an object detection model for identifying illegal behaviors of mine workers.
[0029] A data acquisition module, which is used to acquire pictures of the working scenes of mine workers.
[0030] An object detection module, which is used to input the acquired pictures into the object detection model. After extracting the features of the pictures from the bottom layer to the high layer through the backbone network, it focuses on important features through the attention mechanism; performs convolution and downsampling operations on the pictures to extract feature maps containing different levels of semantic information; sends the feature maps through the PAN-FPN structure for feature fusion and then to the DyHead detection head part, and generates object illegal behavior detection results through the DyHead detection head.
[0031] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps described in the method for detecting illegal behaviors of mine workers.
[0032] The present invention also provides a computer-readable storage medium, on which a computer program is stored. The characteristic is that when the computer program is loaded by a processor, it can execute the steps described in the method for detecting illegal behaviors of mine workers.
[0033] The method for detecting illegal behaviors of mine workers provided by the present invention has the following beneficial effects:
[0034] An underground target detection algorithm based on the improved YOLO8 proposed by the present invention introduces an EMA attention mechanism layer after each of the three C2f modules in the backbone network, enhancing the model's feature extraction ability under insufficient light conditions, better adapting to the changes in the scales of underground targets, and using a unified self-attention mechanism detection head DyHead to replace the detection head in the original YOLO8 model. After passing through the DyHead detection head after EMA attention processing, the model's target localization accuracy can be improved, and the feature information of small targets can be prevented from being lost. Therefore, the target detection model obtained after training based on the improved YOLO8 model improves the accuracy of target detection results in the complex underground environment compared with other algorithm models, detects the illegal production of underground miners, and prevents underground operation accidents. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention and its design, the accompanying drawings required for this embodiment will be briefly introduced below. The accompanying drawings in the following description are only partial embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a flowchart of a method for detecting illegal targets of underground miners according to the present invention.
[0037] Figure 2 It is a structural diagram of EMA according to the present invention.
[0038] Figure 3 It is a structural diagram of DyHead according to the present invention.
[0039] Figure 4 It is a network structure diagram of the improved YOLO8 model according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to enable those skilled in the art to better understand the technical solutions of the present invention and implement them, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.
[0041] Embodiment
[0042] The present invention provides a method for detecting illegal targets of underground miners. As Figure 1 shown in the flowchart of a method for detecting illegal targets of underground miners, it includes:
[0043] S1. Use the self-built underground personnel detection dataset and the PASCAL VOC2012 public dataset to obtain the dataset, and select the data in the dataset as the training set according to a certain proportion.
[0044] The self-built underground personnel detection dataset is specifically as follows: Use the Hikvision vertical camera DS-2DC4423IW-D to take pictures of personnel in the constructed scenario. Raise the camera lens to simulate the angle of a real surveillance camera, and walk around randomly in the scenario to simulate real target scale changes. At the same time, reduce the illumination of the scenario to simulate the low-light conditions in a real mine. Then, split the video into frames to obtain pictures, and use the LabelImg tool to annotate the illegal targets. The self-built dataset contains a total of 5 illegal categories and 3000 pictures, including 2100 training set images, 300 validation set pictures, and 600 test set pictures.
[0045] The PASCAL VOC2012 public dataset is a well-known public dataset in the field of computer vision. This dataset has 20 common object categories, such as cars, people, airplanes, etc., as well as some less common categories, such as chairs, birds, bottles, etc. Each image contains annotation information such as the category, location, and size of the object. The dataset format uses XML files to describe the images and their annotation information. Considering that this dataset is relatively similar to the self-built dataset, it is used to test the generalization performance of the proposed algorithm model and increase data authenticity.
[0046] S2. Use the training set to train the improved YOLOv8 model to obtain a target detection model. YOLOv8 itself consists of three parts: Backbone, Neck, and Head. Among them, Backbone is composed of multiple Conv, C2f, and SPPF modules connected in sequence, which is used to extract features from the input image. Conv is a basic convolutional layer, and the C2f module is the main module for learning residual features. The Neck part is generally a simple bidirectional fusion PAN structure, which integrates high semantic information and high-resolution information by fusing high-level features and low-level features. The Head part consists of multiple detection heads. Since Backbone and Neck can only extract and fuse features and cannot complete the localization task, Head needs to decouple the feature information fused by Neck to obtain the category and location of the target object.
[0047] In the present invention, EMA attention modules are respectively embedded after three C2f modules in the backbone network of the original YOLOv8 model. The EMA attention module contains a parallel sub-structure of a CA block, and extracts the attention weight descriptors of the grouped feature maps through three parallel routes. Two parallel paths are on the 1x1 branch, and the third path is on the 3x3 branch. In the 1x1 branch, two 1D global average pooling operations are respectively adopted in two spatial directions to encode the channels. In the 3x3 branch, only one 3x3 kernel is stacked to capture multi-scale feature representations. The output feature maps within each group are calculated as a set of two generated spatial attention weight values. Finally, the Sigmoid function is used, increasing the original feature extraction network from 10 layers to 13 layers for network model optimization. It improves the model's perception ability for targets of different scales, enables the model not to lose the feature information of small targets, and simultaneously improves the model's feature extraction ability.
[0048] Then, the self-attention mechanism dynamic object detection head DyHead is introduced to replace the detection head of the original network. The DyHead detection head takes the input vector of (L, S, C) obtained from the Neck and sends it into three perception-enhanced attention modules. The three perception-enhanced modules obtain enhanced outputs through three attentions applied at different positions. Specifically:
[0049] W(F) = π C (π S (π L (F)·F)·F)·F;
[0050] Among them, F represents the input feature, and π C π S π L respectively represent the channel dimension, the spatial dimension, and the scale dimension. The enhanced spatial perception is achieved by introducing scale-aware attention and dynamically fusing features of different scales according to the semantic importance of different scales. The operation is performed on the input L dimension. Specifically:
[0051]
[0052] Among them, f is a linear function of a 1×1 convolutional layer, and the activation function is hard–sigmoid; represents the global average pooling operation on the feature map F. By summing the spatial dimension and the channel dimension of the feature map, the global average information of the feature map is obtained. S represents the spatial dimension, that is, the height and width of the feature map, and C represents the number of channels.
[0053] The enhanced spatial perception uses deformable convolution to make attention learning sparse, and then aggregates elements at different levels in the same spatial position, as shown in formula (3):
[0054]
[0055] Among them, L represents the number of positions in the spatial dimension, that is, the number of pixel points in the feature map. Indicates normalization processing of the spatial dimension; l represents a specific spatial position index in the feature map; W l,k is the spatial weighting term, indicating the weight occupied by the feature value at a specific spatial position l and offset k; K is the number of sparse sampling positions, p k +Δp k is the position moved by the self-learned spatial offset, Δp k is learned by deformable convolution and focuses on the ambiguous area, Δm k represents the importance of self-learning at position Δp k
[0056] Enhanced task awareness automatically selects to turn on or off channels by controlling whether the hyperparameter of the learning activation threshold is turned on or off to adapt to different tasks, as shown in formula (4):
[0057] π C (F)·F = max(α 1 (F)·F c +β 1 (F), α 2 (F)·F c +β 2 (F));
[0058] Among them, max() is a hyperparameter function to learn the activation threshold of the channel, F c is the feature split at the c-th channel; α is the weighting coefficient related to the input feature F; β is an offset term, a learnable parameter related to the input feature F.
[0059] S3. Input the picture of the mine operation scenario into the target detection model to obtain the target behavior detection result in the mine. The target behavior detection result judges whether the behavior of the mine workers is illegal through the bounding box and class label. Specifically: the position information of the illegal behavior in the picture is marked through the bounding box, and the class label is used to output the type of illegal behavior of the worker.
[0060] S4. The main evaluation metrics for the model detection accuracy are two categories: detection accuracy and model complexity. Among them, the detection accuracy is mainly reflected by the accuracy P (precision), recall rate R (recall) and mean average precision mAP (mean average precision) of the model. The model complexity of the algorithm is reflected by the number of model parameters and the model calculation amount, and the larger the value, the higher the model complexity.
[0061] Accuracy can be used to evaluate the detection accuracy of the model, which is defined as the proportion of correctly predicted positives among all predicted positives. Recall can be used to evaluate the comprehensiveness of the model detection, which is expressed as the proportion of correctly predicted positives among all actual positives. And mAP is one of the most important model performance evaluation indicators in the field of object detection, which is used to measure the detection accuracy and comprehensiveness of the model in multiple categories. In this invention, mAP@0.5 with an IOU threshold of 0.5 is used as the main evaluation indicator. Assuming the number of true positive and false positive samples in the prediction result are TP and FP respectively, and the number of positive samples predicted as false is FN, the calculation formulas for accuracy P, recall R, and mean average precision mAP can be obtained as follows:
[0062]
[0063] The model complexity of the algorithm is reflected by the number of model parameters and the model computational amount. The larger the value, the higher the model complexity.
[0064] The mine object detection algorithm based on the improved YOLOv8n proposed in this invention introduces an EMA attention mechanism layer in the backbone network to enhance the feature extraction ability of the model. The unified self-attention mechanism detection head DyHead is used to replace the detection head of the original network, which simultaneously enhances the model's perception ability of scale, space, and tasks, and improves the feature expression ability of the detection head. It not only has lower complexity but also higher detection accuracy. Compared with other models, the model of this invention is more suitable for object detection in mines.
[0065] To verify the effectiveness and rationality of the improvement method of the YOLOv8n network model in this invention, ablation experiments are carried out on the self-built dataset to compare the influence of each module on the model, as shown in Table 1. At the same time, in order to enhance the reliability and authenticity of the data, the same experiment is also carried out on the VOC2012 public dataset, as shown in Table 2.
[0066] Table 1 Ablation experiment on the self-built dataset
[0067]
[0068] Table 2 Ablation experiment on VOC2012
[0069]
[0070] As can be seen from the results in Table 1 and Table 2, on the self-built dataset, by adding the two improvements of EMA and DyHead respectively, the detection accuracy can be improved by 1.1% and 2.3% respectively. At the same time, after adding both improvements, the improvement in the model detection accuracy is as high as 3.2%. On the public dataset VOC2012, by adding the two improvements of EMA and DyHead respectively, the detection accuracy can be improved by 1.4% and 1.5% respectively. At the same time, after adding both improvements, the model detection accuracy is improved by 2.0%. Therefore, compared with the original model, the model improvement of the present invention not only improves the model detection ability, but also has good generalization performance.
[0071] In order to further evaluate the performance of the present invention in underground mine target detection, a comparative experiment was conducted on the self-built dataset with the model of the present invention, YOLOv3-tiny, YOLOv5n, and YOLOv7-tiny models. As can be seen from Table 3, compared with the original models of YOLOv5n and YOLOv8n, although the complexity of the model of the present invention is relatively large, it has better performance for underground target detection. Compared with YOLOv3-tiny and YOLOv7-tiny, not only is the model complexity lower, but also it has higher detection accuracy. Compared with other mainstream algorithms, it is more suitable for underground mine target detection.
[0072] Table 3 Comparison with other mainstream algorithms
[0073]
[0074] The present invention also provides an underground mine worker violation target detection system, which specifically includes:
[0075] A model construction module, which is used to embed an EMA attention module after each of the three C2f modules in the backbone network of the original YOLOv8 model, and introduce a self-attention mechanism dynamic target detection head DyHead to replace the detection head of the original network to form an improved YOLOv8 model; train the improved YOLOv8 model to obtain a target detection model for identifying underground mine worker violation behaviors.
[0076] A data acquisition module, which is used to acquire pictures of the underground mine worker operation scene.
[0077] A target detection module, which is used to input the acquired pictures into the target detection model. After extracting the features of the pictures from the bottom layer to the high layer through the backbone network, it focuses on important features through the attention mechanism; performs convolution and downsampling operations on the pictures to extract feature maps containing different levels of semantic information; after fusing the feature maps through the PAN-FPN structure, it sends them to the DyHead detection head part, and generates target violation behavior detection results through the DyHead detection head.
[0078] Each module in the above-mentioned mine worker violation target detection system can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of the processor in a computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0079] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps in an embodiment of a mine worker violation target detection method. The specific implementation method can be referred to the method embodiment and will not be elaborated here.
[0080] Furthermore, the present invention also provides a non-transitory computer-readable storage medium containing instructions, on which a computer program is stored. For example, a memory containing instructions, and the above instructions can be executed by the processor of the computer device to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. When the computer program is executed by the processor, it can implement the steps in an embodiment of a mine worker violation target detection method. The specific implementation method can be referred to the method embodiment and will not be elaborated here.
[0081] Those skilled in the art should understand that the embodiments of the present invention can provide a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0082] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0083] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in the process(es) Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks.
[0084] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the process(es) Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks.
[0085] It should be noted that the above-described specific embodiments can enable those skilled in the art to more comprehensively understand the present invention, but do not limit the present invention in any way. Therefore, although the present specification and embodiments have described the present invention in detail, those skilled in the art should understand that the present invention can still be modified or equivalently replaced; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are covered by the protection scope of the patent of the present invention. Any reference numeral in the claims should not be regarded as limiting the claimed claim. Any simple variations or equivalent replacements of technical solutions that can be obviously obtained by those skilled in the art within the technical scope disclosed by the present invention all fall within the protection scope of the present invention.
Claims
1. A method for detecting illegal behaviors of mine workers, characterized in that: The following steps are involved: The EMA attention modules are respectively embedded after the second to fourth C2f modules of the original YOLOv8 model backbone network, and the dynamic target detection head DyHead with self-attention mechanism is introduced to replace the detection head of the original network to form an improved YOLOv8 model; the improved YOLOv8 model is trained to obtain a target detection model for identifying illegal behaviors of mine workers; Get pictures of mine workers working in the mine; The acquired image is input into the target detection model, and after extracting the features of the image from the bottom layer to the high layer through the backbone network, the important features are focused on through the attention mechanism; Convolution and downsampling operations are performed on the image to extract feature maps containing semantic information at different levels; the feature maps are sent to the DyHead detection head after feature fusion through the PAN-FPN structure, and the target violation detection result is generated through the DyHead detection head.
2. A method for detecting illegal behaviors of mine workers according to claim 1, characterized in that: The method also includes obtaining a data set by using a self-built underground personnel detection data set and a target detection data set PASCAL VOC2012. Specifically, the self-built underground personnel detection data set is: obtaining multiple pictures of mine operations through a camera in a constructed mine operation simulation scene; the target detection data set PASCAL VOC2012 has multiple object images, which include annotations of the category, position and size information of the objects; and selecting data in the data set as a training set in proportion; The improved YOLOv8 model is trained using the data of the training set.
3. A method for detecting illegal behaviors of mine workers according to claim 1, characterized in that: The target violation detection result determines whether the mine worker's behavior is in violation of the law through a bounding box and a category label. Specifically, the position information of the violation in the image is marked through the bounding box; and the type of the worker's violation is output through the category label.
4. A method for detecting illegal mining workers according to claim 1, characterized in that: The processing process of the DyHead detection head specifically includes the following steps: The input vector obtained from Neck is sent to three perception-enhanced attention modules. The three perception-enhanced modules obtain enhanced outputs through three attentions applied at different positions, specifically: W(F)=π C (p S (p L (F)·F)·F)·F; Among them, F represents the input feature, π C , π S , π L They represent the channel dimension, spatial dimension, and scale dimension respectively; according to the semantic importance of different scales, the features of different scales are dynamically fused and operated on the input L dimension, specifically: Among them, f is a linear function composed of 1×1 convolutional layers, and the activation function is hard-sigmoid; Indicates a global average pooling operation on the feature map F. By summing the spatial dimension and channel dimension of the feature map, the global average information of the feature map is obtained. S represents the spatial dimension, that is, the height and width of the feature map, and C represents the number of channels. Use deformable convolution to sparse attention learning and aggregate different levels of elements at the same spatial position, specifically: Among them, L represents the number of positions in the spatial dimension, that is, the number of pixels in the feature map. Indicates the normalization of the spatial dimension; l represents a specific spatial position index in the feature map; W l,k is the spatial weighting term, which indicates the weight of the eigenvalue at a specific spatial position l and offset k; K is the number of sparse sampling positions; p k +Δp k is the position moved by the self-learned spatial offset, Δp k Learned by deformable convolution, focusing on ambiguous areas; Δm k Represents the position Δp k The importance of self-learning; The switch controls whether the hyperparameter learns the activation threshold to automatically choose to open or close the channel to adapt to different tasks. Specifically: p C (F)·F=max(α 1 (F)·F c +b 1 (F),a 2 (F)·F c +b 2 (F)); Among them, max() is a hyperparameter function to learn the threshold of activating channel; F c is the feature split at the cth channel; α is the weighting coefficient related to the input feature F; β is an offset term, a learnable parameter related to the input feature F.
5. A method for detecting illegal behaviors of mine workers according to claim 1, characterized in that: The evaluation indicators of the improved YOLOv8 model include detection accuracy and model complexity. The detection accuracy is measured by the model's accuracy P, recall R and average accuracy mAP; the model complexity is measured by the number of model parameters and the amount of model calculation.
6. A mine worker violation target detection system, characterized in that: include: A model building module is used to embed EMA attention modules after the second to fourth C2f modules of the original YOLOv8 model backbone network, and introduce a self-attention mechanism dynamic target detection head DyHead to replace the detection head of the original network to form an improved YOLOv8 model; the improved YOLOv8 model is trained to obtain a target detection model for identifying illegal behaviors of mine workers; A data acquisition module is used to obtain pictures of the working scenes of mine workers; The target detection module is used to input the acquired image into the target detection model, extract the features of the image from the bottom layer to the high layer through the backbone network, and focus on the important features through the attention mechanism; Convolution and downsampling operations are performed on the image to extract feature maps containing semantic information at different levels; the feature maps are sent to the DyHead detection head after feature fusion through the PAN-FPN structure, and the target violation detection result is generated through the DyHead detection head.
7. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is loaded into a processor, it can execute the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
YOLO-based deep learning industrial field production working area violation detection method
CN114913606A
Enteroscope polyp real-time detection method, system and model training method
CN118134887A