Method, system and equipment for detecting personal protective equipment of miner and storage medium
By using visual transformation network model and multi-scale spatial attention mechanism in the detection of miners' personal protective equipment for feature extraction and fusion, combined with the detection model of spatial reconstruction unit and channel reconstruction unit, the existing target detection algorithm is solved, and efficient and real-time detection effect is achieved.
Patent Information
- Application Number
- CN202510010929.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
Existing object detection algorithms, such as Faster R-CNN, YOLOv7, YOLOv7-tiny, LMIENet image enhancement algorithm, and YOLOv8, have problems with slow detection speed or complex calculations, which affect the detection results.
The backbone network model based on the visual transformation network model is used for feature extraction, and the neck network model combined with the multi-scale spatial attention mechanism is used for feature fusion and enhancement. The detection model is constructed using spatial reconstruction units and channel reconstruction units to improve the accuracy and efficiency of detection.
While ensuring detection accuracy, reducing calculation complexity, improving the real-time nature of the algorithm, achieving lightweight, balancing calculation speed and recognition accuracy, and providing strong support for coal mine safety production.
Smart Images

Figure CN119942591A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of mining production safety technology, and in particular to a method, system, equipment and storage medium for detecting miners' personal protective equipment. Background Art
[0002] Miners working underground is a high-risk working environment, so they need to comply with a series of regulations and safety measures to ensure the safety and health of miners. Before entering the mine, it is necessary to test the miners' personal protective equipment and determine whether the miners are fully equipped, such as wearing safety helmets. In the mine, there are problems such as uneven lighting underground and serious interference from coal dust, which leads to high noise and blurred images in the monitoring video. It is necessary to identify whether the miners' current personal protection is in place. The above problems will affect the accuracy of recognition.
[0003] Personal protective equipment (PPE) refers to devices or appliances worn or used to protect individuals from hazards. The purpose of PPE testing is to ensure that these devices can effectively protect users from potential hazards, such as physical, chemical, biological and other sources. PPE testing is essential to ensure workplace safety. It can help identify and correct design defects and ensure that products meet relevant safety standards and regulatory requirements. According to the latest research progress, PPE testing technology is constantly developing to improve the accuracy and efficiency of testing. For example, PPE detection systems based on artificial intelligence and computer vision have been proposed. These systems can automatically analyze images and video frames to check whether employees are wearing personal protective equipment such as masks, helmets and gloves correctly.
[0004] The wearing of miners' safety equipment is mainly checked by full-time personnel. The inspection by full-time personnel is easily affected by human subjective factors and requires a large amount of human resources, which results in a large amount of management cost and time cost, poor inspection effect and low efficiency. Summary of the invention
[0005] The technical problem to be solved by this application is that the existing target detection algorithms and related models, such as Faster R-CNN, YOLOv7, YOLOv7-tiny, LMIENet image enhancement algorithm, YOLOv8, etc., have problems of slow detection speed or complex calculation, which affects the detection results.
[0006] In order to solve the above problems, in order to solve the above technical problems or at least partially solve the above technical problems, the present application provides a method, system, device and storage medium for detecting miners' personal protective equipment.
[0007] In a first aspect, the present invention discloses a method for detecting personal protective equipment for miners, which comprises the following steps:
[0008] Acquire real-time miner image data, construct a backbone network model based on the visual transformation network model, perform feature extraction on the real-time miner image data through the backbone network model, and extract a first feature image;
[0009] Based on the multi-scale spatial attention mechanism, a neck network model is constructed, and feature fusion and enhancement processing are performed on the first feature image to obtain a second feature image;
[0010] A detection model is constructed based on the space reconstruction unit and the channel reconstruction unit, and the features in the second feature image are detected and processed to obtain the detection results of the miner's personal equipment.
[0011] Preferably, real-time miner image data is acquired, a backbone network model is constructed based on a visual transformation network model, and features of the real-time miner image data are extracted through the backbone network model to extract a first feature image, specifically including the following steps:
[0012] Acquire real-time miner image data, perform image segmentation and feature extraction on the miner image data, and obtain multiple feature maps;
[0013] Inputting multiple feature maps into the visual transformation network model in sequence to perform target feature detection and extraction processing to obtain processed feature maps;
[0014] A variable-size pooling model is constructed to perform pooling and integration processing on all processed feature maps to obtain the first feature image.
[0015] Preferably, the step of sequentially inputting a plurality of feature maps into a visual transformation network model for target feature detection and extraction to obtain a processed feature map specifically comprises the following steps:
[0016] Multiple feature maps are input into the efficient visual conversion module for grouping and self-attention calculation, and the self-attention feature map is output;
[0017] The self-attention feature map is input into the efficient visual converter downsampling module for layering and downsampling to obtain the downsampled feature maps at each level.
[0018] The downsampled feature maps at each level are subjected to feature fusion processing to obtain the processed feature maps.
[0019] Preferably, the multiple feature maps are input into the efficient visual conversion module for grouping and self-attention calculation, and the self-attention feature map is output, which specifically includes the following steps:
[0020] Multiple feature maps are sequentially input into the feedforward neural network layer for nonlinear transformation processing to obtain nonlinear transformation feature maps;
[0021] Multiple nonlinear change feature maps are grouped, attention calculation is performed, and they are combined in multiple cascades to obtain a self-attention feature map.
[0022] Preferably, constructing a neck network model based on a multi-scale spatial attention mechanism, performing feature fusion and enhancement processing on the first feature image to obtain a second feature image specifically includes the following steps:
[0023] The first feature image is subjected to feature extraction to generate a series of feature maps of different scales;
[0024] The feature fusion model combined with the multi-scale spatial attention mechanism performs feature enhancement and feature fusion processing to obtain the second feature image.
[0025] Preferably, the feature fusion model combined with the multi-scale spatial attention mechanism performs feature enhancement and feature fusion processing to obtain a second feature image, specifically comprising the following steps:
[0026] The processed feature map is input into the convolution module for convolution processing to obtain a convolution feature map;
[0027] The convolution feature map is divided into two parts, one part is directly output to the combination module, and the other part is input to the bottleneck module for feature extraction and enhancement processing to obtain the enhanced feature map;
[0028] The enhanced feature map is divided into two parts again, one part is directly output to the combination module, and the other part is input to the bottleneck module again for feature extraction and enhancement processing to obtain a further enhanced feature map;
[0029] The enhanced feature map is subjected to multiple branch processing and feature enhancement processing, and finally all the enhanced feature maps are output to the combination module for feature fusion to output the second feature image.
[0030] Preferably, the detection model is constructed based on the space reconstruction unit and the channel reconstruction unit, and the features in the second feature image are detected and processed to obtain the miner's personal equipment detection result, which specifically includes the following steps:
[0031] The second feature image is input into the spatial reconstruction unit for spatial refinement processing to obtain a spatial refinement feature map;
[0032] The spatially refined feature map is input into the channel reconstruction unit for channel refinement processing to obtain a channel refinement feature map;
[0033] The channel refinement feature map performs target detection through the detection layer to obtain the miner and personal equipment features and feature parameters in the image. The feature parameters include the category of the target feature, the location of the target feature bounding box, and the confidence score;
[0034] A preset judgment threshold is set, and the feature parameters are compared with the preset judgment threshold to generate a detection result.
[0035] In a second aspect, the present invention discloses a miner's personal protective equipment detection system, which implements the above-mentioned miner's personal protective equipment detection method steps.
[0036] In a third aspect, the present invention discloses a computer device, which includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0037] Memory, used to store computer programs;
[0038] The processor is used to implement the steps of the miner's personal protective equipment detection method when executing the program stored in the memory.
[0039] In a fourth aspect, the present invention discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a method for detecting personal protective equipment for miners.
[0040] The above technical solution provided by this application has the following advantages compared with the prior art:
[0041] The present application provides a method, system, device and storage medium for detecting personal protective equipment for miners. The method mentions that the backbone network model constructed based on the visual transformation network model has better feature extraction capability, improves the detection accuracy and efficiency, reduces the computational complexity while ensuring the detection accuracy, and improves the real-time performance of the algorithm. The neck network is constructed based on a multi-scale spatial attention mechanism, which can realize multi-scale feature processing, improve feature representation capability and target detection performance, and construct a detection model based on a spatial reconstruction unit and a channel reconstruction unit to reduce network complexity and improve detection accuracy. The detection method achieves lightweight while ensuring detection accuracy, and achieves an ideal balance between computational speed and recognition accuracy, providing strong support for safe production in coal mines.
[0042] Furthermore, the protective equipment on the miners is tested to determine whether they are wearing protective equipment, ensuring the safety of the miners when going down the mine, ensuring the detection accuracy and achieving the lightweight of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0045] Figure 1 A process diagram of a method for detecting personal protective equipment for miners provided in this application Figure 1 ;
[0046] Figure 2 A process diagram of a method for detecting personal protective equipment for miners provided in this application Figure 2 ;
[0047] Figure 3 A schematic flow chart of step S1 of a method for detecting personal protective equipment for miners provided in this application;
[0048] Figure 4 A flowchart of step S12 of a method for detecting personal protective equipment for miners provided in this application;
[0049] Figure 5 A flowchart of step S121 of a method for detecting personal protective equipment for miners provided in this application;
[0050] Figure 6 A flow chart of step S2 of a method for detecting personal protective equipment for miners provided in this application Figure 1 ;
[0051] Figure 7 A flow chart of step S2 of a method for detecting personal protective equipment for miners provided in this application Figure 2 ;
[0052] Figure 8 A schematic flow chart of step S3 of a method for detecting personal protective equipment for miners provided in this application. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0054] First, see Figure 1-2 The present invention discloses a method for detecting personal protective equipment for miners, which comprises the following steps:
[0055] Step S1: acquiring real-time miner image data, constructing a backbone network model based on a visual transformation network model, performing feature extraction on the real-time miner image data through the backbone network model, and extracting a first feature image;
[0056] Step S2: constructing a neck network model based on a multi-scale spatial attention mechanism, performing feature fusion and enhancement processing on the first feature image to obtain a second feature image;
[0057] Step S3: construct a detection model based on the space reconstruction unit and the channel reconstruction unit, detect and process the features in the second feature image, and obtain the detection result of the miner's personal equipment.
[0058] Specifically, in step S1, the backbone network model introduces a visual change network model, namely the EfficientVit model. By performing physical sign extraction processing through the backbone network model, different types of protective equipment can be distinguished, such as helmets of different colors and styles, to improve the accuracy of detection. In addition, the miner image may have layout changes caused by factors such as different shooting angles and lighting conditions. The backbone network model can better adapt to these changes and reduce the detection error caused by data changes. In addition, the backbone network model constructed based on the visual transformation network model reduces the computational complexity while ensuring the detection accuracy, and improves the real-time performance of the algorithm. In addition, the real-time miner image data can be photos captured from real-time video or taken by a camera.
[0059] Specifically, in step S2, in the detection of personal protective equipment, personal protective equipment may appear in different sizes in the image, for example, the personal protective equipment worn by workers in the distance is relatively small, while the personal protective equipment worn by workers in the near distance is larger. The neck network model constructed based on the multi-scale spatial attention mechanism (MSDA) can effectively handle this multi-scale situation and ensure that personal protective equipment of different sizes can be accurately detected. In complex industrial scenes, there may be a large amount of background interference. Spatial attention can highlight the area where the personal protective equipment is located and suppress the influence of irrelevant background, thereby improving the accuracy and efficiency of detection, dynamically adjusting the degree of attention to different spatial positions, and adaptively adjusting according to the actual position and distribution of personal protective equipment in the image to better adapt to different detection scenarios, that is, it can improve the network's feature representation ability and target detection performance.
[0060] Specifically, in step S3, the spatial reconstruction unit can optimize the spatial layout of the features, and can adjust the spatial features related to personal protective equipment in the feature map, so that the spatial information such as the shape and position of the personal protective equipment can be more accurately represented, which helps to improve the positioning accuracy of the detection frame. The channel reconstruction unit processes the channel dimension of the features, can redistribute the information between channels, emphasize the channel features related to personal protective equipment detection, improve the classification ability of different types of personal protective equipment, and improve the detection accuracy while reducing the network complexity.
[0061] It can be understood that constructing a backbone network model based on the visual transformation network model has better feature extraction capabilities, improves detection accuracy and efficiency, reduces computational complexity while ensuring detection accuracy, and improves the real-time performance of the algorithm. The neck network is constructed based on a multi-scale spatial attention mechanism, which can realize multi-scale feature processing, improve feature representation capabilities and target detection performance, and construct a detection model based on spatial reconstruction units and channel reconstruction units to reduce network complexity and improve detection accuracy. The detection method achieves lightweight while ensuring detection accuracy, and achieves an ideal balance between computing speed and recognition accuracy, providing strong support for safe production in coal mines.
[0062] Furthermore, the protective equipment on the miners is tested to determine whether they are wearing protective equipment, ensuring the safety of the miners when going down the mine, ensuring the detection accuracy and achieving the lightweight of the network.
[0063] See also Figure 2-5 , the step S1 specifically comprises the following steps:
[0064] Step S11: acquiring real-time miner image data, performing image segmentation and feature extraction processing on the miner image data, and obtaining a plurality of feature maps;
[0065] Step S12: inputting the multiple feature maps into the visual transformation network model in sequence to perform target feature detection and extraction processing to obtain a processed feature map;
[0066] Step S13: construct a variable-size pooling model, perform pooling and integration processing on all processed feature maps, and obtain a first feature image.
[0067] Specifically, the input dimension is 640×640×3 three-channel RGB image, and the image segmentation and feature extraction processing adopts Overlap PatchEmbed model. The Overlap Patch Coding Model can divide the image into overlapping small blocks and convert them into fixed-dimensional vectors, output multiple feature maps, and the EfficientVit model replaces the original YOLOv8 backbone network, which reduces the computational complexity while ensuring the detection accuracy and improves the real-time performance of the algorithm. The variable-size pooling model adopts the Fast Spatial Pyramid Pooling (SPPF) model, which helps to handle targets of different sizes, simplifies the pooling steps, and reduces the amount of calculation in the model. It is very critical for real-time target detection tasks because it can help accelerate the inference speed, while reducing the depth and width of the overall model and reducing the complexity of the calculation.
[0068] See also Figure 2-5 , the step S12 specifically includes the following steps:
[0069] Step S121: multiple feature maps are input into an efficient visual conversion module (EffcientViT Block) for grouping and self-attention calculation, and a self-attention feature map is output;
[0070] Step S122: the self-attention feature map is input into the efficient visual converter downsampling module (EfficientViTsubsample) for layering and downsampling processing to obtain downsampled feature maps at each level;
[0071] Step S123: The downsampled feature maps at each level are subjected to feature fusion processing to obtain processed feature maps.
[0072] Specifically, the feature map is grouped and self-attention is calculated in the efficient visual conversion module, and layered and downsampled in the efficient visual converter downsampling module. The feature map is processed multiple times in the efficient visual conversion module and the efficient visual converter downsampling module. The processing in the efficient visual converter downsampling module reduces the information loss during downsampling, thereby improving the memory efficiency of the model. The efficient visual conversion module allows each head to focus on different feature segmentation, reducing redundancy and increasing the diversity of attention maps, thereby increasing the capacity of the model without introducing additional parameters. While ensuring detection accuracy, the computational complexity is reduced and the real-time performance of the algorithm is improved.
[0073] See also Figure 2-5 , the step S121 specifically includes the following steps:
[0074] Step S1211: multiple feature maps are sequentially input into the feedforward neural network layer for depth-separable convolution processing and nonlinear transformation processing to obtain nonlinear transformation feature maps;
[0075] Step S1212: multiple nonlinear change feature maps are grouped, attention calculation is performed, and they are combined in multiple cascades to obtain a self-attention feature map.
[0076] Specifically, the token interaction layer and the efficient feed forward network (FFN) layer are calculated and processed in sequence. An additional token interaction layer is introduced before each feed forward network. This layer uses a depth-separable convolution. By introducing the inductive bias of local structural information, the token interaction layer can enhance the performance of the model. There are multiple token interaction layers and feed forward network layers, that is, the image needs to be processed by multiple token interaction layers and efficient feed forward network layers in succession. A cascaded group attention is used between the previous efficient feed forward network layer and the next token interaction layer to reduce the memory and time consumption caused by the self-attention layer, while allowing more FFN layers to interact to promote effective communication between different feature channels. The input feature map is divided into multiple parts, each of which is sent to an independent attention head. Each attention head calculates its self-attention map, and then the outputs of all heads are cascaded and projected back to the input dimension through a linear layer. This approach not only reduces computational redundancy in multi-head attention, but also improves model capacity by increasing network depth. By providing different input splits to each head, CGA is able to capture different aspects of the input features, thereby increasing the diversity of the attention map. Since each head only focuses on a portion of the input features, CGA reduces the number of input and output channels in the QKV layer, saving computational resources. The output of each head is added to the input of the next head in a cascaded manner, thereby gradually refining the feature representation, increasing the network depth, and improving the model capacity.
[0077] See also Figure 6-7 , the step S2 specifically comprises the following steps:
[0078] Step S21: extracting features from the first feature image to generate a series of feature maps of different scales;
[0079] Step S22: A feature fusion model combined with a multi-scale spatial attention mechanism is used to perform feature enhancement and feature fusion processing to obtain a second feature image.
[0080] Specifically, the first feature image is feature extracted, and a series of feature maps of different scales are generated according to the scale. For the feature maps of different scales, a part of them is enhanced for different times, and the feature maps after feature enhancement are fused with the feature maps of another part. After the multi-scale spatial attention mechanism is processed, the second feature images of multiple scales are obtained. Multi-scale hole attention (MSDA) is integrated into the neck network to improve the feature representation ability and target detection performance of the network, which can simulate the local and sparse image block interactions in a small range. These findings are derived from the analysis of the image block interactions in the global attention of the conversion model (EfficientVit model) at a shallow level. At a shallow level, the attention matrix has two key properties: locality and sparsity, which indicates that in shallow semantic modeling, most of the blocks far away from the query block are irrelevant, so there is a lot of redundancy in the global attention module. Therefore, using MSDA can effectively capture multi-scale semantic information and reduce the redundancy of the self-attention mechanism.
[0081] Multi-scale dilated attention (MSDA) involves processing feature maps of different scales into a linear transformation, then processing them through a SWDA model (Sliding window dilated attention), and finally fusing all the processed results in a combination module, and finally fusing the fused feature maps through a linear transformation. In the MSDA module, the SWDA model is a key technology for capturing semantic information at different scales. The main idea of the SWDA model is to obtain the query, key, and value of the feature map through linear projection, and then execute the SWDA model with different dilation rates in different heads to improve the processing efficiency and detection accuracy of the model.
[0082] See also Figure 6-7 , the step S22 specifically includes the following steps:
[0083] Step S211: the processed feature map is input into the convolution module for convolution processing to obtain a convolution feature map;
[0084] Step S212: The convolution feature map is divided into two parts, one part is directly output to the combination module, and the other part is input to the bottleneck module for feature extraction and enhancement processing to obtain an enhanced feature map;
[0085] Step S213: The enhanced feature map is divided into two parts again, one part is directly output to the combination module, and the other part is input to the bottleneck module again for feature extraction and enhancement processing to obtain a re-enhanced feature map;
[0086] Step S214: perform multiple branch processing and feature enhancement processing on the feature map that has been enhanced again, and finally output all the enhanced feature maps to the combination module for feature fusion to output a second feature image.
[0087] Specifically, the feature map is first convolved to extract features from the image, and the convolution process divides the feature map into two parts. One part is enhanced multiple times, and each enhancement process is branched. After all enhancement processes are completed, the enhanced feature map is combined with the branch image to obtain a second feature image.
[0088] It can be understood that the neck network model includes an upsampling module, a combination module, a bottleneck module and a cross-stage local model (C2F_MSDA) with dual fusion combined with a multi-source domain adaptation mechanism. In the C2F_MSDA model, the feature map is enhanced by performing feature conversion, branch processing, and feature fusion processing on the feature map, thereby improving the representation ability of the feature map. Feature conversion can help extract features of different levels and abstract degrees in the input feature map. Branch processing can increase the nonlinear ability and representation ability of the network, thereby improving the network's modeling ability for complex data. The features of different branches are spliced in the channel dimension to achieve feature fusion, enriching the expression ability of the features, and integrating MSDA into the C2F module to improve the network's feature representation ability and target detection performance.
[0089] See also Figure 8 , the step S3 specifically comprises the following steps:
[0090] Step S31: the second feature image is input into a spatial reconstruction unit (SRU) for spatial refinement processing to obtain a spatial refinement feature map;
[0091] Step S32: the spatially refined feature map is input into a channel reconstruction unit (CRU) for channel refinement processing to obtain a channel refinement feature map;
[0092] Step S33: The channel-refined feature map performs target detection through the detection layer to obtain the miner and personal equipment features and feature parameters in the image, wherein the feature parameters include the category of the target feature, the position of the target feature bounding box, and the confidence score;
[0093] Step S34: preset a judgment threshold, compare the characteristic parameter with the preset judgment threshold, and generate a detection result.
[0094] Specifically, after the second feature image is obtained through the processing of the neck network, the detection model is used to detect features, including two parts, the spatial reconstruction unit and the channel reconstruction unit. The feature X is input, and the spatial refinement feature Xw is first obtained by the spatial reconstruction unit operation, and then the channel refinement feature Y is obtained by the channel reconstruction unit operation. The spatial redundancy and channel redundancy between the features are utilized, and it can be seamlessly integrated into the architecture of any neural network model to reduce the redundancy between intermediate feature maps and enhance feature representation. The spatial reconstruction unit is used for spatial redundancy, and separation and reconstruction operations are utilized. The purpose of the separation operation is to separate the feature map with rich information from the feature map with less information corresponding to the spatial content. The scale factor in the group normalization (GN) layer is used to evaluate the information content of different feature maps. The channel reconstruction unit adopts the strategy of separation transformation fusion to reduce channel redundancy.
[0095] In a second aspect, the present invention discloses a miner's personal protective equipment detection system, which implements the above-mentioned miner's personal protective equipment detection method steps.
[0096] It can be understood that the system mentions that the backbone network model constructed based on the visual transformation network model has better feature extraction capabilities, improves the detection accuracy and efficiency, reduces the computational complexity while ensuring the detection accuracy, and improves the real-time performance of the algorithm. The neck network is constructed based on the multi-scale spatial attention mechanism, which can realize multi-scale feature processing and improve the feature representation ability and target detection performance of Saint Luo. The detection model constructed based on the spatial reconstruction unit and the channel reconstruction unit can reduce the network complexity and improve the detection accuracy. The detection method achieves lightweight while ensuring the detection accuracy, and achieves an ideal balance between computing speed and recognition accuracy, providing strong support for coal mine safety production.
[0097] In a third aspect, the present invention discloses a computer device, which includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0098] Memory, used to store computer programs;
[0099] The processor is used to implement the steps of the miner's personal protective equipment detection method when executing the program stored in the memory.
[0100] It can be understood that the processor of the device implements the program in the memory. In the program, the backbone network model constructed based on the visual transformation network model has better feature extraction ability, improves the detection accuracy and efficiency, reduces the computational complexity while ensuring the detection accuracy, and improves the real-time performance of the algorithm. The neck network is constructed based on the multi-scale spatial attention mechanism, which can realize multi-scale feature processing, improve feature representation ability and target detection performance, and construct a detection model based on the spatial reconstruction unit and the channel reconstruction unit to reduce network complexity and improve detection accuracy. The detection method achieves lightweight while ensuring detection accuracy, and achieves an ideal balance between computing speed and recognition accuracy, providing strong support for safe production in coal mines.
[0101] In a fourth aspect, the present invention discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a method for detecting personal protective equipment for miners.
[0102] It can be understood that the computer program stored in the storage medium can implement the method mentioned in the first aspect. The backbone network model constructed based on the visual transformation network model has better feature extraction ability, improves the detection accuracy and efficiency, reduces the computational complexity while ensuring the detection accuracy, and improves the real-time performance of the algorithm. The neck network is constructed based on the multi-scale spatial attention mechanism, which can realize multi-scale feature processing, improve feature representation ability and target detection performance, and construct a detection model based on the spatial reconstruction unit and the channel reconstruction unit. It can reduce network complexity and improve detection accuracy. The detection method is lightweight while ensuring detection accuracy, and achieves an ideal balance between computing speed and recognition accuracy, providing strong support for coal mine safety production.
[0103] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0104] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0105] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0106] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be connected, detachably connected, or integrated; it can be mechanically connected or electrically connected; it can be directly connected or indirectly connected through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0107] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may include that the first and second features are in direct contact, or may include that the first and second features are not in direct contact but are in contact through another feature between them. Moreover, a first feature being "above", "above" and "above" a second feature includes that the first feature is directly above and obliquely above the second feature, or simply indicates that the first feature is higher in level than the second feature. A first feature being "below", "below" and "below" a second feature includes that the first feature is directly below and obliquely below the second feature, or simply indicates that the first feature is lower in level than the second feature.
[0108] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms should not be understood as necessarily referring to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.
[0109] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
[0110] The above is a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A method for detecting personal protective equipment for miners, characterized in that: The following steps are included: Acquire real-time miner image data, construct a backbone network model based on the visual transformation network model, perform feature extraction on the real-time miner image data through the backbone network model, and extract a first feature image; Based on the multi-scale spatial attention mechanism, a neck network model is constructed, and feature fusion and enhancement processing are performed on the first feature image to obtain a second feature image; A detection model is constructed based on the space reconstruction unit and the channel reconstruction unit, and the features in the second feature image are detected and processed to obtain the detection results of the miner's personal equipment.
2. The method according to claim 1, characterized in that Real-time miner image data is obtained, a backbone network model is constructed based on a visual transformation network model, and features of the real-time miner image data are extracted through the backbone network model to extract a first feature image, specifically including the following steps: Acquire real-time miner image data, perform image segmentation and feature extraction on the miner image data, and obtain multiple feature maps; Inputting multiple feature maps into the visual transformation network model in sequence to perform target feature detection and extraction processing to obtain processed feature maps; A variable-size pooling model is constructed to perform pooling and integration processing on all processed feature maps to obtain the first feature image.
3. The method according to claim 2, characterized in that The step of sequentially inputting a plurality of feature maps into a visual transformation network model for target feature detection and extraction to obtain a processed feature map specifically includes the following steps: Multiple feature maps are input into the efficient visual conversion module for grouping and self-attention calculation, and the self-attention feature map is output; The self-attention feature map is input into the efficient visual converter downsampling module for layering and downsampling to obtain the downsampled feature maps at each level. The downsampled feature maps at each level are subjected to feature fusion processing to obtain the processed feature maps.
4. The method according to claim 3, characterized in that The multiple feature maps are input into the efficient visual conversion module for grouping and self-attention calculation, and the self-attention feature map is output, which specifically includes the following steps: Multiple feature maps are sequentially input into the feedforward neural network layer for nonlinear transformation processing to obtain nonlinear transformation feature maps; Multiple nonlinear change feature maps are grouped, attention calculation is performed, and they are combined in multiple cascades to obtain a self-attention feature map.
5. The method according to claim 1, characterized in that The method constructs a neck network model based on a multi-scale spatial attention mechanism, performs feature fusion and enhancement processing on the first feature image, and obtains a second feature image, specifically comprising the following steps: The first feature image is subjected to feature extraction to generate a series of feature maps of different scales; The feature fusion model combined with the multi-scale spatial attention mechanism performs feature enhancement and feature fusion processing to obtain the second feature image.
6. The method according to claim 5, characterized in that The feature fusion model combined with the multi-scale spatial attention mechanism performs feature enhancement and feature fusion processing to obtain a second feature image, specifically comprising the following steps: The processed feature map is input into the convolution module for convolution processing to obtain a convolution feature map; The convolution feature map is divided into two parts, one part is directly output to the combination module, and the other part is input to the bottleneck module for feature extraction and enhancement processing to obtain the enhanced feature map; The enhanced feature map is divided into two parts again, one part is directly output to the combination module, and the other part is input to the bottleneck module again for feature extraction and enhancement processing to obtain a further enhanced feature map; The enhanced feature map is subjected to multiple branch processing and feature enhancement processing, and finally all the enhanced feature maps are output to the combination module for feature fusion to output the second feature image.
7. The method according to claim 1, characterized in that The detection model is constructed based on the space reconstruction unit and the channel reconstruction unit, and the features in the second feature image are detected and processed to obtain the miner's personal equipment detection result, which specifically includes the following steps: The second feature image is input into the spatial reconstruction unit for spatial refinement processing to obtain a spatial refinement feature map; The spatially refined feature map is input into the channel reconstruction unit for channel refinement processing to obtain a channel refinement feature map; The channel refinement feature map performs target detection through the detection layer to obtain the miner and personal equipment features and feature parameters in the image. The feature parameters include the category of the target feature, the location of the target feature bounding box, and the confidence score; A preset judgment threshold is set, and the feature parameters are compared with the preset judgment threshold to generate a detection result.
8. A miner's personal protective equipment detection system, characterized in that: Implement the steps of the method for detecting miners' personal protective equipment according to any one of claims 1 to 7 above.
9. A computer device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the steps of any one of the miner's personal protective equipment detection methods of claims 1-7 when executing the program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting personal protective equipment for miners according to any one of claims 1 to 7 are implemented.