A dangerous behavior detection method for personnel in chemical enterprises based on improved YOLOv8
Patent Information
- Application Number
- CN202410226816.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-02-29
AI Technical Summary
然而,现有研究普遍偏向于某一或两类特定的危险行为,这在解决多样化的实际作业场景需求时存在一定挑战
[0026]有益效果:与现有技术相比,本发明的有益效果:本发明针对传统模型存在检测精度不高的问题,提出了一种组内点卷积残差模块,能够在通道维度上进行信息交流和特征融合;同时引入CSP_EALA模块替换原模型的C2f模块来改进YOLOv8网络的主干和颈部,有助于模型在保证轻量化的同时获得更加丰富的梯度流信息,以解决原模型对危险行为目标定位准确性差及抽烟打电话等小目标的判别能力较差问题;引入IPSPP模块至模型的颈部能够通过多尺度特征捕获和特征融合,有效提升危险行为的识别能力,使模型准确地理解并区分涉及不同尺度的危险行为目标;另外,使用DCSENet能够使模型能够更好地理解不同通道之间的相互作用,解决了通道关联性不足的问题,从而提高模型对复杂场景下的危险行为检测性能;综上本发明相较于原YOLOv8模型检测精度有所提升,有效提高了模型的检测效果、鲁棒性和泛化性能,为化工业领域的危险行为检测提供了有力支持。
Smart Images

Figure CN118038555B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hazardous behavior detection and identification, and in particular to a method for detecting hazardous behavior of personnel in chemical enterprises based on an improved YOLOv8. Background Technology
[0002] With rapid economic development, chemical enterprises have become an important part of modern society. However, this rapid development has also brought significant safety hazards to people's lives. These enterprises produce chemicals widely used in daily life; however, the chemical production process itself is accompanied by various potential risks, including fires, explosions, and leaks, posing serious threats to human life and the environment. Safety in production has a long history, and with industrial development and worsening working conditions, there is a growing demand for improved safety and hygiene. Safety accidents cause harm and loss to people's lives and property, making it imperative to strengthen safety management in chemical enterprises. Therefore, safety management and monitoring in chemical enterprises have become crucial to ensuring the safety of workers and the public.
[0003] However, traditional monitoring and safety management methods often rely on manual inspections, which are inefficient, costly, and subject to subjectivity. Given this issue, technologies such as computer vision and deep learning offer new opportunities to address this challenge. In recent years, object detection technology has made significant progress, particularly with the application of deep learning models such as YOLO and SSD. These models possess excellent real-time performance, enabling rapid identification and localization of different objects, including people and physical objects, in video streams. This invention aims to develop and improve a deep learning-based method for detecting hazardous behaviors in chemical enterprises. It is used to automatically detect and identify various hazardous behaviors, such as operating without a helmet, smoking, or making phone calls during work. However, existing research generally focuses on one or two specific types of hazardous behaviors, which presents challenges in addressing the diverse needs of real-world work scenarios. By introducing computer vision and deep learning technologies, this invention not only promises to improve overall workplace safety but also helps to effectively reduce the potential risks faced by chemical enterprises. Summary of the Invention
[0004] Purpose of the invention: This invention proposes a method for detecting hazardous behaviors of personnel in chemical enterprises based on an improved YOLOv8 model. Compared with the original YOLOv8 model, the detection accuracy is improved, effectively enhancing the model's detection performance, robustness, and generalization ability, thus providing strong support for hazardous behavior detection in the chemical industry.
[0005] Technical Solution: The present invention provides a method for detecting hazardous behaviors of personnel in chemical enterprises based on an improved YOLOv8, comprising the following steps:
[0006] (1) Data augmentation methods were used to augment the self-made dangerous behavior dataset, and the dataset was divided into training set, test set and validation set;
[0007] (2) Construct a dangerous behavior detection model based on improved YOLOv8; the dangerous behavior detection model introduces a cross-stage partial enhancement adaptive layer aggregation CSP_EALA module in the backbone and neck part of the YOLOv8 model; the CSP_EALA module is composed of an intra-group point convolutional residual module IPCR and an enhancement adaptive layer aggregation module EALA; an intra-group point convolutional spatial pyramid pooling IPSPP module is introduced in the neck part of the YOLOv8 target detection model to replace the SPPF module of the original YOLOv8 network; a dual convolutional squeeze excitation network DCSENet is introduced in the Neck part of YOLOv8;
[0008] (3) The features processed by the double convolutional squeeze excitation network are input into a series of CBS, CSP_EALA modules and upsampling operations to fuse the features; then the features fused by the Neck part are input into the Head part of the detection model to detect dangerous behavior targets;
[0009] (4) Input the training set of the dataset into the dangerous behavior detection model to obtain the trained dangerous behavior detection model;
[0010] (5) Evaluate the trained dangerous behavior detection model; input the dangerous behavior images to be detected into the trained dangerous behavior detection model and output the final detection results.
[0011] Furthermore, the dangerous behaviors described in step (1) include not wearing a helmet, smoking, and making phone calls.
[0012] Further, the implementation process of step (1) is as follows: the collected image data containing dangerous behaviors are manually labeled to ensure that the dangerous behaviors in each image are accurately labeled; random rotation, scaling, translation, brightness adjustment and mirroring are used to augment the images, and the augmented dataset is divided into training set, test set and validation set according to the proportion.
[0013] Furthermore, the CSP_EALA module described in step (2) replaces the last C2f module of Backbone.
[0014] Furthermore, the IPCR module in step (2) contains two independent branches, each employing different convolutional operations to capture features of different scales and spatial structures; one branch achieves high-level feature learning in the channel dimension through Pointwise convolution, while the other branch extracts local structural information in the spatial dimension through Group convolution; the outputs of the two branches are concatenated so that the network simultaneously takes into account both channel information and spatial structure, achieving richer and more comprehensive feature expression; Leaky_ReLU is introduced as the activation function and combined with Batch Normalization for normalization processing, accelerating training convergence and improving the stability of the model.
[0015] Further, step (2) EALA module is composed of two IPCR modules connected in series; in CSP_EALA module, the input is first processed by IPCR module to extract key features; then, the processed features are divided into two parts by split operation, one part is directly connected to the output to retain part of the original information; the other part is processed by multiple serial EALA modules to further refine and enhance the features; finally, all these parts are spliced together and the final output is obtained through IPCR module.
[0016] Further, the IPSPP module in step (2) includes an intra-group point convolutional residual module (IPCR), 1×1 and 3×3 convolutional layers, and three different sizes of MaxPools. First, the input features are divided into two parts and then fed into two branches respectively. In the first branch, the features are first extracted and transformed by 1×1 and 3×3 convolutional layers, and then the feature expression ability is enhanced by the IPCR module. The IPCR module includes group convolution, pointwise convolution, and batch normalization operations. Subsequently, the extracted features enter the max pooling layers of multiple sizes to capture feature information at different scales. Then, these features are further processed and fused through a series of convolutional operations to enhance their representation ability. The other branch also processes feature information through the IPCR module. Finally, the features extracted by the two branches are merged and further processed and adjusted by the final convolutional layer to generate an output feature map. The sizes of the max pooling layers are 5×5, 9×9, and 13×13, respectively.
[0017] Furthermore, the implementation process of introducing the dual convolutional squeeze excitation network DCSENet into the Neck part of YOLOv8 in step (2) is as follows:
[0018] For a given input tensor X, the convolution result after grouped convolution is passed through the activation function σ(·) to obtain the activation function tensor A′ for that part. gc =σ(Z′) gcSimilarly, the convolution result after point convolution is passed through an activation function to obtain A′. wpc =σ(Z′) wpc Next, the two are summed and then normalized to obtain the output Y′; then, a Squeeze operation is performed to perform average pooling on the feature map, resulting in a 1*1*C vector, i.e.:
[0019]
[0020] The vector z obtained in the previous step is processed by fully connected layers F1 and F2 to obtain the channel weight values W. After passing through two fully connected layers, different values in W represent the weight information of different channels, assigning different weights to the channels:
[0021] W=Γex(z,F)=μ(F2δ(F1z))
[0022] Where δ(·) represents the ReLU activation function and μ(·) represents the Sigmoid activation function, after two fully connected layers, we get 1*1*C; then, after the ScalerMultiplication operation, that is, multiplying the corresponding channels of the feature map W and the input feature map Y, we get the output feature map Output:
[0023] Output = Γsc(Y′) c W c ).
[0024] Furthermore, the Neck section in step (3) outputs shallow feature maps, mid-level feature maps, and deep feature maps respectively, which are used to detect the location and category of the target.
[0025] Furthermore, the Head part described in step (3) is responsible for predicting the target category, location, and confidence level. The detection task is divided into two branches: classification and regression. The classification branch includes a convolutional layer and a Sigmoid activation function, which are used to output the probability of each pixel corresponding to each category. The regression branch includes a convolutional layer and a Softmax activation function, which are used to output the probability distribution of each pixel corresponding to each dimension.
[0026] Beneficial Effects: Compared with existing technologies, the present invention offers the following benefits: Addressing the issue of low detection accuracy in traditional models, this invention proposes an intra-group point convolutional residual module capable of information exchange and feature fusion across channels. Simultaneously, it introduces the CSP_EALA module to replace the original model's C2f module, improving the backbone and neck of the YOLOv8 network. This helps the model obtain richer gradient flow information while maintaining lightweight design, resolving the original model's poor accuracy in locating dangerous behavior targets and its weak ability to distinguish small targets such as smoking and making phone calls. Introducing the IPSPP module to the model's neck enables multi-scale feature capture and fusion, effectively enhancing the recognition of dangerous behaviors and allowing the model to accurately understand and distinguish dangerous behavior targets at different scales. Furthermore, using DCSENet allows the model to better understand the interactions between different channels, solving the problem of insufficient channel correlation and thus improving the model's performance in detecting dangerous behaviors in complex scenarios. In summary, this invention improves detection accuracy compared to the original YOLOv8 model, effectively enhancing the model's detection performance, robustness, and generalization capabilities, providing strong support for dangerous behavior detection in the chemical industry. Attached Figure Description
[0027] Figure 1 This is a flowchart of the present invention;
[0028] Figure 2 This is a schematic diagram of the dangerous behavior detection model based on the improved YOLOv8.
[0029] Figure 3 The diagrams show the structure of Group convolution and Pointwise convolution, where (a) is Group convolution and (b) is Pointwise convolution.
[0030] Figure 4 This is a schematic diagram of the IPCR module.
[0031] Figure 5 The diagram below shows the CSP_EALA module, where (a) is a schematic diagram of the EALA structure, (b) is the definition of each icon, and (c) is a schematic diagram of the CSP_EALA network structure.
[0032] Figure 6 This is a schematic diagram of the IPSPP module structure;
[0033] Figure 7 This is a schematic diagram of the DCSENet network structure. Detailed Implementation
[0034] The present invention will now be described in further detail with reference to the accompanying drawings.
[0035] As attached Figure 1As shown, this invention provides a method for detecting hazardous behaviors of personnel in chemical enterprises based on an improved YOLOv8, with the following specific steps:
[0036] Step 1: Image Preprocessing. First, the original images in the hazardous behavior dataset are processed once to generate images of uniform size. Then, these uniform-sized images undergo a second processing step. The hazardous behavior dataset contains 5063 images of hazardous behaviors, categorized into four types: personnel, personnel wearing safety helmets, smoking during work, and making phone calls. These categories represent common hazardous behaviors within chemical enterprises, covering a variety of hazardous scenarios. To ensure the reliability and accuracy of the experiment, this invention adopts a strict dataset partitioning strategy, dividing the entire dataset into a training set, a test set, and a validation set in a 6:2:2 ratio.
[0037] The initial images of dangerous behaviors were enhanced using a mosaic technique. This process involves segmenting the input image into small patches and then recombining them to transform the overall image. Each patch undergoes translation, scaling, cropping, and stitching, while simultaneously being transformed in terms of hue, brightness, and saturation, ultimately forming a new image and thus constructing a dangerous behavior dataset. This operation simulates real-world visual changes, such as object rotation, scaling, and translation, helping the model extract more information from the original image and improving its generalization ability across different scenarios.
[0038] Step 2: As Figure 2As shown, an improved YOLOv8-based hazardous behavior detection model is constructed. First, an innovative Cross-Stage Partial Enhanced Adaptive Layer Aggregation (CSP_EALA) network is introduced into the backbone and neck region of the YOLOv8 model. The CSP_EALA network consists of an Intra-group Pointwise Convolutional Residual Module (IPCR) and an Enhanced Adaptive Layer Aggregation (EALA) network. Second, an Intra-group Pointwise Spatial Pyramid Pooling (IPSPP) module is introduced into the neck region of the YOLOv8 object detection model to replace the original SPPF module. Finally, a Double Convolution Squeeze-and-Excitation Network (DCSENet) is introduced into the neck region of the YOLOv8 object detection model.
[0039] The preprocessed image is input into the backbone of the network model. The input image passes through a series of CBS and C2f modules to extract local features from the input features; and in the last C2f operation of the backbone, the CSP_EALA module replaces the C2f module at that point.
[0040] like Figure 5 As shown, the CSP_EALA module is composed of the IPCR module and the EALA module; the IPCR module is composed of Group convolution and Pointwise convolution. Specifically, Group convolution is a variant of convolution operation that divides the input channels into several groups, and the channels in each group are only convolved with other channels in the same group. Thus, the parameters of the convolution kernel are divided into multiple groups, and each group is only convolved with a subset of the input channels. Specifically, if the number of input channels is C, and they are divided into G groups, then the number of channels in each group is C / G. This type of convolution reduces the computational cost and number of parameters per convolution operation compared to convergent convolution. Pointwise convolution is a convolution operation with a 1x1 kernel size. It considers only the value of a single pixel at each position of the input, mainly used for linear combinations in the channel dimension. Although it does not have a receptive field in space, it introduces a non-linear transformation in the channel dimension, which helps the network learn the relationships between different channels. A comparison of the two convolutions is shown in the figure below. Figure 3 As shown.
[0041] like Figure 4 As shown, the IPCR module structure comprises two independent branches, each employing different convolutional operations to capture features of varying scales and spatial structures. Specifically, one branch utilizes Pointwise Convolution to learn high-level features along the channel dimension, while the other branch extracts local structural information along the spatial dimension through Group Convolution. By concatenating the outputs of these two branches, the network can simultaneously consider both channel information and spatial structure, achieving richer and more comprehensive feature representation. To further enhance the network's nonlinearity and robustness, Leaky_ReLU is introduced as the activation function, combined with BatchNormalization for normalization, accelerating training convergence and improving model stability. This innovative network structure not only adapts to features of different scales and complexities but also possesses parameter efficiency, demonstrating excellent performance in scenarios with limited computational resources. The intra-group pointwise convolutional residual module enables information exchange and feature fusion along the channel dimension, facilitating the extraction of richer feature representations. These enhanced feature representation capabilities contribute to improving the accuracy of object recognition and localization in object detection tasks. The specific implementation process is as follows:
[0042] For input X∈RN×C in ×H in ×W in Input tensors, where N is the batch size and C is the input tensor. in H is the number of input channels. in and W in These are the input height and width, respectively; X∈RN×C out ×H out ×W out The output tensor, where Cout is the number of output channels, and Hout and Wout are the height and width of the output, respectively.
[0043] Group convolution operation: The convolution operation between the input tensor X and the group convolution kernel tensor Wgc is defined as follows:
[0044]
[0045] in, Xi represents the weight tensor of the grouped convolution kernel, and Xi represents the i-th group of the input tensor.
[0046] Pointwise convolution operation: The convolution operation between the input tensor X and the pointwise convolution kernel tensor Wpwc is defined as follows:
[0047] Zpwc=X*Wpwc
[0048] in, This represents the weight tensor of the point convolution kernel.
[0049] Element-wise addition: Element-wise addition of the activation feature map tensors Agc and Apwc yields the summed feature map tensor.
[0050] Asum = Agc + Apwc
[0051] The activation function is Leaky_Relu(x,α)=max(αx,x), where α=0.2. The activation feature tensor is obtained as ALeaky_relu=LaekyRelu(Asum,0.2).
[0052] Y=BatchNorm(ALeaky_relu,γ,β,ε)
[0053] Here, γ and β represent scaling and bias parameters, respectively, and ε prevents division by zero error. Batch normalization is applied to the activation feature map tensor to obtain the final output tensor Y.
[0054] The EALA module consists of two cascaded IPCR modules. The CSP_EALA module first performs an IPCR on the input, then splits the result into two parts. One part is directly connected to the output, and the other part is processed through multiple cascaded EALA modules. Finally, all these parts are concatenated and passed through another IPCR module to obtain the final output. This processing can obtain richer gradient flow information while maintaining lightweight design.
[0055] The preprocessed image is processed by the Backbone section, which outputs a series of feature maps, each corresponding to an image region at a different scale.
[0056] The features processed through the backbone are input into the neck of the network model for further processing of the feature maps passed from the backbone to extract richer semantic information. The extracted features are then fused from feature maps at different levels using the IPSPP network. For example... Figure 6As shown, the IPSPP network structure consists of multiple components, including an within-group pointwise convolutional residual module (IPCR), 1×1 and 3×3 convolutional layers, and three different sizes of max-pooling layers. First, the input features are divided into two parts and fed into two branches. In the first branch, the input data undergoes feature extraction and transformation through 1×1 and 3×3 convolutional layers, and then is processed by the IPCR module, which includes group convolution, pointwise convolution, and batch normalization to enhance feature representation. Next, the extracted features enter max-pooling layers of multiple sizes (5×5, 9×9, and 13×13) to capture feature information at different scales. Subsequently, these features are further processed and fused through a series of convolutional operations to further enhance their representational power. Simultaneously, the other input branch also processes feature information through the IPCR module. Finally, the features extracted from the two branches are merged and further processed and adjusted by the final convolutional layer to generate the output feature map. This network structure has the ability to capture and fuse features at multiple scales, significantly improving the detection performance of small targets (such as smoking, making phone calls, etc.).
[0057] In dangerous behavior detection tasks, multiple different dangerous behaviors may coexist within a single image frame, leading to missed or false detections due to occlusion. To address this issue, this invention introduces a DCSENet network into the neck region of the original model. Specifically, feature maps processed by the IPSPP module are fed into the DCSENet network, increasing the model's focus on channel importance and thus enhancing its performance. Figure 7 As shown, DCSENet is an innovative convolutional neural network architecture designed to enhance object detection performance and improve the model's ability to perceive key features. The core innovation of DCSENet lies in the fusion of its intra-group point convolutional residual module (IPCR) and squeeze excitation operation. The IPCR module, by introducing intra-group point convolution operations, emphasizes modeling the internal correlations of channels, enabling the network to capture more detailed information about the internal structure of the image. Simultaneously, to further enhance the model's focus on key information, Squeeze and Excitation operations are introduced to adaptively adjust the weights of each channel. Furthermore, through global average pooling and a series of fully connected layers, dynamic attention weights are generated for each channel, allowing the network to focus more intently on the most informative channels in the image. The specific implementation process is as follows:
[0058] For a given input tensor X, the convolution result after grouped convolution is passed through the activation function σ(·) to obtain the activation function tensor A′ for that part. gc =σ(Z′) gc Similarly, the convolution result after point convolution is passed through an activation function to obtain A′. wpc =σ(Z′) wpcNext, the two are summed and then normalized to obtain the output Y′. Then, a Squeeze operation is performed to perform average pooling on the feature map, resulting in a 1*1*C vector, i.e.:
[0059]
[0060] Next, the vector z obtained in the previous step is processed through fully connected layers F1 and F2 to obtain the desired channel weight values W. After passing through two fully connected layers, different values in W represent the weight information of different channels, assigning different weights to the channels.
[0061] W=Γex(z,F)=μ(F2δ(F1z))
[0062] Where δ(·) represents the ReLU activation function and μ(·) represents the Sigmoid activation function, after two fully connected layers, a 1*1*C structure is obtained. Then, a Scaler Multiplication operation is performed, which multiplies the corresponding channels of the input feature map Y with the feature map W to obtain the output feature map.
[0063] Output = Γsc(Y′) c W c ).
[0064] Step 3: The features processed by the dual convolutional compression excitation network are input into a series of CBS, CSP_EALA modules and upsampling operations to fuse the features and obtain richer gradient flow information. The Neck part outputs shallow feature maps, mid-level feature maps and deep feature maps respectively, which are used to detect the location and category of the target.
[0065] The features fused by the Neck part are then input into the Head part of the network model for target detection of dangerous behaviors. The Head part is the key part of the network structure responsible for predicting the target category, location, and confidence level. It divides the detection task into two branches: classification and regression. The classification branch includes a convolutional layer and a sigmoid activation function, which outputs the probability of each pixel for each category. The regression branch includes a convolutional layer and a softmax activation function, which outputs the probability distribution of each pixel for each dimension. This design allows the network to handle classification and regression tasks separately and ultimately obtain complete detection results.
[0066] Step 4: Train the hazardous behavior detection network. Input the images processed in Step 1 into the hazardous behavior detection network for training. Generate a hazardous behavior detection model. Save the weights of the trained hazardous behavior detection network to generate a hazardous behavior detection model for detecting hazardous behaviors in chemical industry scenarios.
[0067] A Windows 10-based computer equipped with an NVIDIA Tesla V100SXM2 graphics card with 16GB of video memory was used. PyTorch deep learning framework version 1.8.0 was chosen as the primary development tool, and CUDA 10.2.89 was utilized to accelerate the training process. Furthermore, Python interpreter version 3.8 was used, and the SGD optimizer was employed to tune the model parameters. Table 1 provides detailed configurations of the key experimental parameters.
[0068] Table 1 Experimental parameter configuration
[0069]
[0070] This invention uses precision (P), recall (R), and mean average precision (mAP) as the criteria for model evaluation. Specifically, a positive sample accurately identified as a positive sample is denoted as True Positive (TP), a negative sample accurately identified as a negative sample is denoted as True Negative (TN), a negative sample incorrectly identified as a positive sample is denoted as False Positive (FP), and a positive sample incorrectly identified as a negative sample is denoted as False Negative (FN).
[0071]
[0072]
[0073]
[0074] The dangerous behavior dataset was input into the YOLOv8 model and the improved YOLOv8 model respectively for comparative experiments, and the results are shown in the table below:
[0075]
[0076] This invention relates to safety systems for chemical enterprises and can be used to detect hazardous behaviors of workers in chemical enterprises.
[0077] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for detecting hazardous behaviors of personnel in chemical enterprises based on an improved YOLOv8, characterized in that, Includes the following steps: (1) Data augmentation methods were used to augment the self-made dangerous behavior dataset, and the dataset was divided into training set, test set and validation set; (2) Construct a dangerous behavior detection model based on improved YOLOv8; the dangerous behavior detection model introduces the cross-stage partial enhancement adaptive layer aggregation CSP_EALA module in the backbone and neck part of the YOLOv8 model; The CSP_EALA module is composed of an intra-group point convolutional residual module IPCR and an enhanced adaptive layer aggregation module EALA; an intra-group point convolutional spatial pyramid pooling (IPSPP) module is introduced into the neck of the YOLOv8 target detection model to replace the original YOLOv8 network's SPPF module; and a dual convolutional squeeze excitation network DCSENet is introduced into the Neck part of YOLOv8. (3) The features processed by the double convolutional compression excitation network are input into a series of CBS, CSP_EALA modules and upsampling operations to fuse the features; The features fused through the Neck section are then input into the Head section of the detection model for target detection of dangerous behaviors; (4) Input the training set of the dataset into the dangerous behavior detection model to obtain the trained dangerous behavior detection model; (5) Evaluate the trained dangerous behavior detection model; Input the image of the dangerous behavior to be detected into the trained dangerous behavior detection model, and output the final detection result; In step (2), the CSP_EALA module replaces the last C2f module of Backbone; Step (2) The IPCR module contains two independent branches, each employing different convolutional operations to capture features of different scales and spatial structures. One branch uses Pointwise convolution to learn advanced features in the channel dimension, while the other branch uses Group convolution to extract local structural information in the spatial dimension. The outputs of the two branches are concatenated so that the network can simultaneously take into account both channel information and spatial structure, achieving richer and more comprehensive feature representation. Leaky_ReLU is introduced as the activation function and combined with Batch Normalization for normalization, accelerating training convergence and improving the stability of the model. Step (2) The EALA module consists of two IPCR modules connected in series. In the CSP_EALA module, the input is first processed by the IPCR module to extract key features. Then, the processed features are divided into two parts by the split operation. One part is directly connected to the output to retain part of the original information. The other part is processed by multiple serial EALA modules to further refine and enhance the features. Finally, all these parts are spliced together and the final output is obtained through the IPCR module. Step (2) The IPSPP module includes an intra-group point convolutional residual module (IPCR), 1×1 and 3×3 convolutional layers, and three different sizes of MaxPools. First, the input features are divided into two parts and then fed into two branches. In the first branch, the features are first extracted and transformed by 1×1 and 3×3 convolutional layers, and then the feature expression ability is enhanced by the IPCR module. The IPCR module includes group convolution, point-wise convolution, and batch normalization operations. Subsequently, the extracted features enter the max pooling layers of multiple sizes to capture feature information at different scales. Then, these features are further processed and fused through a series of convolutional operations to enhance their representation ability. The other branch also processes feature information through the IPCR module. Finally, the features extracted from the two branches are merged and further processed and adjusted by the final convolutional layer to generate the output feature map; the sizes of the max pooling layers are 5×5, 9×9 and 13×13, respectively. The implementation process of introducing the dual convolutional squeeze excitation network DCSENet into the Neck part of YOLOv8 in step (2) is as follows: For a given input tensor The convolution result after grouped convolution is then activated by an activation function. , thus obtaining the activation function tensor Similarly, the convolution result after point convolution is passed through an activation function to obtain... Next, the two are summed and then normalized to obtain the output. Then, a Squeeze operation is performed to average pool the feature map, resulting in... The vector, that is: Through the fully connected layer and The obtained vector z is processed to obtain the channel weight values w. After passing through two fully connected layers, different values in w represent the weight information of different channels, thus assigning different weights to the channels. in, Represents the activation function ReLU. The activation function is Sigmoid. After passing through two fully connected layers, we get... Then, after the Scaler Multiplication operation, the weight w is compared with... Perform multiplication on the corresponding channels to obtain the output feature map: 。 2. The method for detecting hazardous behaviors of personnel in chemical enterprises based on improved YOLOv8 according to claim 1, characterized in that, The dangerous behaviors described in step (1) include not wearing a helmet, smoking, and making phone calls.
3. The method for detecting hazardous behaviors of personnel in chemical enterprises based on improved YOLOv8 according to claim 1, characterized in that, The implementation process of step (1) is as follows: the collected image data containing dangerous behaviors are manually labeled to ensure that the dangerous behaviors in each image are accurately labeled; random rotation, scaling, translation, brightness adjustment and mirroring are used to augment the images, and the augmented dataset is divided into training set, test set and validation set according to the proportion.
4. The method for detecting hazardous behaviors of personnel in chemical enterprises based on improved YOLOv8 according to claim 1, characterized in that, In step (3), the Neck section outputs shallow feature maps, mid-level feature maps, and deep feature maps respectively, which are used to detect the location and category of the target.
5. The method for detecting hazardous behaviors of personnel in chemical enterprises based on improved YOLOv8 according to claim 1, characterized in that, The Head section in step (3) is responsible for predicting the target category, location, and confidence level. The detection task is divided into two branches: classification and regression. The classification branch includes a convolutional layer and a Sigmoid activation function, which are used to output the probability of each pixel corresponding to each category. The regression branch includes a convolutional layer and a Softmax activation function, which are used to output the probability distribution of each pixel corresponding to each dimension.
Citation Information
Patent Citations
Falling person target detection method based on optimized YOLOv8s network structure
CN116863539A
Road target detection method and system based on improved YOLOv8
CN117037119A