Lightweight Object Detection Method and System Based on Improved YOLOv8
The SHM-YOLO model addresses the challenge of detecting small targets in complex tower crane environments by using a lightweight architecture with an attention mechanism and optimized loss function, improving detection accuracy and efficiency.
Patent Information
- Application Number
- CN202411417998.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-10-11
AI Technical Summary
The images collected by the monitoring equipment on the top of the tower set are difficult to accurately locate and identify small targets in construction site scenarios due to factors such as small size, low resolution, complex background and extreme weather interference, and the detection effect is poor.
The interference data set is constructed, and the extreme weather is simulated by introducing factors such as Gaussian noise, and a lightweight object detection model SHM-YOLO is established. The HWD convolution module, C2f feature fusion module and SEAM attention mechanism are adopted, and the Inner-MPDIoU loss function optimization model is combined to improve detection performance.
Under extreme weather conditions, the detection accuracy of small targets and the lightweighting of the model are improved, adapting to the complex environment on the top of the tower set, and enhancing the ability to identify small targets.
Smart Images

Figure CN119323668B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the technical field of object detection, and in particular, to a lightweight object detection method and system based on improved YOLOv8. Background Art
[0002] Construction sites are high-risk locations for safety accidents. As large mechanical equipment, tower cranes often cause accidents when workers operate around them due to hook collisions. Collisions can also occur between multiple tower cranes due to scheduling problems, affecting the safety of the workers at the bottom. These phenomena are more likely to occur in sudden harsh rain and snow weather. At the same time, the operating location of tower crane operators is at the top of the equipment, and their vision is blocked during operation, making it difficult to observe whether the workers at the bottom of the tower crane are in a safe working environment. Transmitting the real-time images captured by the camera on the top of the tower crane to the object detection model and visualizing the detection results, the tower crane operator can grasp the positions of the personnel within the detection range in real time and avoid and handle potential safety problems in a timely manner.
[0003] Object detection methods in computer vision have been widely used after long-term development, such as the YOLO series. However, the excellent performance of these networks is mostly affected by the quality of the training set. The detected objects in these datasets usually have characteristics such as a large proportion of large target objects, clear images, and simple backgrounds, so these targets are easily detected. In real-world object detection scenarios, due to the influence of different image specifications, acquisition perspectives, lighting factors, weather factors, etc., the object detection model cannot achieve its original effect. The monitoring equipment on the tower crane is located at a high altitude, and the captured images have the characteristics of small size, low resolution, and little information. For small target detection, the characteristics of the targets themselves are few, and it is difficult to extract favorable feature information during network training. After multiple downsampling and pooling operations, most of the feature information of small targets will be deleted, resulting in the model being difficult to accurately locate and identify small targets. At the same time, the image background in the construction site scenario is extremely complex, and dust will also make the detected images more blurred. These factors have a great interference on the detection. The angles and sizes of the images captured when the tower crane is lifting and rotating are also very different, which leads to unbalanced samples of the detected images and increases the difficulty of detection. Sudden extreme rain and snow weather may cover the detection targets due to perspective reasons, resulting in missed detections and misclassifications. In summary, for the scenario of object detection of the images of the workers under the tower crane captured by the monitoring equipment on the top of the tower crane, the detection targets are interfered and affected by multiple aspects and factors, and the detection difficulty is far beyond that of the general scenario of object detection, posing higher requirements for the detection method and model. Summary of the Invention
[0004] One or more embodiments of this specification provide a lightweight object detection method based on improved YOLOv8, including:
[0005] Collect construction worker operation images through the tower crane top monitoring device, obtain a normal dataset under normal working conditions, construct an interference dataset by introducing interference factors, label and classify the normal dataset and the interference dataset, and divide them into a training set, a validation set, and a test set according to a predetermined ratio;
[0006] Establish a lightweight object detection model SHM-YOLO. The SHM-YOLO network architecture includes a backbone network Backbone, a neck network Neck, and a head network Head. The backbone network Backbone includes a combined module HWD-C2f of an HWD convolutional module and a C2f feature fusion module. An attention mechanism module SEAM is introduced in the neck network Neck and combined with the HWD convolutional module and the C2f feature fusion module. SEAM is placed after each C2f feature fusion module;
[0007] Set training parameters, use the training set to train the lightweight object detection model SHM-YOLO, use the Inner-MPDIoU loss function to optimize the constructed object detection model, and evaluate the object detection results through evaluation metrics;
[0008] Use the test set to test the effect of the trained object detection model.
[0009] Further, the specific method for constructing the interference dataset by introducing interference factors is:
[0010] Add Gaussian noise to the normal dataset to construct the interference dataset.
[0011] Further, the HWD convolutional module includes a lossless feature encoding module and a feature representation learning module;
[0012] The lossless feature encoding module is used to perform Haar wavelet transform on the input image, convert the image features and retain the image information while reducing the spatial resolution;
[0013] The feature representation learning module consists of a standard convolutional layer, a batch normalization layer, and a ReLU activation layer. The 1x1 convolutional layer is used to process the number of channels of the feature map to keep the input and output channels of the front and back layers consistent. The input image further reduces redundant information through the batch normalization layer and the Relu activation layer.
[0014] Further, the attention mechanism module SEAM is specifically used for:
[0015] In the SEAM attention mechanism module, the input image first passes through the Channel and Spatial Mixing Module (CSMM) to adjust the channel and spatial dimensions of the image. The Patch Embedding layer can be changed according to the set size and number of channels, and then passes through the GELU loss function. After being processed by the BatchNorm layer, depthwise separable convolution is used to learn the importance of different channels, retaining important channels and removing irrelevant parameters.
[0016] The input before depthwise separable convolution and the input after channel separation are combined through 1x1 convolution. After the input image is processed by three CSMM modules of different scales, the main channels of the occluded image in different dimensions are highlighted while compensating for the detail loss between channels. Then, Average Pooling is performed to remove redundant information, and all information is fused through two fully connected layers to enhance the connection of channel information. Finally, the output of the SEAM module is used as the attention and multiplied by the original features.
[0017] Furthermore, when training the constructed object detection model, the loss function used is the MPDIoU loss function Inner-MPDIoU that fuses Inner-IoU, as shown below:
[0018] L inner-MPDIoU = L MPDIoU + IoU - IoU inner .
[0019] Furthermore, the set training parameters include:
[0020] Set the number of training epochs epoch = 300, the batch size batch size = 16, set the initial learning rate of the network to 0.001, use the Stochastic Gradient Descent (SGD) optimizer, and uniformly scale the size of the input pictures to 640×640.
[0021] Furthermore, the evaluation metrics include precision, recall, mAP@0.5, Params, and GFLOPs; the precision is used to evaluate the ratio of positive samples that are correctly identified as positive samples among the identified pictures. The higher the precision, the better the detection effect of the detection model; the recall is used to evaluate the proportion of samples that are actually positive among all samples detected as positive; the mAP is used to evaluate the average value of the average precision (AP) of multiple classes. The larger the value of mAP, the better the performance of the detection model; Params and GFLOPs are used to evaluate the lightweight characteristics of the improved model.
[0022] One or more embodiments of this specification provide a lightweight object detection system based on improved YOLOv8, including:
[0023] Data processing module: It is used to collect operation images of construction workers through the monitoring device at the top of the tower crane, obtain a normal data set under normal working conditions, construct an interference data set by introducing interference factors, label and classify the normal data set and the interference data set, and divide them into a training set, a validation set and a test set according to a predetermined ratio;
[0024] Model construction module: It is used to establish a lightweight object detection model SHM-YOLO. The SHM-YOLO network architecture includes a backbone network Backbone, a neck network Neck and a head network Head. The backbone network Backbone includes a combined module HWD-C2f composed of an HWD convolution module and a C2f feature fusion module. An attention mechanism module SEAM is introduced into the neck network Neck and combined with the HWD convolution module and the C2f feature fusion module. SEAM is placed after each C2f feature fusion module;
[0025] Model training module: It is used to set training parameters, train the lightweight object detection model SHM-YOLO using the training set, optimize the constructed object detection model using the Inner-MPDIoU loss function, and evaluate the object detection results through evaluation metrics;
[0026] Model testing module: It is used to test the effect of the trained object detection model using the test set.
[0027] One or more embodiments of this specification provide an electronic device, including:
[0028] A processor; and,
[0029] A memory arranged to store computer-executable instructions, which when executed cause the processor to implement the steps of the above-mentioned lightweight object detection method based on improved YOLOv8.
[0030] One or more embodiments of this specification provide a storage medium for storing computer-executable instructions, which when executed implement the steps of the above-mentioned lightweight object detection method based on improved YOLOv8.
[0031] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0033] Figure 1 Flowchart of a lightweight object detection method based on improved YOLOv8 provided for one or more embodiments of this specification;
[0034] Figure 2 Schematic diagram of the SHM-YOLO network structure of a lightweight object detection method based on improved YOLOv8 provided for one or more embodiments of this specification;
[0035] Figure 3 Schematic diagram of the HWD module structure of a lightweight object detection method based on improved YOLOv8 provided for one or more embodiments of this specification;
[0036] Figure 4 Schematic diagram of the SEAM attention mechanism module structure of a lightweight object detection method based on improved YOLOv8 provided for one or more embodiments of this specification;
[0037] Figure 5 Schematic diagram of the MPDIoU structure of a lightweight object detection method based on improved YOLOv8 provided for one or more embodiments of this specification;
[0038] Figure 6 Schematic diagram of the composition of a lightweight object detection system based on improved YOLOv8 provided for one or more embodiments of this specification;
[0039] Figure 7 Schematic diagram of the structure of an electronic device provided for one or more embodiments of this specification. Detailed implementation manners
[0040] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification in conjunction with the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.
[0041] Method Embodiment
[0042] According to an embodiment of the present invention, a lightweight object detection method based on improved YOLOv8 is provided. Figure 1 The flowchart of a lightweight object detection method based on improved YOLOv8 provided for one or more embodiments of this specification is as follows. Figure 1 As shown, the lightweight object detection method based on improved YOLOv8 according to an embodiment of the present invention specifically includes:
[0043] S1. Collect construction worker operation images through the monitoring device at the top of the tower crane to obtain a normal data set under normal working conditions. By introducing interference factors, construct an interference data set, mark and classify the normal data set and the interference data set, and divide them into a training set, a validation set, and a test set according to a predetermined ratio.
[0044] In this embodiment, the operation images of construction workers in a non-interference state under normal weather are captured by the monitoring device at the top of the tower crane. These images represent the normal data set OP under normal working conditions. In order to simulate various interference factors that may be encountered in the real environment, such as weather changes, light condition changes, or other external interferences, these interference factors are introduced. In this embodiment, Gaussian noise is used in the normal data set to simulate the operation images captured under extreme rain and snow weather, and a data set containing interference factors is constructed, which is called the interference data set.
[0045] Specifically, first, add Gaussian noise with a probability percentage of 70% to the normal data set. Second, add Gaussian blur with a radius of 2 pixels and set the color level to 75. Finally, add dynamic blur with an angle of 70 degrees and a distance of 50 pixels to construct the interference data set RSOP. The OP data set and the RSOP data set consist of the same number of pictures, a total of 4118 pictures. The YOLO format is used to annotate and classify the data set so that the training model can recognize and understand these objects and behaviors, and divide the training set, validation set, and test set according to a ratio of 8:1:1.
[0046] S2. Establish a lightweight object detection model SHM-YOLO. The SHM-YOLO network architecture includes a backbone network Backbone, a neck network Neck, and a head network Head. The backbone network Backbone includes a combined module HWD-C2f composed of an HWD convolutional module and a C2f feature fusion module. An attention mechanism module SEAM is introduced in the neck network Neck and combined with the HWD convolutional module and the C2f feature fusion module. SEAM is placed after each C2f feature fusion module.
[0047] The SHM-YOLO network structure is a lightweight, high-performance target detection network based on the YOLOv8n network structure and the addition of HWD downsampling in the Haar wavelet. At the same time, the SEAM attention mechanism is introduced to further improve the detection performance, and the Inner-MPDIoU loss function is used to improve the convergence speed and detection accuracy.
[0048] The SHM-YOLO network structure is as follows Figure 2 As shown in the figure, it includes a backbone network Backbone, a neck network Neck and a head network Head. The collected image data is input into the backbone network Backbone. The backbone network includes a combination module HWD-C2f of a HWD convolution module and a C2f feature fusion module. C2, C3, C4, and C5 are output feature factor maps, and the number of output channels is 128, 256, 512, and 1024, respectively. SPPF is a spatial pyramid pooling layer for aggregating features at multiple scales. The attention mechanism module SEAM is introduced into the neck network Neck. The Attention-Neck part is an efficient aggregation network that can pay more attention to image details, including an HWD convolution module, an attention mechanism module SEAM, and a C2f feature fusion module. The attention mechanism module SEAM is placed after each C2f feature fusion module to further focus on the image details after feature fusion. The Attention-Neck part inputs 3 feature maps and outputs 3 feature maps at the same time, and the number of input and output channels remains consistent.
[0049] Haar wavelet transform is a simple and efficient vector transformation method, which can be divided into two steps: decomposition and reconstruction. Haar wavelet transform can perform lossless decomposition and reconstruction of image features at multiple scales in image processing. The more image detail features are retained in the encoder module, the more detection features can be extracted later, thereby improving the detection performance. The HWD module proposed based on Haar wavelet transform retains the original spatial feature information as much as possible while using Haar wavelet transform to increase the number of channels of feature mapping. Then, in order to reduce redundant information features, the resolution of feature mapping is reduced before convolution operation is performed on representative learning features. The structure of HWD module is as follows: Figure 3 As shown in the figure, the module mainly consists of two parts: a feature representation learning module and a lossless feature encoding module. The lossless feature encoding module uses Haar wavelet transform to convert image features and retain image information while reducing spatial resolution. The feature representation learning module consists of a standard convolutional layer, a batch normalization layer, and a ReLU activation layer. H, W, and C are the height, width, and number of channels of the image, respectively.
[0050] In the lossless feature encoding module, the input image is first subjected to Haar wavelet transform. Through two decompositions and reconstructions, the image features are transformed while retaining the image information and reducing the spatial resolution at the same time. The Haar wavelet transform can be understood as a band-pass filter that only allows signals with frequencies close to those of the wavelet basis function to pass through. The processing process is as follows: First, calculate the average value of adjacent pixel pairs in the input image to obtain a new image with a relatively low resolution. The resolution of the new image is In this step, some information of the image has been lost. Secondly, in order to be able to reconstruct the original image composed of original pixels from the image composed of average pixels, it is necessary to store some detail coefficients of the image so as to retrieve the lost information during reconstruction. The method used in this step is to divide the difference of this pixel pair by 2. Finally, splice the average pixels in the first step and the detail information in the second step. In this pixel pair, the overall information and detail information of the image are included. Subsequently, repeat the first two steps to decompose the overall information again to obtain the second-level decomposition result. During this process, images of various resolutions can be reconstructed from the recorded pixel pairs, and no information is lost during the transformation because the original image can be reconstructed from the recorded data. After the Haar wavelet transform, the image generates low-frequency information and high-frequency information. The low-frequency information stores the contour information and approximate information of the picture, corresponding to the calculation of the mean value. The high-frequency information stores the detail information, local information, and contains noise, corresponding to the calculation of the difference. Hhigh is a high-pass filter that allows high-frequency information to pass through, and Hlow is a low-pass filter that allows low-frequency information to pass through.
[0051] ↓2 indicates downsampling applied to the approximation component and the detail component. After the two-dimensional image is filtered by the high-pass and low-pass filters twice, four components are generated. The resolution of each component becomes half of the original, and the number of channels increases to four times the original. The Haar wavelet transform can encode some information in the spatial dimension into the channel dimension without losing information.
[0052] In the feature representation learning module, the 1x1 convolutional layer is used to process the number of channels of the feature map to keep the input and output channels of the front and back layers consistent. After the image passes through the batch normalization layer and the Relu activation layer, redundant information can be further reduced, facilitating more efficient feature learning for subsequent layers.
[0053] Such as Figure 4As shown, on the left is the overall structure diagram of the SEAM module, and on the right is the CSMM (Channel and Spatial Mixing Module), which is used to learn the correlation between the spatial dimension and channels. In the attention mechanism module SEAM, the input image first passes through the CSMM of the channel and spatial mixing module to adjust the channel and spatial dimensions of the image. The PatchEmbedding layer can change according to the set size and number of channels, and then passes through the GELU loss function to enhance the fitting; after passing through the BatchNorm layer, the depthwise separable convolution is used to learn the importance of different channels, retain the important channels, and remove the irrelevant parameters;
[0054] To reduce the loss of detailed information, the input before the depthwise separable convolution and the input after channel separation are combined through 1x1 convolution, so that the attention mechanism can learn the main feature information and complete detailed information, that is, the main features and detailed information of the occluded image. After the input image is processed by three CSMM modules of different scales, it highlights the main channels of the occluded image in different dimensions while making up for the detailed loss between channels. Then, AveragePooling is used to remove redundant information, and then two fully connected layers are used to fuse all information to enhance the connection of channel information. Finally, the output of the SEAM module is multiplied by the attention and the original features, so that the model can more effectively process the occluded targets and is also more suitable for small target detection.
[0055] In terms of the improvement idea of SHM-YOLO, since convolutional layers are filled in various parts of the network structure, using more lightweight and efficient convolutions can effectively reduce the overall network volume and optimize the network performance at the same time. The Neck layer is located between the backbone network and the detection head in the overall network and can perform feature fusion and enhancement. Adding an attention mechanism in the Neck can further explore the features of the detection object and help small target detection achieve better results. SHM-YOLO also optimizes the loss function to enable the network training and inference to have better speed and accuracy.
[0056] S3. Set the training parameters, use the training set to train the lightweight object detection model SHM-YOLO, use the Inner-MPDIoU loss function to optimize the constructed object detection model, and evaluate the object detection results through evaluation indicators.
[0057] At present, the loss function of Bounding Box Regression (BBR) is constantly evolving. However, the existing means of accelerating convergence based on IoU in BBR still remain at adding new loss terms, resulting in the neglect of the limitations of IoU itself. IoU cannot be adaptively adjusted according to different detection model performances and different detection tasks, nor does it have a certain generalization ability. Distinguishing different regression samples and calculating losses using auxiliary bounding boxes of different scales can effectively accelerate the bounding box regression process. For high-IoU samples, using smaller auxiliary bounding boxes to calculate losses can accelerate convergence, while larger auxiliary bounding boxes are suitable for low-IoU samples. When the predicted box and the ground truth box have the same aspect ratio, better regression results can be obtained. However, in most cases, the aspect ratios are different, and the number loss function cannot achieve good results. MPDIoU is a metric based on the intersection over union. Inspired by the geometric characteristics of the bounding box, it uses the coordinates of the upper left and lower right corners to define a unique rectangle, minimizing the distance between the upper left and lower right corners of the predicted box and the ground truth box. The calculation formula is as follows:
[0058]
[0059]
[0060] In the above formula, A and B are any two input bounding boxes, and w and h are the width and height of the input image respectively. For bounding boxes A and B, represents the coordinates of the upper left and lower right points of A. Similarly for bounding box B. MPDIoU can evaluate the similarity degree between two bounding boxes and can better adapt to the overlapping or non-overlapping phenomena in bounding box regression. The structural diagram of the loss function based on MPDIoU is as Figure 5 shown. The blue box is the ground truth box, and the red box is the predicted box. The calculation formula of the loss function based on MPDIoU is as follows:
[0061]
[0062]
[0063] In the above formula, the coordinates of the upper left and lower right corners of the ground truth box are d1 and d2 are the distances between the upper left and lower right corners of the ground truth box and the predicted box. A gt is the area of the ground truth box region, and A prd is the area of the predicted box region. Subsequently, the intersection point β of the ground truth box and the predicted box is calculated, and the formula is as follows:
[0064]
[0065]
[0066]
[0067]
[0068] In this embodiment, when training the constructed object detection model, the number of training epochs is set to epoch = 300, the batch size is set to batch size = 16, the initial learning rate of the network is set to 0.001, the Stochastic Gradient Descent (SGD) optimizer is used, the size of the input images is uniformly scaled to 640×640, and the loss function used is the MPDIoU loss function Inner-MPDIoU that incorporates Inner-IoU. The ratio scaling factor in Inner-IoU is used to control the size of the auxiliary bounding boxes for calculating the loss. MPDIoU defines a unique rectangle using the coordinates of the upper left and lower right corners of the rectangle, and minimizes the distance between the upper left and lower right corners of the predicted bounding box and the ground truth bounding box. Compared with the IoU loss, when the ratio is less than 1 and the size of the auxiliary bounding box is smaller than the actual bounding box, the effective range of regression is smaller than the IoU loss, but the absolute value of the gradient is greater than the gradient obtained from the IoU loss, which can accelerate the convergence of high-IoU samples. On the contrary, when the ratio is greater than 1, the larger-scale auxiliary bounding box expands the effective range of regression and has an enhanced effect on the regression of low-IoU samples, as shown below:
[0069] L inner-MPDIoU = L MPDIoU + IoU - IoU inner ;
[0070] The fused Inner-MPDIoU loss function has the advantages of strong generalization adaptability, high precision improvement, and fast regression speed, and can further enhance the detection effect in small object detection.
[0071] In this embodiment, the evaluation metrics for evaluating the object detection results include precision, recall, mAP@0.5, Params, and GFLOPs. The precision is used to evaluate the ratio of positive samples that are correctly identified as positive samples among the identified images. The higher the precision, the better the detection effect of the detection model. The recall is used to evaluate the proportion of positive samples that are actually positive among all samples detected as positive. The mAP is used to evaluate the average value of the average precision (AP) of multiple categories. The larger the value of mAP, the better the performance of the detection model. mAP@0.5 indicates the value of mAP when the IOU threshold is 0.5. When the IOU between the predicted bounding box and the annotated bounding box is greater than 0.5, the object is considered to be predicted correctly, and on this premise, the average value of AP, mAP, is calculated. The Params and GFLOPs are used to evaluate the lightweight characteristics of the improved model. GFLOPs can measure the amount of floating-point operations, and thus compare the computational complexity of different models. The larger the GFLOPs, the more complex the model, and at the same time, the more data it can "handle", enabling it to complete complex tasks, but with higher requirements for hardware computing resources. A smaller GFLOPs value indicates a lower computational requirement for the task or model, which may be more efficient or lightweight. Most of the hardware devices on tower cranes are industrial control computers, single-chip microcomputers, etc., and the hardware computing level is not high. Therefore, a detection model with low GFLOPs is more suitable for use on tower crane equipment.
[0072] S4. Use the test set to test the effect of the trained object detection model.
[0073] System embodiment
[0074] According to an embodiment of the present invention, a lightweight object detection system based on improved YOLOv8 is provided. Figure 6 This is a schematic diagram of the composition of a lightweight object detection system based on improved YOLOv8 provided for one or more embodiments of this specification, as Figure 6 shown. The lightweight object detection system based on improved YOLOv8 according to an embodiment of the present invention specifically includes:
[0075] The data processing module 60: is used to collect the operation images of construction workers through the monitoring equipment at the top of the tower crane, obtain the normal data set under normal working conditions, construct a interference data set by introducing interference factors, label and classify the normal data set and the interference data set, and divide them into a training set, a validation set, and a test set according to a predetermined ratio;
[0076] Model construction module 62: It is used to build a lightweight object detection model SHM-YOLO. The SHM-YOLO network architecture includes a backbone network Backbone, a neck network Neck, and a head network Head. The backbone network Backbone includes a combined module HWD-C2f of an HWD convolution module and a C2f feature fusion module. An attention mechanism module SEAM is introduced in the neck network Neck and combined with the HWD convolution module and the C2f feature fusion module. SEAM is placed after each C2f feature fusion module;
[0077] Model training module 64: It is used to set training parameters, train the lightweight object detection model SHM-YOLO using a training set, optimize the constructed object detection model using an Inner-MPDIoU loss function, and evaluate the object detection results through evaluation metrics;
[0078] Model testing module 66: It is used to test the effect of the trained object detection model using a test set.
[0079] The embodiment of the present invention is a system embodiment corresponding to the above method embodiment. The specific operations of each module can be understood with reference to the description of the method embodiment and will not be elaborated here.
[0080] Device embodiment 1
[0081] The embodiment of the present invention provides an electronic device, as Figure 7 shown, including: a memory 70, a processor 72, and a computer program stored on the memory 70 and executable on the processor 72. When the computer program is executed by the processor 72, the following method steps are implemented:
[0082] S1. Collect construction worker operation images through the tower crane top monitoring device, obtain a normal data set under normal working conditions, construct an interference data set by introducing interference factors, label and classify the normal data set and the interference data set, and divide them into a training set, a validation set, and a test set according to a predetermined ratio;
[0083] S2. Establish a lightweight object detection model SHM-YOLO. The SHM-YOLO network architecture includes a backbone network Backbone, a neck network Neck, and a head network Head. The backbone network Backbone includes a combined module HWD-C2f of an HWD convolution module and a C2f feature fusion module. An attention mechanism module SEAM is introduced in the neck network Neck and combined with the HWD convolution module and the C2f feature fusion module. SEAM is placed after each C2f feature fusion module;
[0084] S3. Set the training parameters, train the lightweight object detection model SHM-YOLO using the training set, optimize the constructed object detection model using the Inner-MPDIoU loss function, and evaluate the object detection results through evaluation metrics;
[0085] S4. Test the effect of the trained object detection model using the test set.
[0086] Device Embodiment II
[0087] An embodiment of the present invention provides a computer-readable storage medium, on which an implementation program for information transmission is stored. When the program is executed by a processor 72, the following method steps are implemented:
[0088] S1. Collect the operation images of construction workers through the monitoring device at the top of the tower crane, obtain the normal data set under normal working conditions, construct a interference data set by introducing interference factors, label and classify the normal data set and the interference data set, and divide them into a training set, a validation set and a test set according to a predetermined ratio;
[0089] S2. Establish a lightweight object detection model SHM-YOLO. The SHM-YOLO network architecture includes a backbone network Backbone, a neck network Neck, and a head network Head. The backbone network Backbone includes a combined module HWD-C2f composed of an HWD convolutional module and a C2f feature fusion module. An attention mechanism module SEAM is introduced into the neck network Neck and combined with the HWD convolutional module and the C2f feature fusion module. SEAM is placed after each C2f feature fusion module;
[0090] S3. Set the training parameters, train the lightweight object detection model SHM-YOLO using the training set, optimize the constructed object detection model using the Inner-MPDIoU loss function, and evaluate the object detection results through evaluation metrics;
[0091] S4. Test the effect of the trained object detection model using the test set.
[0092] The computer-readable storage medium described in this embodiment includes, but is not limited to: ROM, RAM, magnetic disk or optical disk, etc.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A lightweight object detection method based on improved YOLOv8, characterized in that, Including: Collect the operation images of construction workers through the monitoring equipment at the top of the tower crane, obtain the normal data set under normal working conditions, construct the interference data set by introducing interference factors, label and classify the normal data set and the interference data set, and divide them into training set, validation set and test set according to a predetermined ratio; Build a lightweight object detection model SHM-YOLO. The SHM-YOLO network architecture includes a backbone network Backbone, a neck network Neck and a head network Head. The backbone network Backbone includes a combined module HWD-C2f of an HWD convolutional module and a C2f feature fusion module. An attention mechanism module SEAM is introduced into the neck network Neck and combined with the HWD convolutional module and the C2f feature fusion module. SEAM is placed after each C2f feature fusion module; The attention mechanism module SEAM is specifically used for: In the attention mechanism module SEAM, the input image first passes through a channel and spatial mixing module CSMM to adjust the channel and spatial dimensions of the image. The Patch Embedding layer can change according to the set size and number of channels, and then passes through the GELU loss function; after being processed by the BatchNorm layer, the importance of different channels is learned through depthwise separable convolution, the important channels are retained, and the irrelevant parameters are removed; The input before depthwise separable convolution and the input after channel separation are combined through 1x1 convolution. After the input image is processed by three CSMM modules of different scales, the main channels of the occluded image in different dimensions are highlighted while compensating for the detail loss between channels. Then, Average Pooling is performed to remove redundant information, and all information is fused through two fully connected layers to enhance the connection of channel information. Finally, the output of the SEAM module is used as the attention and multiplied by the original features; Set the training parameters, use the training set to train the lightweight object detection model SHM-YOLO, use the Inner-MPDIoU loss function to optimize the constructed object detection model, and evaluate the object detection results through evaluation indicators; When training the constructed object detection model, the loss function used is the MPDIoU loss function Inner-MPDIoU that fuses Inner-IoU, which is specifically as follows: L inner-MPDIoU = L MPDIoU + IoU - IoU inner ; Use the test set to test the effect of the trained object detection model.
2. The method according to claim 1, wherein The specific method for constructing the interference data set by introducing interference factors is as follows: Add Gaussian noise to the normal data set to construct the interference data set.
3. The method according to claim 1, characterized in that, The HWD convolutional module includes a lossless feature encoding module and a feature representation learning module; The lossless feature encoding module is used to perform Haar wavelet transform on the input image, convert the image features and retain the image information while reducing the spatial resolution; The feature representation learning module consists of a standard convolutional layer, a batch normalization layer, and a ReLU activation layer. The 1x1 convolutional layer is used to process the number of channels of the feature map, so that the input and output channels of the front and back layers are consistent. The input image is further reduced of redundant information through the batch normalization layer and the ReLU activation layer.
4. The method according to claim 1, wherein The set training parameters include: Set the number of training epochs epoch = 300, the batch size batch size = 16, set the initial learning rate of the network to 0.001, use the Stochastic Gradient Descent (SGD) optimizer, and uniformly scale the size of the input image to 640×640.
5. The method according to claim 1, characterized in that The evaluation metrics include precision, recall, mAP@0.5, Params, and GFLOPs; the precision is used to evaluate the ratio of positive samples correctly identified as positive samples among the identified images. The higher the precision, the better the detection effect of the detection model; the recall is used to evaluate the proportion of truly positive samples among all samples detected as positive; the mAP is used to evaluate the average value of the average precision (AP) of multiple classes. The larger the value of mAP, the better the performance of the detection model; the Params and GFLOPs are used to evaluate the lightweight characteristics of the improved model.
6. A lightweight object detection system based on improved YOLOv8, characterized in that, Include: Data processing module: used to collect construction worker operation images through the tower crane top monitoring device, obtain a normal dataset under normal working conditions, construct an interference dataset by introducing interference factors, label and classify the normal dataset and the interference dataset, and divide them into a training set, a validation set, and a test set according to a predetermined ratio; Model construction module: used to establish a lightweight object detection model SHM-YOLO. The SHM-YOLO network architecture includes a backbone network Backbone, a neck network Neck, and a head network Head. The backbone network Backbone includes a combined module HWD-C2f composed of an HWD convolutional module and a C2f feature fusion module. An attention mechanism module SEAM is introduced in the neck network Neck and combined with the HWD convolutional module and the C2f feature fusion module. SEAM is placed after each C2f feature fusion module; The attention mechanism module SEAM is specifically used for: In the attention mechanism module SEAM, the input image first passes through a channel and spatial mixing module CSMM to adjust the channel and spatial dimensions of the image. The Patch Embedding layer can change according to the set size and number of channels, and then passes through the GELU loss function; after being processed by the BatchNorm layer, the depthwise separable convolution is used to learn the importance of different channels, retain the important channels, and remove irrelevant parameters; The input before depthwise separable convolution and the input after channel separation are combined through 1x1 convolution. After the input image is processed by three CSMM modules of different scales, the main channels of the occluded image in different dimensions are highlighted while the detail loss between channels is compensated. Subsequently, Average Pooling is performed to remove redundant information, and then all information is fused through two fully connected layers to enhance the connection of channel information. Finally, the output of the SEAM module is used as the attention and multiplied by the original features; Model training module: used to set training parameters, train the lightweight object detection model SHM-YOLO using the training set, optimize the constructed object detection model using the Inner-MPDIoU loss function, and evaluate the object detection results through evaluation metrics; When training the constructed object detection model, the loss function used is the MPDIoU loss function Inner-MPDIoU that fuses Inner-IoU, which is specifically as follows: L inner-MPDIoU = L MPDIoU + IoU - IoU inner ; Model testing module: used to test the effect of the trained object detection model using the test set.
7. An electronic device, characterized in that, Including: A processor; And, A memory arranged to store computer-executable instructions that, when executed, cause the processor to implement the steps of the lightweight object detection method based on improved YOLOv8 as described in any one of claims 1 to 5.
8. A storage medium, characterized in that, Used to store computer-executable instructions that, when executed, implement the steps of the lightweight object detection method based on improved YOLOv8 as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Unmanned aerial vehicle aerial photography small target detection method based on background noise weakening
CN118229965A
Lightweight PCB defect detection method based on improved YOLOv8 algorithm
CN118608509A