Method and system for detecting helmet wearing of electric vehicle rider based on improved YOLOv8s
By improving the YOLOv8s model, the StarNet, C2f-FusionStar and LSC-Detector modules were introduced, and combined with the InnerWIoU loss function, the problem of difficulty in both speed and accuracy in electric vehicle helmet detection is solved, and efficient and accurate detection is achieved in complex environments, which is suitable for resource-constrained equipment deployment.
Patent Information
- Application Number
- CN202510624586.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-26
AI Technical Summary
The existing electric vehicle helmet detection methods are difficult to take into account the detection speed and accuracy, and are prone to missed and missed detection in complex environments, which is difficult to meet the real-time detection needs. At the same time, the computing resource demand is high, which limits its application in resource-constrained environments.
Using the improved YOLOv8s model, by introducing the StarNet module, C2f-FusionStar module and LSC-Detector module, combined with the InnerWIoU loss function, feature extraction and bounding box regression are optimized, calculation complexity and parameter quantity are reduced, detection accuracy and real-time performance are improved.
In complex scenarios, the detection accuracy and real-time performance are significantly improved, the model parameter quantity and computing resource requirements are reduced, making it suitable for deployment on resource-constrained edge devices, and the detection and generalization capabilities of small targets are improved.
Smart Images

Figure CN120544233A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a detection method, in particular to a method for detecting the wearing of a helmet by an electric vehicle rider based on an improved YOLOv8s, and belongs to the technical field of target detection. Background Art
[0002] With the rapid growth of electric bicycle ownership, they have become a primary means of transportation for short-distance travel in cities. However, electric bicycle accidents are frequent, with head injuries caused by not wearing a helmet accounting for 68.7% of electric bicycle fatalities (WHO, 2024). Therefore, automated helmet detection for electric bicycle riders is of great practical significance.
[0003] Currently, deep learning-based object detection algorithms have surpassed traditional object recognition image processing methods in many fields. Object detection algorithms are mainly divided into two categories: traditional object detection algorithms and deep learning-based object detection algorithms. Deep learning-based object detection algorithms are mainly divided into single-stage object detection algorithms and two-stage object detection algorithms. Two-stage detection algorithms such as Faster R-CNN and Mask R-CNN have high accuracy but slow speed. Single-stage algorithms (such as the YOLO series, SSD, and RetinaNet) effectively address the slow speed of two-stage algorithms.
[0004] While existing methods for detecting electric bike helmets have made some progress, their performance remains unsatisfactory. Common challenges include balancing detection speed and accuracy, a relatively limited number of detection categories, and the tendency for missed and false detections to occur, making it difficult to meet the real-time detection requirements in complex road scenarios.
[0005] Zhu Zhouhua et al. improved the detection of small targets in helmets by introducing the CBAM and CA attention modules, combined with an improved DIoUNMS algorithm. However, this method's high computational resource requirements limit its application in resource-constrained environments. Xie Boxuan et al. added an efficient channel ECA attention mechanism and a Bi-FPN module to YOLOv5, and introduced Alpha-CIOULoss, further balancing the importance of features at different levels and enhancing localization performance. However, this method still had a high miss detection rate when targets were occluded. Wang Guicheng et al. optimized the anchor box parameters of YOLOv5 using a clustering algorithm, aggregating the output network to improve feature resolution, achieving a mAP of 94.53% at 30 FPS. However, the experimental data focused on a single scene, and the generalization ability to complex environments such as rain, fog, and backlight was insufficient.
[0006] To address the above problems, the present invention proposes a method and system for detecting the wearing of helmets by electric vehicle riders based on improved YOLOv8s. Through modular improvement and loss function optimization, the detection accuracy and real-time performance of the model in multiple scenarios are improved, while the number of model parameters is reduced, making it more suitable for deployment on resource-constrained edge devices. Summary of the Invention
[0007] In response to the problems in the prior art, the present invention provides a method and system for detecting the wearing of helmets by electric vehicle riders based on improved YOLOv8s.
[0008] The purpose of the present invention can be achieved through the following technical solutions:
[0009] An electric vehicle rider helmet wearing detection system based on improved YOLOv8s, including:
[0010] The data acquisition and preprocessing module is used to collect image data of electric vehicle riders and preprocess and annotate the images;
[0011] The model training module is used to train the improved YOLOv8s model based on the preprocessed dataset and compare the optimal model after training;
[0012] The detection module is used to input real-time detection images into the trained improved YOLOv8s model to obtain detection results.
[0013] Preferably, the data acquisition and preprocessing module includes:
[0014] The data acquisition unit is used to crawl image data using Python scripts and take field photos in various scenes and weather conditions to obtain high-quality original images. It also filters the required images from public datasets to enrich the dataset diversity.
[0015] The data preprocessing unit is used to remove low-quality images, perform data enhancement processing, and label them into three categories of targets. Finally, the dataset is divided into training set, validation set, and test set to ensure diversity and representativeness.
[0016] Preferably, the improved YOLOv8s model includes:
[0017] S1: Introducing the StarNet module into the backbone network, achieving efficient fusion of input features and learnable weights through "star operations" to reduce computational complexity;
[0018] S2: Neck network. The C2f-FusionStar module is used in the neck network. By replacing the Bottleneck structure in the original C2f module with StarBlocks, the multi-scale feature fusion capability is enhanced. The input features are extracted through StarBlocks, and the feature expression capability is enhanced through channel splicing and convolution operations.
[0019] S3: The LSC-Detector module is introduced into the detection head. The convolution branch with shared parameters and group normalization (GN) replace the traditional detection head to reduce the number of parameters and improve positioning accuracy. Group normalization (GN) is used to replace the traditional normalization layer to improve the small target detection capability. The calculation formula of group normalization is as follows:
[0020] S4: In model training, the InnerWIoU loss function is used to adjust the gradient gain through a dynamic non-monotonic focusing mechanism to optimize low-quality sample learning, and then the gradient gain is dynamically adjusted by evaluating the abnormality of the anchor box.
[0021] Preferably, the StarNet module realizes efficient fusion of input features and learnable weights based on “star operation”, and its calculation formula is:
[0022]
[0023] Where: star is the result after the star operation; W is the weight matrix of the linear layer; T represents the transpose of the matrix; B is the bias of the linear layer; X is the input; * represents the star operation. The channel attention mechanism is applied to the fused feature map to further enhance the expression of important features. The calculation formula of channel attention is as follows:
[0024] F attention =δ(MLP(GAP(F out )));
[0025] Among them, σ is the sigmoid activation function, MLP is the multi-layer perceptron, and GAP is the global average pooling operation;
[0026] Finally, the feature map enhanced by the attention mechanism is passed through the convolution layer to generate the final feature map, which is used by the subsequent detection head for target classification and positioning.
[0027] Preferably, the feature map F from the backbone network is received in , and perform dimensionality reduction processing on the channel dimension through the convolution layer to extract preliminary features. The feature map after dimensionality reduction is split into two parts F1 and F2, and feature extraction and transmission are performed separately. Then, dynamic feature interaction is performed on F1 and F2 through StarBlocks to enhance the expression ability of multi-scale features. The calculation formula of feature fusion is as follows:
[0028] F fuse =Concat(F1, F2)·W fuse ;
[0029] Among them, W fuse is the fusion weight matrix, used to adjust the importance of features;
[0030] The channel attention mechanism is applied to the fused feature map to further enhance the expression of important features; the calculation formula is as follows:
[0031] F attention =δ(MLP(GAP(F fuse )));
[0032] Among them, δ is the sigmoid activation function, MLP is the multi-layer perceptron, and GAP is the global average pooling operation;
[0033] The feature map enhanced by the attention mechanism is restored through the convolution layer to generate the final feature map F out , F out It is passed to the subsequent detection head module for target classification and positioning. The formula for feature transfer is as follows:
[0034] F out =Conv(F attention );
[0035] Among them, Conv represents the convolution operation, which is used to restore the channel dimension of the feature map and enhance the feature expression capability;
[0036] Through the above steps, the C2f-FusionStar module can effectively enhance the fusion capability of multi-scale features and improve the model's detection accuracy for targets in complex scenes.
[0037] Preferably, the feature map F is received from the neck network in And adjust the channel dimension through the convolution layer to extract preliminary features, and split the adjusted feature map into two parts F loc and F cls , respectively used for positioning and classification tasks, for F loc and F cls Applying the convolution branch with shared parameters can reduce the number of parameters and improve computational efficiency. The calculation formula for the shared parameter convolution is as follows:
[0038] For the input feature map F∈R C×H×W ,The shared parameter convolution operation is defined as:
[0039]
[0040] Among them, W i Represents the convolution kernel parameters. Under the shared parameter mechanism, different positions or branches use the same W i ; F i represents the i-th channel of the input feature map; b represents the bias term. Similarly, in the shared parameter mechanism, different positions or branches use the same b; σ represents the activation function, which is used to introduce nonlinearity; Conv shared represents the convolution operation with shared parameters;
[0041] The shared features are normalized by group normalization (GN) to improve the small target detection capability. The calculation formula of group normalization is as follows:
[0042]
[0043] Among them, x ij is the input feature, μ ij and are the mean and variance of the j-th group of features, γ j and β j are learnable scaling and offset parameters, and ∈ is a parameter to prevent numerical instability;
[0044] Positioning branch: F loc Perform bounding box regression, calculate the deviation between the predicted box and the true box and optimize it. The calculation formula for bounding box regression is as follows:
[0045] Δ bbox =Conv reg (F loc );
[0046] Among them, Conv reg represents the convolution operation used for regression;
[0047] F cls Perform classification and predict the target category. The classification calculation formula is as follows:
[0048] p cls =Conv cls (F cls );
[0049] Among them, Conv cls Represents the convolution operation used for classification;
[0050] The features of the positioning branch and the classification branch are fused to generate the final detection result. The formula for feature fusion is as follows:
[0051] F out =Convcat(F loc ,F cls );
[0052] Among them, Concat represents the feature concatenation operation;
[0053] Output the detection results (including bounding boxes, categories, and confidence levels) for subsequent visualization and application. The format of the detection results is as follows:
[0054] Result={bbox, class, confidence};
[0055] Among them, bbox is the bounding box coordinate, class is the target category, and confidence is the confidence.
[0056] Through the above steps, the LSC-Detector module can effectively reduce the number of parameters, while improving the accuracy of positioning and classification, and enhancing the model's ability to detect small targets.
[0057] Preferably, IoU calculation: calculate the intersection over union (IoU) between the predicted bounding box and the true bounding box. As a basic indicator, the calculation formula of IoU is as follows:
[0058]
[0059] Among them, x gt and y gt is the coordinate of the real box, x pred and y pred is the coordinate of the prediction box;
[0060] The loss value is adjusted through the dynamic gradient gain mechanism to optimize the learning of low-quality samples. The calculation formula of dynamic gradient gain is as follows:
[0061]
[0062] Among them, ε is a parameter to prevent numerical instability;
[0063] The ratio of the distance between the anchor box and the ground-truth box to the average distance is the outlier of the anchor box, which is used to evaluate the quality of the anchor box. The focus coefficient is then dynamically adjusted according to the outlier to optimize the gradient distribution. The calculation formula is as follows:
[0064] τ=β·exp(-α·outlier)
[0065] Among them, β is the hyperparameter, α is the adjustment coefficient, and τ is the dynamic focusing coefficient;
[0066] Gradient adjustment adjusts the gradient gain based on the dynamic focusing coefficient to reduce the adverse effects of low-quality samples on model training. The formula for gradient adjustment is as follows:
[0067]
[0068] in, represents the gradient operator;
[0069] The adjusted loss value is used to update the model parameters through back propagation to optimize the bounding box regression. The back propagation formula is as follows:
[0070]
[0071] Among them, θ new is the updated model parameter, θ old are the model parameters before updating, is the learning rate.
[0072] Through the above steps, the InnerWIoU loss function can effectively optimize the learning of low-quality samples and improve the model's detection accuracy and positioning performance for targets in complex scenes.
[0073] A method for detecting whether an electric vehicle helmet is worn is provided. The method is based on the above-mentioned electric vehicle helmet wearing detection system, and further includes the following steps:
[0074] S1: Dataset for building an electric vehicle helmet detection model;
[0075] S2: Preprocess the dataset;
[0076] S3: Train the electric vehicle helmet detection model based on the dataset;
[0077] S4: Input the real-time detection image into the electric vehicle helmet detection model to obtain the detection results.
[0078] Beneficial effects of the present invention:
[0079] 1. The C3 module in the backbone layer is replaced with a lightweight StarNet module; the C2f module in the neck layer is replaced with a C2f-FusionStar module. This reduces network redundancy and memory loss while maintaining network feature extraction capabilities.
[0080] 2. The introduction of the lightweight LSC-Detector module reduces the redundancy of convolution kernels and improves the ability to extract spatial semantic features, reduces the number of network parameters and computational complexity, and improves detection accuracy.
[0081] 3. Optimize the bounding box loss function of the YOLOv8s network and construct an InnerWIoU loss function with a dynamic focusing mechanism. This increases the network model's attention to low-quality data and reduces the penalty for high-quality data. This improves the model's generalization ability for the dataset and is more conducive to model deployment for street detection in complex backgrounds.
[0082] Compared with the standard YOLOv8s network, the improved network reduces the requirements for device memory and processors required for deployment, is more capable of edge deployment, and can better adapt to image data detection tasks in generalized environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0084] Figure 1 is a flow chart of the present invention;
[0085] Figure 2 This is the StarNet module structure diagram;
[0086] Figure 3 This is the network structure diagram of the C2f_Star module;
[0087] Figure 4 This is the network structure diagram of the LSC-Detector module;
[0088] Figure 5 Improved YOLOv8s network structure diagram;
[0089] Figure 6 Here are three image examples of how to obtain sample data;
[0090] Figure 7 middle;
[0091] Figure 7 a is a scene with one person in one car;
[0092] Figure 7 b is the original YOLOv8s detection image of a scene with one car and one person;
[0093] Figure 7 c is the detection image of the improved model for the scene with one car and one person;
[0094] Figure 8 middle;
[0095] Figure 8 a is a scene with multiple people in one car;
[0096] Figure 8 b is the original YOLOv8s detection image of a scene with a car and multiple people occluded;
[0097] Figure 8 c is the detection image of the improved model in a scene with occlusion of one car and multiple people;
[0098] Figure 9 middle;
[0099] Figure 9 a is a scene with occlusion and dense vehicles;
[0100] Figure 9 b is the original YOLOv8s detection image of a scene with occlusion and dense vehicles;
[0101] Figure 9 c is the detection image of the improved model in a scene with occlusion and dense vehicles;
[0102] Figure 10 middle;
[0103] Figure 10 a is an intersection scene with complex vehicle types and small targets;
[0104] Figure 10 b is the original YOLOv8s detection image of the intersection scene with complex vehicle types and small targets;
[0105] Figure 10 c is the detection image of the improved model for an intersection scene with complex vehicle types and small targets. DETAILED DESCRIPTION
[0106] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0107] See also Figure 1-10 As shown, a method for detecting the wearing of a helmet by an electric vehicle rider based on an improved YOLOv8s is provided; the method comprises the following steps:
[0108] Step 1: Collect image data of electric bike riders, build a dataset of images of electric bike riders wearing helmets, and pre-process and annotate the images.
[0109] Furthermore, image preprocessing includes removing blur and ghost images, and performing HSV color interference, Gaussian noise, and salt and pepper noise data enhancement on the remaining image data.
[0110] Step 2: Build an improved YOLOv8s network and train it on the training set to obtain the optimal weights;
[0111] Furthermore, the improved YOLOv8s network includes: semantic information extraction by the backbone layer, which includes the StarNet module, C2f-FusionStar module, LSC-Detector module and an SPPF spatial pyramid pooling module; semantic fusion by the neck layer, which includes the C2f-FusionStar module, upsampling module and Involution inner convolution module; and finally, through the Conv2d convolution layer, generating feature maps of three different scales for prediction.
[0112] Furthermore, the StarNet module achieves efficient fusion of input features and learnable weights based on the “star operation”. The calculation formula is as follows:
[0113]
[0114] Where: star is the result after the star operation; W is the weight matrix of the linear layer; T represents the transpose of the matrix; B is the bias of the linear layer; X is the input; * represents the star operation.
[0115] Furthermore, the C2f-FusionStar module enhances the multi-scale feature fusion capability by replacing the Bottleneck structure in the original C2f module with StarBlocks.
[0116] Step 3: Use InnerWIoU loss function;
[0117] Furthermore, the InnerWIoU loss function is as follows:
[0118] L WIoU =1-L IoU +R WIoU ×R WIoU ;
[0119]
[0120] Where: L WIoU is the WIoU loss; γ is the gradient gain; R WIoU is the weight coefficient penalty term; L IoU is the bounding box loss; (x, y) is the center coordinate of the anchor box; (x gt ,y gt ) is the center coordinate of the target frame; W g , H g are the width and height of the minimum bounding box respectively; * represents the separation operation, which no longer tracks gradient information; β is the outlier degree; δ and α are hyperparameters (in this paper, δ is 3 and δ is 1.9); is the actual bounding box loss; is the average bounding box loss;
[0121] Furthermore, a lightweight electric vehicle rider helmet wearing detection system includes: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement a lightweight electric vehicle rider helmet wearing detection method.
[0122] Furthermore, a computer-readable medium stores computer program code, which, when executed by a processor, implements a lightweight electric vehicle rider helmet wearing detection method.
[0123] like Figure 1 As shown, the present invention provides a method for detecting the wearing of a helmet by an electric vehicle rider based on an improved YOLOv8s, comprising the following steps:
[0124] Step 1: Data collection and preprocessing
[0125] First, we used web crawlers to obtain image data of scooter riders from web pages, using keywords such as "scooter rider," "scooter helmet," and "traffic safety helmet," yielding a total of 5,000 images. Next, we captured 2,356 images in various locations, including traffic intersections, side streets, and residential gates. We also integrated image data from public datasets such as Kaggle to select 644 high-quality images. The final dataset contains 8,000 images, covering a variety of scenes and weather conditions.
[0126] Preprocess the image: Image screening: Remove low-quality images such as blur and ghosting. Data enhancement: Perform HSV color interference, Gaussian noise, salt and pepper noise processing on the image to enhance the model's adaptability to complex environments. Labeling and classification: Use the Labelimg tool to label the image and divide the detection targets into three categories: electric vehicle (twowheeler), wearing a helmet (helmet) and without a helmet (withouthelmet). The annotation format is a txt file in YOLO format. Dataset division: The dataset is divided into training set, validation set and test set in a ratio of 7:2:1, with 5,600 training sets, 1,600 validation sets and 800 test sets.
[0127] Step 2: Improvement of the original YOLOv8s network
[0128] (1) The YOLOv8s backbone network is introduced into the StarNet module, and the “star operation” is used to achieve efficient fusion of input features and learnable weights, thereby reducing computational complexity;
[0129] like Figure 2As shown in the figure, the StarNet module uses the "star operation" to achieve efficient fusion of input features and learnable weights, reducing computational complexity. The calculation formula is as follows:
[0130]
[0131] Where: star is the result after the star operation; W is the weight matrix of the linear layer; T represents the transpose of the matrix; B is the bias of the linear layer; X is the input; * represents the star operation.
[0132] First, feature extraction is performed. The input feature map is downsampled through a convolutional layer. Convolution kernels of varying sizes are used to extract multi-scale features to capture objects of varying sizes. Next, a star operation is performed on the extracted features. This efficient fusion of the input features with learnable weights is achieved through element-by-element multiplication, enhancing feature representation. After the star operation, important features in the feature map are enhanced, while less important features are suppressed, improving the model's ability to detect objects, particularly in complex backgrounds and multi-target scenarios. Finally, the output fused feature map serves as input to subsequent modules for further object detection tasks.
[0133] (2) Introducing the YOLOv8s neck network into the StarNet module
[0134] like Figure 3 As shown in the figure, the C2f-FusionStar module significantly enhances multi-scale feature fusion capabilities by replacing the Bottleneck structure in the original C2f module with StarBlocks. Input features first enter the module, where they are extracted using StarBlocks to capture feature information at different scales. Subsequently, the features undergo channel concatenation and convolution operations within the module to achieve feature interaction and fusion, ensuring the consistency and richness of feature maps at different scales. Finally, the module outputs a fused feature map that integrates the advantages of multi-scale features, retains the detailed information of the original features, and improves feature expression. This provides high-quality feature input for subsequent object detection tasks, ensuring the model's high performance in complex scenarios.
[0135] The mathematical expression of feature fusion is as follows:
[0136]
[0137] Where: F input Represents the feature map input to the C2f-FusionStar module, StarBlocks: represents the StarBlocks operation, used to extract multi-scale features, F extracted Represents the feature map extracted by StarBlocks.
[0138] F residual Represents the residual features, usually from the input of the module or the features of the previous layer. Concat represents the channel concatenation operation, which concatenates the two feature maps in the channel dimension. fuse Represents the fusion weight matrix, which is used to adjust the importance of the spliced features. fused Represents the concatenated and fused feature map. out Represents the final output feature map, which is used for subsequent target detection tasks.
[0139] (3) Change the YOLOv8s detection head to the LSC-Detector module
[0140] The LSC-Detector module replaces the traditional detection head with a convolution branch and group normalization (GN) with shared parameters, effectively reducing the number of parameters and improving positioning accuracy. The specific implementation is as follows: First, a convolution branch with shared parameters is used to extract features from the input feature map. This sharing mechanism significantly reduces the number of parameters while maintaining the expressiveness of the features. Next, group normalization (GN) is used to replace the traditional normalization layer to enhance the small target detection capability, because GN can better handle the normalization problem of small batches of data, thereby improving the stability of feature expression. Finally, the module separates the positioning branch and the classification branch, performing bounding box regression and classification tasks respectively. This separation design enables the model to more accurately locate and classify targets, ultimately outputting high-quality detection results.
[0141]
[0142] Where: F in is the input feature map, Conv shared represents the convolution operation with shared parameters, F shared is the feature map after shared parameter convolution, x ijk is the input feature, μ G and are the mean and variance of the G-th group features, γ and β are learnable scaling and offset parameters, ∈ is a parameter to prevent numerical instability, and y ijk is the normalized feature.
[0143] (4) Change YOLOv8s loss function to InnerWIoU
[0144] The InnerWIoU loss function adjusts gradient gain through a dynamic non-monotonic focusing mechanism to optimize learning for low-quality examples. Specifically, this mechanism uses "outlier" instead of the traditional IoU to evaluate anchor box quality and provides a judicious gradient gain allocation strategy. This strategy reduces the competitiveness of high-quality anchor boxes while also reducing harmful gradients generated by low-quality examples. This allows the model to focus more on anchor boxes of average quality during training, thereby improving overall detection performance.
[0145] L WIoU =1-L IoU +R WIoU ×R WIoU ;
[0146]
[0147] Where: L WIoU is the WIoU loss; γ is the gradient gain; R WIoU is the weight coefficient penalty term; L IoU is the bounding box loss; (x, y) is the center coordinate of the anchor box; (x gt ,y gt ) is the center coordinate of the target frame; W g , H g are the width and height of the minimum bounding box respectively; * represents the separation operation, which no longer tracks gradient information; β is the outlier degree; δ and α are hyperparameters (in this paper, δ is 3 and δ is 1.9); is the actual bounding box loss; is the average bounding box loss.
[0148] Model training and deployment
[0149] Step 3: Model training and deployment
[0150] Model training:
[0151] Training environment: Use NVIDIA GeForce RTX 3090 GPU and train based on the PyTorch framework.
[0152] Data augmentation: Mosaic data augmentation strategy is used during training.
[0153] The learning rate adjustment formula is as follows:
[0154]
[0155] η t is the learning rate for the tth training step. η min and η min are the minimum and maximum learning rate, respectively. t is the current training step. T is the total number of training steps.
[0156] Model deployment: Deploy the trained model to the traffic monitoring system for real-time detection. The lightweight design ensures that the model runs efficiently on resource-constrained edge devices.
[0157] Step 4: Experimental verification
[0158] Table 1 Experimental environment platform configuration:
[0159] Table 1 Model parameters
[0160]
[0161] Experimental results:
[0162] In terms of performance indicators, the improved model showed significant improvements. Specifically, the mAP@0.5 of the improved model successfully reached 89.3%, an increase of 3.4 percentage points compared to the baseline model. This shows that the model's ability to locate and classify targets in detection tasks has been significantly enhanced. At the same time, the number of model parameters has been reduced by 54%, which not only reduces the storage requirements of the model, but also reduces the consumption of computing resources, making the model more lightweight. GFLOPs (billion floating-point operations) has been reduced to 13.0, further demonstrating the optimization of the model's computational efficiency. In the robustness test, the improved model's detection performance in multiple perspectives and complex environments (such as rain, fog, backlighting, and occlusion) has been significantly improved, and both the missed detection rate and the false detection rate have been greatly reduced. These results show that the improved model significantly improves detection accuracy and robustness while maintaining high efficiency, and can better adapt to various challenges in actual application scenarios.
[0163] Results analysis shows that the addition of the StarNet module effectively improves feature extraction efficiency and reduces computational complexity. StarNet implicitly maps from a low-dimensional space to a high-dimensional feature space through star operations, achieving efficient feature extraction without increasing network width or complex design. This property enables the model to obtain richer and more expressive feature representations while maintaining high efficiency, which is particularly important for capturing fine-grained features in object detection tasks. The introduction of the C2f-FusionStar module enhances multi-scale feature fusion capabilities and improves the model's ability to detect occluded objects. By replacing the Bottleneck structure in the original C2f module with StarBlocks, the C2f-FusionStar module better integrates feature information from different scales, resulting in excellent performance when handling occluded objects. The LSC-Detector module replaces the traditional detection head with a parameter-shared convolutional branch and group normalization (GN), reducing the number of parameters while improving localization accuracy. The parameter-shared convolutional branch reduces the model's parameter count, while GN enhances the ability to detect small objects, allowing the model to achieve more accurate object localization while maintaining its lightweight design. The InnerWIoU loss function optimizes learning from low-quality samples and improves small object detection by dynamically adjusting gradient gains. This loss function helps the model better handle small object detection in complex scenes, thereby improving overall detection performance. These improvements significantly enhance detection accuracy and robustness while maintaining high efficiency.
[0164] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. An electric vehicle rider helmet wearing detection system based on improved YOLOv8s, characterized in that: include: The data acquisition and preprocessing module is used to collect image data of electric vehicle riders and preprocess and annotate the images; The model training module is used to train the improved YOLOv8s model based on the preprocessed dataset and compare the optimal model after training; The detection module is used to input real-time detection images into the trained improved YOLOv8s model to obtain detection results.
2. The detection system according to claim 1, characterized in that The data acquisition and preprocessing module includes: The data acquisition unit is used to crawl image data using Python scripts and take field photos in various scenes and weather conditions to obtain original images. It also filters the required images from public datasets to enrich the dataset diversity. The data preprocessing unit is used to remove low-quality images, perform data enhancement processing, and label them into three categories of targets. Finally, the dataset is divided into training set, validation set, and test set.
3. The detection system according to claim 2, characterized in that The improved YOLOv8s model includes: S1: Introducing the StarNet module into the backbone network, achieving efficient fusion of input features and learnable weights through "star operations" to reduce computational complexity; S2: Neck network. The C2f-FusionStar module is used in the neck network. By replacing the Bottleneck structure in the original C2f module with StarBlocks, the multi-scale feature fusion capability is enhanced. The input features are extracted through StarBlocks, and the feature expression capability is enhanced through channel splicing and convolution operations. S3: The LSC-Detector module is introduced into the detection head. The convolution branch with shared parameters and group normalization replace the traditional detection head to reduce the number of parameters and improve positioning accuracy. Group normalization is used to replace the traditional normalization layer to improve the small target detection capability. The calculation formula of group normalization is as follows: S4: In model training, the InnerWIoU loss function is used to adjust the gradient gain through a dynamic non-monotonic focusing mechanism to optimize low-quality sample learning, and then the gradient gain is dynamically adjusted by evaluating the abnormality of the anchor box.
4. The detection system according to claim 3, characterized in that The StarNet module achieves efficient fusion of input features and learnable weights based on the "star operation", and its calculation formula is: Where: star is the result after the star operation; W is the weight matrix of the linear layer; T represents the transpose of the matrix; B is the bias of the linear layer; X is the input; * represents a star operation. The channel attention mechanism is applied to the fused feature map to further enhance the expression of important features. The calculation formula of channel attention is as follows: F attention δ(MLP(GAP(F out ))) Among them, σ is the sigmoid activation function, MLP is the multi-layer perceptron, and GAP is the global average pooling operation; Finally, the feature map enhanced by the attention mechanism is passed through the convolution layer to generate the final feature map, which is used by the subsequent detection head for target classification and positioning.
5. The detection system according to claim 4, characterized in that: Receive the feature map F from the backbone network in , and perform dimensionality reduction processing on the channel dimension through the convolution layer to extract preliminary features. The feature map after dimensionality reduction is split into two parts F1 and F2, and feature extraction and transmission are performed separately. Then, dynamic feature interaction is performed on F1 and F2 through StarBlocks to enhance the expression ability of multi-scale features. The calculation formula of feature fusion is as follows: F fuse =Concat(F1,F2)·W fuse ; Among them, W fuse is the fusion weight matrix, used to adjust the importance of features; The channel attention mechanism is applied to the fused feature map to further enhance the expression of important features; the calculation formula is as follows: F attention δ(MLP(GAP(F fuse ))) Among them, δ is the sigmoid activation function, MLP is the multi-layer perceptron, and GAP is the global average pooling operation; The feature map enhanced by the attention mechanism is restored through the convolution layer to generate the final feature map F out , F out It is passed to the subsequent detection head module for target classification and positioning. The formula for feature transfer is as follows: F out =Conv(F attention ); Among them, Conv represents the convolution operation, which is used to restore the channel dimension of the feature map and enhance the feature expression capability.
6. The detection system according to claim 5, characterized in that: Receive feature map F from the neck network in And adjust the channel dimension through the convolution layer to extract preliminary features, and split the adjusted feature map into two parts F loc and F cls , respectively used for positioning and classification tasks, for F loc and F cls Applying the convolution branch with shared parameters can reduce the number of parameters and improve computational efficiency. The calculation formula for the shared parameter convolution is as follows: For the input feature map F∈R C×H×W ,The shared parameter convolution operation is defined as: Among them, W i Represents the convolution kernel parameters. Under the shared parameter mechanism, different positions or branches use the same W i ; F i represents the i-th channel of the input feature map; b represents the bias term. Similarly, in the shared parameter mechanism, different positions or branches use the same b; σ represents the activation function, which is used to introduce nonlinearity; Conv shared represents the convolution operation with shared parameters; The shared features are normalized through group normalization to improve the small target detection capability. The calculation formula of group normalization is as follows: Among them, x ij is the input feature, μ ij and are the mean and variance of the j-th group of features, γ j and β j are learnable scaling and offset parameters, and ∈ is a parameter to prevent numerical instability; Positioning branch: F loc Perform bounding box regression, calculate the deviation between the predicted box and the true box and optimize it. The calculation formula for bounding box regression is as follows: Δ bbox =Conv reg (F loc ); Among them, Conv reg represents the convolution operation used for regression; F cls Perform classification and predict the target category. The classification calculation formula is as follows: p cls =Conv cls (F cls ); Among them, Conv cls Represents the convolution operation used for classification; The features of the positioning branch and the classification branch are fused to generate the final detection result. The formula for feature fusion is as follows: F out =Convcat(F loc ,F cls ); Among them, Concat represents the feature concatenation operation; Output the test results for subsequent visualization and application. The format of the test results is as follows: Result={bbox, class, confidence}; Among them, bbox is the bounding box coordinate, class is the target category, and confidence is the confidence.
7. The detection system according to claim 6, characterized in that IoU calculation: Calculate the intersection-over-union ratio between the predicted bounding box and the true bounding box. As a basic indicator, the IoU calculation formula is as follows: Among them, x gt and y gt is the coordinate of the real box, x pred and y pred is the coordinate of the prediction box; The loss value is adjusted through the dynamic gradient gain mechanism to optimize the learning of low-quality samples. The calculation formula of dynamic gradient gain is as follows: Among them, ε is a parameter to prevent numerical instability; The ratio of the distance between the anchor frame and the true frame to the average distance is the abnormality of the anchor frame, which is used to evaluate the quality of the anchor frame. Then, the focus coefficient is dynamically adjusted according to the abnormality to optimize the gradient distribution. The calculation formula is as follows: τ = β·exp(-α·outlier); Among them, β is the hyperparameter, α is the adjustment coefficient, and τ is the dynamic focusing coefficient; Gradient adjustment adjusts the gradient gain based on the dynamic focusing coefficient to reduce the adverse effects of low-quality samples on model training. The formula for gradient adjustment is as follows: in, represents the gradient operator; The adjusted loss value is used to update the model parameters through back propagation to optimize the bounding box regression. The back propagation formula is as follows: Among them, θ new is the updated model parameter, θ old are the model parameters before updating, is the learning rate.
8. A method for detecting the wearing of an electric vehicle helmet, based on the electric vehicle helmet wearing detection system according to any one of claims 1 to 7, characterized in that: The method further comprises the steps of: S1: Dataset for building an electric vehicle helmet detection model; S2: Preprocess the dataset; S3: Train the electric vehicle helmet detection model based on the dataset; S4: Input the real-time detection image into the electric vehicle helmet detection model to obtain the detection results.