Helmet wearing detection method and device for two-wheeled vehicle, electronic equipment and medium
By improving the YOLOv5 model and combining it with K-means clustering, SE attention, and EIOU loss function, the efficiency and accuracy issues of two-wheeled vehicle helmet wearing detection are solved, adapting to diverse riding scenarios and improving the detection effect of light two-wheeled vehicles.
Patent Information
- Application Number
- CN202510770393.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies are inefficient in detecting helmet wearing on two-wheeled vehicles, are difficult to adapt to diverse riding scenarios, require high computing resources, and are insufficient for detecting light two-wheeled vehicles, resulting in frequent missed detections and false detections.
An improved YOLOv5 model is used to generate a helmet wearing detection model by introducing the K-means clustering algorithm, SE attention mechanism and EIOU loss function, combined with data enhancement and multi-type dataset training, to improve detection accuracy and efficiency.
While ensuring lightweight, the accuracy and efficiency of helmet wearing detection are significantly improved, adapting to different shooting angles and complex backgrounds, and reducing computing resource requirements.
Smart Images

Figure CN120655951A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a method, device, electronic device and storage medium for detecting helmet wearing on a two-wheeled vehicle. Background Art
[0002] With the acceleration of urbanization and the widespread adoption of green mobility, two-wheeled vehicles (including motorcycles, electric vehicles, bicycles, and electric scooters) have become an essential means of transportation for daily commuting and recreation. However, two-wheeled vehicle accidents are frequent, with head injuries being a leading cause of serious injury and even death among riders. Numerous studies have shown that wearing a helmet is an effective way to reduce injuries from two-wheeled vehicle accidents and protect riders' lives. Helmets not only significantly reduce the risk of head injuries from impact or crushing, but also effectively protect riders from natural environmental factors such as wind, sand, and rain, enhancing overall riding comfort and safety. Therefore, wearing a helmet should be a conscious and accepted code of conduct for every cyclist.
[0003] To promote the popularization of helmet wearing, governments around the world have introduced relevant laws and regulations, explicitly requiring the wearing of helmets when riding two-wheeled vehicles, and have raised public awareness of the importance of wearing helmets through mandatory measures and extensive publicity and education activities. However, in the actual supervision of helmet wearing, traditional manual supervision methods face many challenges. Manual supervision is not only inefficient and difficult to cover all riding scenarios, but is also prone to problems such as uneven supervision intensity and frequent blind spots, making it difficult to form a sustained and effective deterrent. Therefore, the introduction of intelligent technical means to achieve automated and real-time monitoring of helmet wearing has become the key to solving the current management dilemma. Obviously, there is an urgent need for a new method for detecting helmet wearing for two-wheeled vehicles to solve at least one of the above problems.
[0004] It should be noted that the above content only provides background technical information related to this application and does not necessarily constitute prior art. Summary of the Invention
[0005] In view of the above-mentioned shortcomings of the prior art, the present application provides a method, device, electronic device and storage medium for detecting helmet wearing on two-wheeled vehicles, so as to improve the efficiency and accuracy of helmet wearing detection on two-wheeled vehicles.
[0006] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0007] According to one aspect of an embodiment of the present application, a method for detecting helmet wearing for a two-wheeled vehicle is provided, comprising: obtaining an initial detection image and a helmet wearing label; annotating the initial detection image according to the helmet wearing label to generate a sample data set with the helmet wearing label, and extracting a sample training set from the sample data set; performing data enhancement on the sample training set by a preset data enhancement method to obtain a target data set; obtaining an improved YOLOv5 model by introducing a K-means clustering algorithm, adding a SE attention mechanism, and introducing an EIOU loss function into the YOLOv5 model; training the improved YOLOv5 model based on the target data set to obtain a helmet wearing detection model; obtaining an image to be tested, the image to be tested including an image of a driver or passenger of a target two-wheeled vehicle; performing helmet wearing detection on the image to be tested by using the helmet wearing detection model to obtain a helmet wearing detection result for the target two-wheeled vehicle.
[0008] In one embodiment of the present application, based on the aforementioned scheme, the method further includes: extracting the width and height of all target boxes from the target data set; randomly selecting a preset number of target boxes as initial clustering centers, and assigning each target box to the initial clustering center with the closest relative distance to obtain a preset number of initial clusters, wherein the relative distance is calculated based on the width and height of the target box; calculating according to the width and height of the target box in each of the initial clusters to obtain a new clustering center for each of the initial clusters; repeating the iterative allocation until a preset end condition is met, and determining the new clustering center as the target anchor box, wherein the preset end condition includes that the change in the new cluster center is less than a preset change, or the number of iterations is greater than or equal to a preset number threshold; applying the target anchor box to the YOLOv5 model to complete the introduction of the K-means clustering algorithm in the YOLOv5 model.
[0009] In one embodiment of the present application, based on the aforementioned scheme, the method further includes: compressing the spatial dimension of each channel in the input feature map through global average pooling to generate a channel descriptor; reducing the number of channels of the input feature map through a first fully connected layer and the channel descriptor, and using a linear rectification function to generate intermediate features; restoring the number of channels of the input feature map through a second fully connected layer and the intermediate features, and using an S-type function to generate normalized channel weights; weighting the normalized channel weights to the input feature map channel by channel to obtain an enhanced feature map, so as to complete the addition of the SE attention mechanism to the YOLOv5 model.
[0010] In one embodiment of the present application, based on the aforementioned scheme, the method further includes: obtaining predicted box information, true box information and minimum bounding box information in the target data set; based on the predicted box information, true box information and minimum bounding box information, respectively calculating the width loss and height loss of the predicted box and the true box, thereby obtaining a target loss function; applying the target loss function to the YOLOv5 model to complete the introduction of the EIOU loss function in the YOLOv5 model.
[0011] In one embodiment of the present application, based on the aforementioned scheme, obtaining an initial detection image includes: obtaining a first preset number of first images and a second preset number of second images, the first preset number being less than the second preset number; upsampling the first image until the difference between the first preset number and the second preset number is less than a preset number threshold; merging the upsampled first image and the second image to obtain the initial detection image.
[0012] In one embodiment of the present application, based on the aforementioned scheme, after the initial detection image is annotated according to the helmet wearing label to generate a sample data set with a helmet wearing label, the method further includes: extracting a sample verification set from the sample data set to verify the helmet wearing detection model according to the sample verification set, and then evaluating the performance of the helmet wearing detection model.
[0013] In one embodiment of the present application, based on the aforementioned scheme, the improved YOLOv5 model includes an input end, a backbone network, a neck network and a prediction head. The input end is used to preprocess the input image and then input it into the backbone network. The backbone network is used to extract the multi-scale features of the input image and then input it into the neck network. The neck network is used to fuse the multi-scale features to generate a target feature map and then input it into the prediction head. The prediction head is used to perform helmet wearing detection according to the target feature map to obtain a helmet wearing detection result.
[0014] According to one aspect of an embodiment of the present application, a helmet wearing detection device for a two-wheeled vehicle is provided, comprising: a first acquisition module for acquiring an initial detection image and a helmet wearing label; a labeling module for labeling the initial detection image according to the helmet wearing label, generating a sample data set with the helmet wearing label, and extracting a sample training set from the sample data set; a data enhancement module for performing data enhancement on the sample training set by a preset data enhancement method to obtain a target data set; an improvement module for obtaining an improved YOLOv5 model by introducing a K-means clustering algorithm, adding a SE attention mechanism, and introducing an EIOU loss function into the YOLOv5 model; a model training module for training the improved YOLOv5 model based on the target data set to obtain a helmet wearing detection model; a second acquisition module for acquiring an image to be tested, the image to be tested including an image of a driver or passenger of a target two-wheeled vehicle; a detection module for performing helmet wearing detection on the image to be tested by using the helmet wearing detection model to obtain a helmet wearing detection result for the target two-wheeled vehicle.
[0015] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the electronic device implements a helmet wearing detection method for a two-wheeled vehicle as described in any one of the above embodiments.
[0016] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor of a computer, the computer is caused to execute the helmet wearing detection method for a two-wheeled vehicle as described in any one of the above embodiments.
[0017] The beneficial effects of the present application are as follows: the present application obtains an initial detection image and a helmet wearing label; annotates the initial detection image according to the helmet wearing label to generate a sample data set with a helmet wearing label, and extracts a sample training set from the sample data set; performs data enhancement on the sample training set through a preset data enhancement method to obtain a target data set; obtains an improved YOLOv5 model by introducing the K-means clustering algorithm, adding the SE attention mechanism, and introducing the EIOU loss function into the YOLOv5 model; trains the improved YOLOv5 model based on the target data set to obtain a helmet wearing detection model; obtains an image to be tested, which includes an image of the driver and passenger of a target two-wheeled vehicle; performs helmet wearing detection on the image to be tested through the helmet wearing detection model to obtain a helmet wearing detection result of the target two-wheeled vehicle, so as to improve the efficiency and accuracy of helmet wearing detection for two-wheeled vehicles, and the YOLOv5 model has low requirements for computing resources and is easy to deploy in various application environments including edge devices.
[0018] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, serving to explain the principles of the present application. It is obvious that the drawings described below are merely some embodiments of the present application, and a person of ordinary skill in the art can derive other drawings based on these drawings without inventive effort. In the drawings: Figure 1 1 is a flow chart of a method for detecting helmet wearing on a two-wheeled vehicle according to an exemplary embodiment of the present application; Figure 2 1 is a schematic diagram of the structure of an improved YOLOv5 model for detecting helmet wearing on two-wheeled vehicles, shown in an exemplary embodiment of the present application; Figure 3 is a flow chart of a method for detecting helmet wearing on a two-wheeled vehicle according to another exemplary embodiment of the present application; Figure 4 is a PR graph of the improved YOLOv5 model shown in an exemplary embodiment of the present application; Figure 5 1 is a detection effect diagram of the improved YOLOv5 model shown in an exemplary embodiment of the present application; FIG6( a ) is a schematic diagram showing a comparison of PR curves before and after the SE attention mechanism is added, as shown in an exemplary embodiment of the present application; FIG6( b ) is a schematic diagram showing a comparison of the detection effects of a faint target sensitivity test before and after the addition of the SE attention mechanism, as shown in an exemplary embodiment of the present application; FIG6( c ) is a schematic diagram showing a comparison of the detection effects of the shooting angle sensitivity test before and after the addition of the SE attention mechanism, shown in an exemplary embodiment of the present application; FIG6( d ) is a schematic diagram showing a comparison of the detection effects of a joint sensitivity test of faint targets and shooting angles before and after the addition of the SE attention mechanism, as shown in an exemplary embodiment of the present application; Figure 7 is a block diagram of a helmet wearing detection device for a two-wheeled vehicle, shown in an exemplary embodiment of the present application; Figure 8 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0020] The following will describe the embodiments of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand the other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for the purpose of illustrating the present application and are not intended to limit the scope of protection of the present application.
[0021] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the illustrations only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0022] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0023] In recent years, with the rapid development of computer vision and deep learning technologies, helmet detection methods based on image or video processing have gradually become a research hotspot. However, these methods still have many shortcomings when applied to the real-time monitoring and accurate judgment of helmet wearing by two-wheeled riders. On the one hand, these methods often require high computing resources and are prone to processing performance bottlenecks when running on edge or embedded devices, making it difficult to meet the needs of real-time monitoring. On the other hand, existing solutions mostly focus on detecting helmets worn by motorcycles (including electric vehicles), while relatively little research has been conducted on helmet detection for riders of light two-wheeled vehicles (such as bicycles and electric scooters), making it difficult to adapt to the diverse riding scenarios of two-wheeled vehicles.
[0024] YOLO has been iterated and updated to 11 generations. Among them, the YOLOv5 model has been widely used in edge object detection due to its easy integration with edge devices and its fast detection speed in the field of general object detection.
[0025] However, when it comes to detecting whether a two-wheeled cyclist is wearing a helmet, the detection effect of the YOLOv5 model is less than ideal, and missed detections and false detections often occur, due to factors such as the small size and easy obstruction of the detected target, susceptibility to interference from weather factors or camera shooting angles, and complex backgrounds. This has caused great trouble for traffic law enforcement police or urban traffic managers. Therefore, in order to improve the real-time detection effect of wearing a bicycle helmet, and in response to the diverse interference factors faced by urban two-wheeled cyclists wearing helmets, this application specifically proposes an improved YOLOv5 model to improve the accuracy and efficiency of two-wheeled cyclist helmet wearing detection. While ensuring the lightweight of the original YOLOv5 model, the model greatly improves the accuracy of two-wheeled cyclist helmet wearing detection under different shooting angles and dense traffic.
[0026] First of all, it should be noted that in image data processing, upsampling is a technique to increase image resolution or the number of samples, which is mainly used to solve category imbalance problems or generate more training data.
[0027] ReLU (Rectified Linear Unit), also known as the rectified linear unit, is an activation function commonly used in artificial neural networks. It usually refers to nonlinear functions represented by ramp functions and their variants.
[0028] The sigmoid function is a common S-shaped function in biology, also known as the S-shaped growth curve. In information science, due to its monotonic increasing properties and the monotonic increasing properties of its inverse function, the sigmoid function is often used as an activation function in neural networks to map variables between 0 and 1.
[0029] The SE attention mechanism dynamically adjusts the weights of feature channels, enhancing the model's ability to focus on key features while suppressing irrelevant ones. Its core steps are: 1. Squeeze: Capturing global contextual information and generating channel descriptors. 2. Excitation: Adaptively learning the dependencies between channels, generating channel weights, and applying them to the feature map.
[0030] The K-means clustering algorithm (k-means clustering algorithm) is an iterative cluster analysis algorithm. Its steps are as follows: first, the data is divided into K groups, then K objects are randomly selected as the initial cluster centers. Then, the distance between each object and each seed cluster center is calculated, and each object is assigned to the cluster center closest to it. The cluster centers and the objects assigned to them represent a cluster. With each assignment of a sample, the cluster center is recalculated based on the existing objects in the cluster. This process is repeated until a termination condition is met. The termination condition can be that no (or a minimum number of) objects are reassigned to different clusters, no (or a minimum number of) cluster centers change again, or the sum of squared errors reaches a local minimum.
[0031] The Efficient Intersection over Union Loss (EIOU) loss, also known as the EIoU loss, is a loss function used in object detection tasks, designed to improve the accuracy and efficiency of the model in object detection. The EIOU loss is an improved version of the CIoU loss. By decomposing the width and height losses into independent terms, it simplifies computation and improves optimization efficiency, while maintaining precise control over the shape of the predicted box.
[0032] Gradient Backforward (GDB) often refers to the backpropagation algorithm, an important algorithm for training neural networks. Backpropagation minimizes the loss function by calculating the gradient of the loss function with respect to the network parameters and using these gradients to update the network parameters.
[0033] Figure 1 FIG1 is a flow chart of a method for detecting helmet wearing on a two-wheeled vehicle according to an exemplary embodiment of the present application. Figure 1 As shown, the method for detecting helmet wearing on a two-wheeled vehicle includes at least steps S110 to S170, which are described in detail as follows: In step S110 , an initial detection image and a helmet wearing tag are acquired.
[0034] In one embodiment of the present application, a first preset number of first images and a second preset number of second images are obtained, where the first preset number is less than the second preset number; the first images are upsampled until the difference between the first preset number and the second preset number is less than a preset number threshold; the upsampled first and second images are merged to obtain an initial detection image.
[0035] In this embodiment, the helmet wearing label includes whether a helmet is worn and whether a helmet is not worn. Two-wheeled vehicles can be divided into motorcycles, bicycles, and electric scooters according to their type, where motorcycles include electric vehicles and bicycles include electric bicycles. Self-made helmet wearing detection dataset for multiple types of two-wheeled vehicles. This application collects relevant image data of a preset number of bicycles (including electric bicycles) and electric scooters to solve the problem that the helmet wearing detection dataset for two-wheeled vehicles in the related art mainly targets motorcycles (including electric vehicles) and ignores light two-wheeled vehicles such as bicycles, electric bicycles, or electric scooters.
[0036] In this embodiment, upsampling techniques are used to upsample the first image to a data volume comparable to that of motorcycle images. A self-developed first image dataset is then merged with a publicly available second image dataset to generate a combined initial detection image. This increases the amount of training data for the model, thereby enhancing its interference tolerance and generalization capabilities. Upsampling techniques include, but are not limited to, random sampling with replacement, synthesizing new bicycle samples in feature space, and using the Synthetic Minority Over-sampling Technique (SMOTE) algorithm to generate new data points by interpolating adjacent samples.
[0037] In step S120, the initial detection image is annotated according to the helmet wearing label to generate a sample data set with the helmet wearing label, and a sample training set is extracted from the sample data set.
[0038] In one embodiment of the present application, the first image can be manually annotated; then, an auxiliary data labeling model based on YOLOv5 is constructed on this annotated data, wherein the auxiliary data labeling model requires manual verification of its detection effect, and any deviations or errors need to be manually corrected.
[0039] In one embodiment of the present application, a sample validation set is extracted from a sample data set to validate a helmet wearing detection model based on the sample validation set, thereby evaluating the performance of the helmet wearing detection model.
[0040] In step S130, data enhancement is performed on the sample training set using a preset data enhancement method to obtain a target data set.
[0041] In one embodiment of the present application, the preset data enhancement methods include but are not limited to mosaic, rotation, cropping and stretching, etc. The sample training set is enhanced by the above-mentioned preset data enhancement methods to increase the diversity of the data set and improve the detection effect of small targets.
[0042] In step S140, an improved YOLOv5 model is obtained by introducing the K-means clustering algorithm, adding the SE attention mechanism, and introducing the EIOU loss function into the YOLOv5 model.
[0043] In one embodiment of the present application, the width and height of all target boxes are extracted from the target data set; a preset number of target boxes are randomly selected as initial cluster centers, and each target box is assigned to the initial cluster center with the closest relative distance to obtain a preset number of initial clusters, wherein the relative distance is calculated based on the width and height of the target box; a new cluster center of each initial cluster is obtained according to the width and height of the target box in each initial cluster; the iterative allocation is repeated until a preset end condition is met, and the new cluster center is determined as the target anchor box, wherein the preset end condition includes that the change in the new cluster center is less than a preset change, or the number of iterations is greater than or equal to a preset number threshold; the target anchor box is applied to the YOLOv5 model to complete the introduction of the K-means clustering algorithm in the YOLOv5 model.
[0044] In this embodiment, K-means clustering anchor frames are introduced into the classic YOLOv5 model. Since the original YOLOv5 model uses three preset anchor frames of different sizes for target frame prediction, for large-scale, small-target detection tasks, the preset anchor frames may not fit the boundaries of two-wheeled vehicles or helmets well in actual detection. In order to better fit the real target frame, the K-means clustering algorithm is introduced to optimize the anchor frame mechanism of the original YOLOv5. The specific approach is as follows: the width and height of each target frame are extracted from the training data; the width and height of these target frames are used as input and optimized using the K-means clustering algorithm. K-means will try to divide the width and height of these target frames into K clusters, where K is the set number of anchor frames. The clustering result will give the sizes of K anchor frames; the anchor frames generated by K-means clustering are used to replace the default anchor frames of YOLOv5, thereby improving the detection accuracy of the model in target detection tasks.
[0045] In some embodiments, the width and height (in pixels) of all target boxes are extracted from the annotation file of the sample training set. The annotation format is ensured to be YOLO-compatible and converted to absolute pixel values. The width and height of the target boxes are normalized to the range [0, 1] based on the image size. K target boxes are randomly selected as initial cluster centers (the number of anchor boxes K can be a preset number, typically consistent with the YOLOv5 default, such as 9 anchor boxes divided into 3 scales). Intersection over Union (IoU) is used as the distance metric instead of the traditional Euclidean distance because it better reflects the degree of match between anchor boxes and target boxes. It is understood that IoU is the ratio of the overlapping area of two bounding boxes to the area of their union. The IoU between bounding boxes A and B is denoted as IoU(A,B), and the IoU distance is defined as d(A,B)=1−IoU(A,B). Each target box is assigned to the cluster center with the closest IoU distance. Recalculate the center (anchor box size) of each cluster by taking the mean width and height of all target boxes in that cluster. Repeat until the cluster center no longer changes significantly or the maximum number of iterations is reached. Sort the K anchor box sizes after clustering (optionally by area, small to large) to ensure compatibility with YOLOv5's multi-scale detection layer. Restore the normalized anchor box sizes to absolute pixel values based on the image size. Modify the YOLOv5 configuration file (e.g., yolov5s.yaml) to replace the default anchor boxes with the target anchor boxes generated by clustering.
[0046] In one embodiment of the present application, the spatial dimension of each channel in the input feature map is compressed by global average pooling to generate a channel descriptor; the number of channels of the input feature map is reduced by a first fully connected layer and the channel descriptor, and a linear rectification function is used to generate intermediate features; the number of channels of the input feature map is restored by a second fully connected layer and the intermediate features, and a S-type function is used to generate normalized channel weights; the normalized channel weights are weighted channel by channel to the input feature map to obtain an enhanced feature map, thereby completing the addition of the SE attention mechanism to the YOLOv5 model.
[0047] In this embodiment, the SE network attention mechanism is introduced into the backbone network of the YOLOv5 model. In its input channel, the features of bicycle helmets worn under different shooting angles, different climate conditions, and various traffic scenarios are compressed (Squeeze) and excited (Excitation) by assigning an attention weight to each feature channel. This allows the convolutional network to pay more attention to important feature channels and suppress feature channels that are less useful for the current task, thereby improving the performance of the model. The operation steps of the SE attention mechanism mainly include two stages: compression and excitation: Among them, the compression stage is used to capture global information. Considering the channel dependency problem of the YOLOv5 model, each learned filter operates within a local receptive field, making it impossible to utilize contextual information outside the area. The compression operation compresses the global spatial information into the channel descriptor and generates channel statistics by using the global average pool. Formally speaking, by H×W Compressing the convolution operator output U of the module can generate a statistic , where the cth element of z The calculation formula is reference formula (1): Formula (1) Among them, z c is the cth element, H is the height, W is the width, u c is the cth feature map.
[0048] Through the global average pooling operation, the two-dimensional features (H×W) of each channel are compressed into a real number, thereby converting the feature map from [h, w, c] to [1, 1, c]. This step realizes the fusion of global context information, and each channel is represented by a numerical value.
[0049] The excitation stage is used for adaptive calibration. To capture inter-channel dependencies while ensuring that multiple channels are emphasized simultaneously, a weight value is generated for each feature channel. Inter-channel correlation is typically established through two fully connected layers. The first fully connected layer reduces the number of channels in the feature map (for example, to 1 / 4 or 1 / 16 of the original number) to reduce computational effort and uses the ReLU activation function. The second fully connected layer restores the number of channels to their original size and uses the Sigmoid activation function to map the weight values to between 0 and 1. The resulting normalized weights are then applied to the features of each channel. This is typically achieved by multiplying each channel by a weight coefficient, resulting in feature maps with different weights. The final output of the block is obtained by rescaling the transformed output U with the activation.
[0050] Formula (2) In formula (2), , Represents the feature map and scaling factor These activation values act as channel weights and are adjusted according to the input-specific descriptor z. In this case, the SE block essentially introduces dynamic characteristics based on the input, which helps to enhance the discriminative ability of the features.
[0051] The addition of the SE attention mechanism further enhances the YOLOv5 model's detection performance for subtle objects and improves the model's robustness to interference from factors such as shooting angle and lighting conditions. The SE network is a lightweight module whose overhead can be neglected and integrated into the backbone network.
[0052] The SE attention mechanism was added to the original YOLOv5 network to address the loss caused by the varying weights of different channels in the feature map during pooling. Adding the SE module to the backbone learns the correlations between channels and selects channel-specific attention, effectively improving detection accuracy. The backbone network consists of the Conv module, the C3 module, and the SPPF module. Its primary function is to convert the original input image into a multi-layer feature map for subsequent object detection tasks.
[0053] In one embodiment of the present application, the predicted box information, the true box information and the minimum bounding box information in the target data set are obtained; based on the predicted box information, the true box information and the minimum bounding box information, the width loss and the height loss of the predicted box and the true box are calculated respectively, and then the target loss function is obtained; the target loss function is applied to the YOLOv5 model to complete the introduction of the EIOU loss function in the YOLOv5 model.
[0054] In this embodiment, the loss function CIoU in the YOLOv5 model is replaced by EIoU. The original YOLOv5 model's loss function is CIoU, which comprehensively considers the overlap between the predicted box and the ground truth box, the center point distance, and the aspect ratio. Its calculation method is shown in the following formula.
[0055] Formula (3) Among them, IoU represents intersection over union, Represent the center points of the predicted box and the real box respectively, represents the Euclidean distance between the center point of the predicted box and the center point of the real box, c represents the diagonal length of the minimum bounding box, is the weight coefficient, v is used to measure the consistency of aspect ratio, Represent the width and height of the real box respectively, Represents the width and height of the predicted box respectively.
[0056] Although CIoU comprehensively considers the overlap, center point distance, and aspect ratio of the predicted box and the true box, its description of the aspect ratio is measured by the length-to-width ratio of the box. EIoU splits the aspect ratio in CIoU and calculates the width and height losses of the predicted box and the true box respectively. This allows the predicted box to fit the true box more closely in width and height, making the predicted box more accurate. The calculation method is shown in the following formula.
[0057] Formula (4) Among them, IoU represents intersection over union, Represents the Euclidean distance between the center point of the predicted box and the center point of the real box, ) represents the square difference between the width of the predicted box and the real box, ) represents the square difference between the height of the predicted box and the real box, Represents the width and height of the minimum bounding box, which is the smallest rectangular box that contains both the predicted box and the true box.
[0058] In one embodiment of the present application, the improved YOLOv5 model includes an input end, a backbone network, a neck network and a prediction head. The input end is used to preprocess the input image and then input it into the backbone network. The backbone network is used to extract multi-scale features of the input image and then input it into the neck network. The neck network is used to fuse the multi-scale features and generate a target feature map, which is then input into the prediction head. The prediction head is used to perform helmet wearing detection based on the target feature map to obtain a helmet wearing detection result.
[0059] The improved YOLOv5 structure is as follows Figure 2 As shown, Figure 2 This is a schematic diagram of the improved YOLOv5 model structure for a two-wheeled vehicle helmet wearing detection method shown in an exemplary embodiment of the present application. After the improvement of the YOLOv5 model structure and loss function is completed, the target data set is input into the improved YOLOv5 model for training. Figure 2 As can be seen, the improved YOLOv5 model has made significant improvements in input, backbone network, and overall model optimization loss compared to the YOLOv5 model. In particular, the SE network attention mechanism module has been added to the backbone network to enhance small target detection and overcome interference from factors such as different shooting angles and different weather conditions, making the model more robust and anti-interference.
[0060] In this embodiment, referring to Figure 2As shown, this application generates K anchor boxes by introducing the K-means clustering algorithm at the input end of the YOLOv5 model, so that the generated K anchor boxes are used for the regression of the auxiliary bounding box in the subsequent prediction stage, and by adding the SE attention mechanism to the backbone network, specifically adding an SE module between the first convolutional layer and the second convolutional layer, adding an SE module between the second convolutional layer and the first C3 module, adding an SE module between the third convolutional layer and the second C3 module, adding an SE module between the fourth convolutional layer and the third C3 module, and also introducing the EIOU loss function to obtain an improved YOLOv5 model. It can also be understood that the SE module is added after each convolutional layer in the backbone network to add the SE attention mechanism to the YOLOv5 model.
[0061] Continue to refer to Figure 2 As shown in the figure, the improved YOLOv5 model has the following significant changes and characteristics compared to the traditional YOLOv5 model. The following is a structural description of the improved YOLOv5 model: Input: In addition to the conventional data enhancements (such as flipping, scaling, etc.) that may be included in the traditional YOLOv5 model, Figure 2 The improved model also incorporates richer data augmentation techniques, such as mosaicing, rotation, cropping, and stretching, which help improve the model's robustness to different scenes and object poses. It also uses the K-Means algorithm to generate K anchor boxes (Anchor1, Anchor2, …, AnchorK). These anchor boxes assist in bounding box regression in the subsequent prediction stage, better adapting to the size and shape distribution of objects in the dataset.
[0062] Backbone network: The convolution layer (Conv) retains the convolution operation for feature extraction, but the parameters and structure of the convolution layer may be adjusted to optimize the feature extraction capability. The SE (Squeeze-and-Excitation) module, also known as the SE Network, is introduced. The SE module can adaptively recalibrate channel feature responses and enhance useful feature channels and suppress useless feature channels by explicitly modeling the interdependence between channels, thereby improving the quality of feature representation. The C3 module is an efficient feature extraction module in YOLOv5. It combines the BottleneckCSP structure and convolution operation to reduce the amount of computation while maintaining good feature extraction capabilities. The SPPF (Spatial Pyramid Pooling-Fast) fast spatial pyramid pooling module can convert feature maps of different scales into fixed-size feature vectors, which helps the model process objects of different scales and can fuse multi-scale feature information.
[0063] The neck network (Neck) is mainly used to fuse feature maps at different levels, fusing the shallow and deep features extracted by the backbone network to provide richer semantic information and position information to the head network for prediction.
[0064] The convolutional layer (Conv) in the prediction head performs further feature transformation and adjustments on the fused feature maps. Its anchor mechanism uses K-Means clustering to generate K anchor boxes, each of which corresponds to a prediction branch. Each branch predicts the bounding box coordinates (x, y, w, h), confidence, and class score.
[0065] The prediction part (Prediction) ultimately outputs multiple prediction results, each of which corresponds to the prediction of an anchor box, including the location information of the bounding box, the confidence level, and the probability score of the category to which the object belongs.
[0066] Loss Function and Gradient Backpropagation. EIoU Loss uses the EIoU loss function to measure the difference between the predicted bounding box and the ground-truth bounding box. The EIoU loss function improves on the traditional IoU loss by considering not only the intersection over union (IoU) but also the center point distance and aspect ratio differences of the bounding box, enabling more accurate model training. Through the gradient backpropagation algorithm, the gradient of the loss function is propagated back to the network, updating the network parameters and continuously optimizing the model during training.
[0067] In step S150, the improved YOLOv5 model is trained based on the target data set to obtain a helmet wearing detection model.
[0068] In one embodiment of the present application, a target dataset is input into an improved YOLOv5 model, which is then trained and optimized to obtain a helmet wearing detection model. The helmet wearing detection model is then validated using a sample validation set to evaluate its performance.
[0069] In step S160 , an image to be tested is acquired.
[0070] In one embodiment of the present application, the image to be tested may be acquired by an image acquisition device in an edge device or an embedded device in which a helmet wearing detection model is deployed, and the image to be tested includes an image of a driver or passenger of a target two-wheeled vehicle.
[0071] In step S170, a helmet wearing detection model is used to perform helmet wearing detection on the image to be tested, and a helmet wearing detection result of the target two-wheeled vehicle is obtained.
[0072] In one embodiment of the present application, a helmet-wearing detection result is used to determine whether the driver or passenger of a target two-wheeled vehicle is wearing a helmet, and the relevant data is stored. Target information, such as driver or passenger information, driving position information, and target vehicle information, can also be detected based on the image to be tested. The image to be tested, showing the driver or passenger not wearing a helmet, and the target information can be sent to a preset user to assist in monitoring.
[0073] In one embodiment of the present application, referring to Figure 3 , Figure 3 This is a flow chart of a helmet wearing detection method for two-wheeled vehicles shown in another exemplary embodiment of the present application, wherein a self-collected dataset and a public dataset are obtained, wherein the self-collected dataset includes first images of bicycles and electric scooters, and the public dataset includes second images of motorcycles and electric vehicles, and the self-collected dataset is annotated with the assistance of manual and model assistance, and the self-collected dataset and the public dataset are merged to obtain merged data, i.e., the initial detection image, and the merged data is enhanced by preset data enhancement methods such as upsampling, mosaicing, rotation, cropping, and stretching. An improved YOLOv5 model is obtained by introducing the KMeans clustering anchor frame and SE network attention mechanism into the YOLOv5 model and replacing the original loss function CIoU with EIoU to perform detection and output the detection results. The specific implementation methods of each step have been described in detail in the aforementioned embodiments and will not be repeated here.
[0074] In one embodiment of the present application, in order to verify the effectiveness of the improved YOLOv5 model in detecting whether helmets are worn when riding multiple types of two-wheeled vehicles, self-constructed bicycle and electric scooter datasets (i.e., self-collected datasets) and open source dataset TWHD (two wheeler helmet dataset), i.e., public datasets, are used to locate and classify the two-wheeled vehicles and their riders as a whole, the heads of people without helmets, and the heads of people wearing helmets in the images.
[0075] For self-collected datasets, a preset number (e.g., 1,000) of first images of bicycles and electric scooters can be downloaded by searching the web for keywords such as "bicycle riding," "electric bicycle riding," and "electric scooter riding," along with similar photos. Manual annotation was then performed using the LabelImg graphical image annotation tool to annotate 300 images (100 each for bicycles, electric bicycles, and electric scooters, with 50 images each indicating whether a helmet was worn). Based on this annotated data, a data-assisted annotation model based on the YOLOv5 model was constructed to output the approximate location of the target and improve annotation efficiency. Finally, manual verification and verification were performed to complete the annotation of the unlabeled data.
[0076] The open source dataset TWHD comes from the OSF dataset, the bike helmet dataset, and web crawlers. For example, 4,710 images were randomly extracted from the OSF dataset and re-labeled. In addition, to enrich the dataset context, 738 images from the bike helmet dataset and web crawlers were integrated, totaling 5,448 images for model training, testing, and verification.
[0077] It should be noted that in order to facilitate subsequent training, testing and verification, the self-collected data can be randomly divided into training set, test set and verification set in a ratio of 8:1:1.
[0078] During YOLOv5 model training, you need to set various hyperparameters and configure image enhancement techniques. These configuration items determine the learning rate adjustment strategy, loss function weights, and input image enhancement methods during model training. By adjusting these parameters, you can optimize model training results and improve model accuracy and generalization capabilities.
[0079] This application organizes two parts of experiments to verify the effectiveness of the proposed method. First, the overall performance of the improved algorithm is introduced. Second, the performance of helmet wearing detection for two-wheeled vehicles is compared before and after the SE attention mechanism is added, demonstrating the effectiveness of the SE attention mechanism.
[0080] Among them, in order to verify the performance of the improved YOLOv5 model, the PR graph of the improved YOLOv5 model was drawn, refer to Figure 4 , Figure 4 yes Figure 4 This is a PR diagram of the improved YOLOv5 model shown in an exemplary embodiment of this application, which also intuitively demonstrates the detection effect of the model. Figure 5 , Figure 5 This is a detection effect diagram of the improved YOLOv5 model shown in an exemplary embodiment of the present application. Figure 4 As can be seen from the figure, the model achieved an average detection accuracy (AP) of 0.965 for multiple types of two-wheeled vehicles, an AP of 0.854 for two-wheeled vehicles wearing helmets, and an AP of 0.719 for two-wheeled vehicles without helmets. The average recognition accuracy (mAP) for whether a two-wheeled vehicle was wearing a helmet (no helmet was detected as not wearing a helmet, and wearing a helmet was detected as wearing a helmet) reached 0.846. Figure 5 It can be seen that the model has a good effect on detecting whether different types of cyclists are wearing helmets in most scenarios.
[0081] In order to verify the effectiveness of adding the SE attention mechanism to the model, the PR curves before and after the addition of the SE attention mechanism were drawn, and the detection effects were compared. Referring to Figures 6(a)-(d), Figure 6(a) is a schematic diagram of the comparison of PR curves before and after the addition of the SE attention mechanism shown in an exemplary embodiment of the present application, Figure 6(b) is a schematic diagram of the comparison of detection effects of the faint target sensitivity test before and after the addition of the SE attention mechanism shown in an exemplary embodiment of the present application, Figure 6(c) is a schematic diagram of the comparison of detection effects of the shooting angle sensitivity test before and after the addition of the SE attention mechanism shown in an exemplary embodiment of the present application, and Figure 6(d) is a schematic diagram of the comparison of detection effects of the joint sensitivity test of faint targets and shooting angles before and after the addition of the SE attention mechanism shown in an exemplary embodiment of the present application. As shown in the PR graph in Figure 6(a), after incorporating the SE attention mechanism, the model's AP for two-wheeled vehicles increased from 0.911 to 0.965, the AP for two-wheeled vehicle riders wearing helmets increased from 0.772 to 0.854, and the AP for two-wheeled vehicle riders without helmets increased from 0.500 to 0.719. The mAP for detecting whether a two-wheeled vehicle rider was wearing a helmet (detecting a vehicle without a helmet as not wearing a helmet and a vehicle rider as wearing a helmet as not wearing a helmet) increased from 0.748 to 0.846. A comparison of the detection results in Figure 6(b) shows that before SE attention was implemented, the two-wheeled vehicle (a faint target) in the distance behind was significantly missed. However, after implementing the SE attention mechanism, the two-wheeled vehicle behind was correctly detected, and the rider was not wearing a helmet (no helmet label). The detection results in Figure 6(d) also show that the SE attention mechanism significantly improves the detection of faint targets. Figure 6(c) shows the comparison of the detection effect before and after the addition of the SE attention mechanism at different shooting angles. It can be seen that after the addition of SE, the model's sensitivity to angles is reduced, and the detection effect is significantly improved. In addition, the addition of the SE attention mechanism has successfully improved the model's confidence in detecting two-wheeled vehicles and helmets. For example, before the addition of the SE attention mechanism, although the two-wheeled vehicle and the rider wearing a helmet were correctly identified, their confidence levels were only 0.88 and 0.07, respectively. However, after the addition of the SE attention mechanism, the confidence levels for the two-wheeled vehicle and the helmet were increased to 0.89 and 0.31, respectively. These results demonstrate that the addition of the SE attention mechanism effectively improves the model's accuracy in detecting two-wheeled vehicles and whether a helmet is worn.
[0082] It can be understood that this embodiment is only for illustration, and the data such as the number of collected images, the ratio of data set division and detection results are all example data. This application does not limit this and should not bring any limitations to the functions and scope of use of the embodiments of this application.
[0083] In addition, the vehicle data and other user data obtained in the embodiments of the present application are all obtained with the user's consent, or are actively submitted after the user's relevant instructions, or are necessarily uploaded when the user uses the corresponding application through the client, web page, etc. In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure and application of the user's personal information involved are in compliance with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and do not violate public order and good morals.
[0084] This application proposes a lightweight, improved YOLOv5 model designed to quickly and accurately detect whether riders of various types of two-wheeled vehicles are wearing helmets. By optimizing the network structure and reducing the amount of computation, this model effectively reduces the demand for computing resources, making it easier to deploy on resource-constrained devices such as edge or embedded devices. This improves the efficiency and coverage of helmet monitoring, providing a strong technical guarantee for two-wheeled riding safety.
[0085] Figure 7 This is a block diagram of a helmet wearing detection device for a two-wheeled vehicle, shown as an exemplary embodiment of the present application. The device can be applied to Figure 1 The implementation environment shown is specifically configured in the computer device 102. The apparatus may also be applicable to other exemplary implementation environments and specifically configured in other devices. This embodiment does not limit the implementation environment to which the apparatus is applicable.
[0086] like Figure 7 As shown, the exemplary helmet wearing detection device for two-wheeled vehicles includes: a first acquisition module 710, a labeling module 720, an enhancement module 730, an improvement module 740, a model training module 750, a second acquisition module 760 and a detection module 770.
[0087] Among them, the first acquisition module 710 is used to obtain the initial detection image and the helmet wearing label; the annotation module 720 is used to annotate the initial detection image according to the helmet wearing label, generate a sample data set with the helmet wearing label, and extract a sample training set from the sample data set; the data enhancement module 730 is used to perform data enhancement on the sample training set through a preset data enhancement method to obtain a target data set; the improvement module 740 is used to obtain an improved YOLOv5 model by introducing the K-means clustering algorithm, adding the SE attention mechanism, and introducing the EIOU loss function into the YOLOv5 model; the model training module 750 is used to train the improved YOLOv5 model based on the target data set to obtain a helmet wearing detection model; the second acquisition module 760 is used to obtain the image to be tested, which includes the image of the driver and passenger of the target two-wheeled vehicle; the detection module 770 is used to perform helmet wearing detection on the image to be tested through the helmet wearing detection model to obtain the helmet wearing detection result of the target two-wheeled vehicle.
[0088] It should be noted that the helmet wearing detection device for a two-wheeled vehicle provided in the above embodiment and the helmet wearing detection method for a two-wheeled vehicle provided in the above embodiment are based on the same concept. The specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here. In actual applications, the helmet wearing detection device for a two-wheeled vehicle provided in the above embodiment can, as needed, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above, and this is not limited here.
[0089] An embodiment of the present application also provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the electronic device implements the helmet wearing detection method for a two-wheeled vehicle provided in the above-mentioned embodiments.
[0090] Figure 8 The following is a schematic diagram showing the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application. Figure 8 The computer system 800 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0091] like Figure 8As shown, computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes, such as executing the methods provided in the various embodiments described above, based on programs stored in read-only memory (ROM) 802 or programs loaded from storage 808 into random access memory (RAM) 803. RAM 803 also stores various programs and data required for system operation. CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0092] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, mouse, and the like; an output section 807 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 808 including a hard disk; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. Removable media 811, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, is installed in the drive 810 as needed, so that computer programs read from the media can be installed in the storage section 808 as needed.
[0093] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 809 and / or installed from removable media 811. When executed by the central processing unit (CPU) 801, the computer program performs the various functions defined in the system of the present application.
[0094] It should be noted that the computer-readable medium described in the embodiments of this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may, for example, be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. This propagated data signal may take a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0096] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0097] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When executed by a computer processor, the computer program causes the computer to perform the two-wheeled vehicle helmet wearing detection method provided in the above-described embodiments. The computer-readable storage medium may be included in the electronic device described in the above-described embodiments, or may exist independently and not be incorporated into the electronic device.
[0098] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0099] Another aspect of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the helmet wearing detection method for a two-wheeled vehicle provided in each of the above embodiments.
[0100] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0101] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
[0102] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, any equivalent modifications or alterations accomplished by a person of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.
Claims
1. A method for detecting the wearing of a helmet for a two-wheeled vehicle, characterized in that: include: Obtain the initial detection image and helmet wearing label; Annotating the initial detection image according to the helmet wearing label to generate a sample data set with the helmet wearing label, and extracting a sample training set from the sample data set; Performing data enhancement on the sample training set using a preset data enhancement method to obtain a target data set; By introducing the K-means clustering algorithm, SE attention mechanism, and EIOU loss function into the YOLOv5 model, an improved YOLOv5 model is obtained; The improved YOLOv5 model is trained based on the target data set to obtain a helmet wearing detection model; Acquiring an image to be tested, wherein the image to be tested includes an image of a driver and passenger of a target two-wheeled vehicle; The helmet wearing detection model is used to perform helmet wearing detection on the image to be tested to obtain a helmet wearing detection result of the target two-wheeled vehicle.
2. The method for detecting helmet wearing for two-wheeled vehicles according to claim 1, characterized in that: The method further comprises: Extract the width and height of all target boxes from the target dataset; Randomly select a preset number of target frames as initial cluster centers, and assign each target frame to the initial cluster center with the closest relative distance to obtain a preset number of initial clusters, where the relative distance is calculated based on the width and height of the target frame; Calculating according to the width and height of the target box in each of the initial clusters to obtain a new cluster center of each of the initial clusters; Repeat the iterative allocation until a preset end condition is met, and determine the new cluster center as the target anchor box, wherein the preset end condition includes that the change of the new cluster center is less than a preset change amount, or the number of iterations is greater than or equal to a preset number threshold; The target anchor box is applied to the YOLOv5 model to complete the introduction of the K-means clustering algorithm in the YOLOv5 model.
3. The method for detecting helmet wearing for two-wheeled vehicles according to claim 1, characterized in that: The method further comprises: Compress the spatial dimension of each channel in the input feature map through global average pooling to generate a channel descriptor; Reducing the number of channels of the input feature map through the first fully connected layer and the channel descriptor, and using a linear rectification function to generate intermediate features; Restoring the number of channels of the input feature map through the second fully connected layer and the intermediate features, and generating normalized channel weights using a sigmoid function; The normalized channel weights are weighted channel by channel to the input feature map to obtain an enhanced feature map, thereby completing the addition of the SE attention mechanism to the YOLOv5 model.
4. The method for detecting helmet wearing for two-wheeled vehicles according to claim 1, characterized in that: The method further comprises: Obtaining predicted box information, true box information, and minimum bounding box information in the target data set; Based on the predicted box information, the true box information and the minimum bounding box information, respectively calculating the width loss and height loss of the predicted box and the true box, thereby obtaining a target loss function; The target loss function is applied to the YOLOv5 model to complete the introduction of the EIOU loss function in the YOLOv5 model.
5. The method for detecting helmet wearing for a two-wheeled vehicle according to any one of claims 1 to 4, characterized in that: Get the initial detection image, including: Acquire a first preset number of first images and a second preset number of second images, wherein the first preset number is smaller than the second preset number; Upsampling the first image until the difference between the first preset number and the second preset number is less than a preset number threshold; The upsampled first image and the second image are merged to obtain the initial detection image.
6. The method for detecting helmet wearing for a two-wheeled vehicle according to any one of claims 1 to 4, characterized in that: After labeling the initial detection image according to the helmet wearing label to generate a sample data set with the helmet wearing label, the method further includes: A sample validation set is extracted from the sample data set to validate the helmet wearing detection model based on the sample validation set, thereby evaluating the performance of the helmet wearing detection model.
7. The method for detecting helmet wearing for a two-wheeled vehicle according to any one of claims 1 to 4, characterized in that: The improved YOLOv5 model includes an input end, a backbone network, a neck network and a prediction head. The input end is used to preprocess the input image and then input it into the backbone network. The backbone network is used to extract the multi-scale features of the input image and then input it into the neck network. The neck network is used to fuse the multi-scale features to generate a target feature map and then input it into the prediction head. The prediction head is used to perform helmet wearing detection according to the target feature map to obtain a helmet wearing detection result.
8. A helmet wearing detection device for a two-wheeled vehicle, characterized in that: include: A first acquisition module is used to acquire an initial detection image and a helmet wearing label; a labeling module, configured to label the initial detection image according to the helmet wearing label, generate a sample data set with the helmet wearing label, and extract a sample training set from the sample data set; A data enhancement module is used to perform data enhancement on the sample training set using a preset data enhancement method to obtain a target data set; Improvement module, which is used to obtain an improved YOLOv5 model by introducing the K-means clustering algorithm, adding the SE attention mechanism, and introducing the EIOU loss function into the YOLOv5 model; A model training module is used to train the improved YOLOv5 model based on the target data set to obtain a helmet wearing detection model; A second acquisition module is used to acquire an image to be tested, wherein the image to be tested includes an image of a driver and passenger of a target two-wheeled vehicle; The detection module is used to perform helmet wearing detection on the image to be tested using the helmet wearing detection model to obtain a helmet wearing detection result of the target two-wheeled vehicle.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the helmet wearing detection method for a two-wheeled vehicle as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the helmet wearing detection method for a two-wheeled vehicle according to any one of claims 1 to 7.
Citation Information
Patent Citations
Helmet wearing detection method for electric vehicle driver and passengers based on improved YOLOv5 algorithm
CN116311359A
Traffic safety helmet wearing detection system and method based on improved YOLOv5
CN116977948A
Electric vehicle helmet wearing detection method based on improved YOLOv5s
CN118762337A
Safety helmet wearing condition detection method and device, electronic equipment and storage medium
CN119360413A