Lightweight electric vehicle rider helmet-wearing detection method and system

By improving the YOLOv5n network structure and loss function, the lightweight problem of the electric vehicle rider helmet wearing detection model was solved, enabling efficient detection on edge devices and improving detection accuracy and adaptability.

CN117079228BActive Publication Date: 2025-11-25CHANGZHOU UNIV

Patent Information

Application Number
CN202311120490.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2025-11-25
Estimated Expiration
2043-08-31

AI Technical Summary

Technical Problem

Existing helmet-wearing detection network models for electric vehicle riders lack lightweight design, making them difficult to deploy effectively on edge devices, and their detection accuracy needs improvement.

Method used

An improved YOLOv5n network is adopted, which replaces the C3 module of the backbone layer with the Faster_x module and the C3 model of the neck layer with the Faster_false_x module. A lightweight Involution inner convolution module is introduced, and the WIoU_v2 Loss bounding box regression loss function is optimized to construct a lightweight model with monotonic focusing mechanism and attention mechanism.

Benefits of technology

While maintaining detection accuracy, it significantly reduces the number of network parameters and computational load, lowers device memory and processor requirements, improves the model's detection capabilities in complex backgrounds, and adapts to edge deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117079228B_ABST
    Figure CN117079228B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, and more particularly to a lightweight electric vehicle rider helmet wearing detection method and system, which comprises collecting electric vehicle rider image data, constructing an electric vehicle rider helmet wearing picture data set, and pre-processing and labeling the image; an improved YOLOv5n network is constructed, and the optimal weight is obtained by training on the training set; a WIoU_v2Loss bounding box regression loss function is used. The present application solves the lightweight problem of YOLOv5n network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a lightweight electric vehicle rider helmet wearing detection method and system. BACKGROUND

[0002] With the prosperity of the country, people's life is getting better and better, and urban road traffic is becoming more and more congested. As a convenient and fast travel tool, electric vehicles have become the choice of more and more people.

[0003] Under the trend of a large number of electric vehicles and rapid growth, electric vehicle-related traffic accidents have also increased significantly, and related riders are more likely to be injured than other drivers in accidents. In the case of rider driver traffic accident death, 80% of the fatal injuries are due to head injury. Related studies have shown that proper wearing of a riding helmet can effectively prevent most head injuries and reduce the risk by 60%-70%.

[0004] Xie Jiafei et al. introduced the coordinate attention mechanism and alpha-IoU loss function to enhance the feature learning of the algorithm based on YOLOv5s algorithm to improve the detection accuracy; Xie Puxuan et al. changed the feature pyramid structure of the neck of the YOLOv5s algorithm to a bidirectional weighted structure, and introduced an alpha-CIoU loss function to improve the detection accuracy of helmet wearing.

[0005] The above method is only limited to considering the improvement of the algorithm for accuracy, but does not consider the actual deployment of the network model. At present, there is still a great demand for network models for detecting helmet wearing on deployment devices, and there is a lack of lightweight models friendly to edge deployment devices. SUMMARY

[0006] In view of the shortcomings of the existing method, the present application solves the lightweight problem of YOLOv5n network.

[0007] The technical scheme adopted by the present application is: a lightweight electric vehicle rider helmet wearing detection method, comprising the following steps:

[0008] Step 1, collect electric vehicle rider image data, construct an electric vehicle rider helmet wearing picture data set, and pre-process and label the images;

[0009] Further, the image preprocessing includes: eliminating blurred and ghosted pictures, and performing data enhancement on the remaining image data, including HSV color interference, Gaussian noise and salt and pepper noise.

[0010] Step 2, construct an improved YOLOv5n network and train the optimal weight on the training set;

[0011] Further, the improved YOLOv5n network comprises: semantic information extraction by a backbone layer, the backbone layer comprising 4 CBS convolution modules, 2 Faster_3 modules, 2 Faster_6 modules, 1 Involution inner convolution module and one SPPF spatial pyramid pooling module; semantic fusion by a neck layer, the neck layer comprising 2 CBS convolution modules, 2 up-sampling modules, 4 Faster_False_3 modules and 2 Involution inner convolution modules, and finally generating three different scale feature maps for prediction through a Conv2d convolution layer.

[0012] Further, the Faster_3 module is composed of a CBS module and 3 Faster blocks, the input tensor is halved in channel dimension through two CBS modules, the output tensor of the first CBS module is subjected to feature extraction through 3 Faster blocks, and then the output tensor of the second CBS module is subjected to Concat splicing and then the third CBS module, so as to obtain an output tensor with the same dimension as the original input tensor.

[0013] Further, the Faster_6 module is composed of a CBS module and 6 Faster blocks, the input tensor is halved in channel dimension through two CBS modules, the output tensor of the first CBS module is subjected to feature extraction through 6 Faster blocks, and then the output tensor of the second CBS module is subjected to Concat splicing and then the third CBS module, so as to obtain an output tensor with the same dimension as the original input tensor.

[0014] Further, the Faster_False_3 module is composed of a CBS module and 3 Faster_False blocks, the input tensor is halved in channel dimension through two CBS modules, the output tensor of the first CBS module is subjected to feature extraction through 3 Faster_False blocks, and then the output tensor of the second CBS module is subjected to Concat splicing and then the third CBS module, so as to obtain an output tensor with the same dimension as the original input tensor.

[0015] Step three, using the WIoU_v2 Loss bounding box regression loss function;

[0016] Further, the WIoU_v2 Loss bounding box regression loss function is as follows:

[0017]

[0018] L WIoUv1 =R WIoU LIoU ;

[0019]

[0020]

[0021] wherein, B gt is the real box, B is the anchor box; L IoU is the predicted bounding box loss, R WIoU is the attention mechanism taking the center point normalized distance as the measurement;(x,y), (x gt ,y gt ) are the center point coordinates of the predicted box and the real box respectively, (W g , H g ) are the width and height of the maximum bounding box of the predicted box and the real box, L WIoUv2 is the bounding box regression loss function of WIoU_v2Loss, is the sliding average value of L IoU ; gamma is an adjustable factor; * indicates that the parameter is separated in the calculation.

[0022] Further, the lightweight electric vehicle rider helmet wearing detection system is characterized in that it comprises: a memory for storing instructions executable by a processor; and a processor for executing instructions to implement a lightweight electric vehicle rider helmet wearing detection method.

[0023] Further, the computer readable medium storing computer program code is characterized in that the computer program code, when executed by a processor, implements a lightweight electric vehicle rider helmet wearing detection method.

[0024] The present application has the following beneficial effects:

[0025] 1. The C3 module of the backbone layer is replaced by the lightweight Faster_x module, and the C3 model of the neck layer is replaced by the Faster_false_x module, thereby reducing network redundancy and memory loss while maintaining network feature extraction capability.

[0026] 2. The lightweight Involution convolution module is introduced, which reduces the redundancy of the convolution kernel, improves the ability to extract spatial semantic features, reduces the parameter quantity and operation quantity of the network, and improves the detection accuracy.

[0027] 3. The bounding box Loss function of the YOLOv5n network is optimized, a WIoU_v2 Loss bounding box regression loss function with monotonic focusing mechanism and attention mechanism is constructed, the attention of the network model to low-quality data is improved, and the punishment to high-quality data is reduced, the generalization ability of the model to the data set is improved, and the model is more conducive to detection in a complex background street;

[0028] 4. Compared with the standard YOLOv5n network, the improved network reduces the demand for device memory and processor required for deployment, has better edge deployment capability, and can better adapt to image data detection tasks in a generalized environment. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 is a flow chart of the light-weight electric vehicle rider helmet wearing detection method and system of the present application;

[0030] Figure 2 is a schematic diagram of the improved YOLOv5n network model of the present application;

[0031] Figure 3 is an image example of three ways to obtain sample data;

[0032] Figure 4 is a schematic diagram of the improved YOLOv5n network sub-module structure of the present application;

[0033] Figure 5 is a schematic diagram of the Involution convolution module;

[0034] Figure 6 is a comparison diagram of the inference speed of the improved YOLOv5n model and the standard YOLOv5n model;

[0035] Figure 7 is an example diagram of the optimized model for rider helmet wearing detection. DETAILED DESCRIPTION

[0036] The present application will be further described below in conjunction with the drawings and examples, which are simplified schematic diagrams and only illustrate the basic structure of the present application in a schematic manner, and therefore only show the components related to the present application.

[0037] As shown in Figure 1 , the light-weight electric vehicle rider helmet wearing detection method comprises the following steps:

[0038] Step 1, collect electric vehicle rider image data, construct electric vehicle rider helmet wearing picture data set, and pre-process and label the image;

[0039] The constructed dataset of images of electric scooter riders wearing helmets was obtained by filtering images from some publicly available Kaggle scooter rider datasets, images captured by web crawlers, and images obtained from real street photography. Examples of some of the image data are shown below. Figure 3 As shown; where, Figure 3 (a) is an example of an image selected from a public dataset. Figure 3 (b) is an example of an image captured on the internet. Figure 3 (c) is an example of street photography image data.

[0040] The logical processing of the dataset involves removing low-quality image data that is blurry, has ghosting, or otherwise does not conform to the meaning of real-world detection, and applying three data enhancement methods—HSV color interference, Gaussian noise, and salt-and-pepper noise—to the remaining image data, ultimately obtaining 5728 image data.

[0041] Then, using labelimg in PascalVOC format, rectangular boxes were used to annotate the cyclists' heads in the image dataset. The labels were divided into two categories: With Helmet and Without Helmet. The annotated dataset was then divided into a training set of 3806 examples and a validation set of 1922 examples.

[0042] Step 2: Construct an improved YOLOv5n network and train it on the training set to obtain the optimal weights;

[0043] like Figure 2 As shown, the improved YOLOv5n network extracts semantic information through a backbone layer, which includes four CBS convolutional modules, two Faster_3 lightweight feature extraction modules, two Faster_6 lightweight feature extraction modules, one Involution inner convolutional module, and one SPPF spatial pyramid pooling module. Semantic information is fused through a neck layer, which includes two CBS convolutional modules, two upsampling modules, four Faster_False_3 modules, and two Involution inner convolutional modules. Finally, the network passes through a Conv2d convolutional layer to generate feature maps of three different scales for prediction.

[0044] The structure of each submodule is as follows: Figure 4 As shown, the CBS convolutional module consists of a Conv convolutional layer, a batch normalization (BN) layer, and a Silu activation function.

[0045] The Faster_x module is lightened, which is composed of a CBS module and x Faster blocks, and x is 3 and 6 in the embodiment of the application, that is, Faster_3 and Faster_6 sub-modules of the backbone layer are obtained; the input tensor is halved in the channel dimension by two CBS modules in the Faster_x module, the output tensor of the first CBS module is extracted by x Faster blocks, and then the output tensor is concatenated with the output tensor of the second CBS module, and then the output tensor is obtained after passing through the third CBS module, and the output tensor has the same dimension as the original input tensor; the Faster block is composed of a lightened PConv convolution, a CBS module and a Conv2d convolution; the input tensor is first passed through a PConv module, then passed through a CBS convolution layer, and then passed through a Conv2d convolution, and the obtained tensor is added to the original input tensor to obtain the output tensor.

[0046] The Faster_False_x module of the neck layer, x is 3 in the embodiment of the application, that is, the Faster_False_3 sub-module is obtained, and Faster_False_x adopts the Faster_False block without a residual structure; the input tensor of the Faster_False block is passed through a Pconv convolution, a CBS convolution module and a Conv2d convolution to obtain the output tensor; the Pconv convolution is different from the ordinary Conv in that the PConv convolution divides the input tensor into C p and C-C p two parts according to the channel dimension, only the channel C p part is extracted by the conventional Conv, and the remaining channel C-C p part remains unchanged. The Faster_False_x module is similar in structure to the Faster_x module, and the Faster block is replaced by the Faster_False block; wherein the x modules are in a series structure.

[0047] The first CBS convolution layer of the SPPF module in the backbone part and the two down-sampling CBS convolution layers in the neck are replaced by the lightened Involution inner convolution module.

[0048] The Involution inner convolution module improves the extraction degree of the convolution layer to different spatial information features by assigning different inner convolution kernels to different spatial positions. Figure 5 As shown in the figure, the Involution inner convolution module includes inner kernel generation and inner convolution calculation:

[0049] Hi,j = φ(X i,j ) = W1δ(W0X i,j ) ;

[0050]

[0051] X i ' ,j = ψ(N(X i,j ,k), H i,j ) ;

[0052] where C denotes the number of channels of the input feature map, φ() denotes an inner kernel generation function, X i,j denotes a data block of size (1*1*C) at the spatial position (i,j) of the input feature map, W0 and W1 denote transformation matrices, r denotes a channel reduction ratio, (k*k*1) denotes the number of output channels, δ denotes a batch normalization operation and a linear activation operation, H i,j is a corresponding inner kernel generated at the spatial position (i,j) of the input feature map, and has a size of (k*k*1); ψ() denotes an inner convolution calculation function, and N(X i,j ,k) denotes a k*k data block centered at X i,j in the input feature map, which has a size of (k*k*C). The inner convolution calculation function performs a convolution operation on the N(X i,j ,k) data block and the corresponding generated inner kernel H i,j , and then performs summation in the channel dimension to generate a target feature X' i,j of size (1*1*C) at the spatial position (i,j).

[0053] The involution inner convolution module first defines the shape of the inner kernel as (k*k*1), and in this example, k=3; then, through the inner kernel generation function φ(), the feature data block X i,j of size (1*1*C) at the spatial position (i,j) of the input feature map is compressed in the channel dimension through the W0 transformation matrix. In this example, a channel compression ratio of r=4 is selected, and the size of the transformed feature data block is After the data block compressed in the channel dimension is normalized and activated through the δ function, the channel is expanded through W1, and the size of the expanded data block is (1*1*k 2 ). The data block is rearranged to have a shape of (k*k*1) to obtain the inner kernel H i,j corresponding to the position (i,j).

[0054] The inner convolution calculation function ψ() performs a convolution operation on the data block N(X i,j ,k) at X i,j and the inner kernel H i,jWhen performing convolution operations, it is worth noting the data block N(X) i,j The C-layer channels of (k) use the same involution kernel H for convolution operations. i,j The data blocks obtained after the convolution operation are then summed according to the channel dimension to obtain the target location feature X' of size (1*1*C). i,j .

[0055] Step 3: Use the WIoU_v2 Loss bounding box regression loss function;

[0056] The WIoU_v2 Loss bounding box regression loss function is as follows:

[0057]

[0058] L WIoUv1 =R WIoU L IoU ;

[0059]

[0060]

[0061] Among them, B gt B is the true bounding box, and L is the anchor bounding box; IoU To predict the bounding box loss, R WIoU This is an attention mechanism that uses the normalized distance from the center point as a metric; (x,y), (x gt ,y gt ) are the center coordinates of the predicted bounding box and the ground truth bounding box, respectively. g H g L represents the width and height of the maximum bounding box between the predicted bounding box and the ground truth bounding box. WIoUv2 WIoU_v2Loss is the bounding box regression loss function. For L IoU The moving average; γ is an adjustable factor, in this example γ = 0.5; * indicates that the parameter is extracted in the calculation and treated as a constant in backpropagation.

[0062] The dataset was trained on the improved YOLOv5n network with a batch size of 32 and a total of 200 iterations. The training platform parameters are shown in Table 1.

[0063] Table 1 Experimental Hardware and Software Platform

[0064]

[0065] The weights of the improved YOLOv5n network are obtained through training and validated using a validation set with a batch size of 1. This yields the time required for the model to detect one frame of data on this platform. After ten iterations, the inference time of the improved YOLOv5n network is obtained and compared with that of the standard YOLOv5n network. Figure 6 As shown, the inference time of this example is significantly better than that of the standard YOLOv5n model.

[0066] The detection accuracy and edge detection capabilities of the improved algorithm were evaluated. The size of the parameters and the number of operator points determine the memory channel requirements of the network model deployment and the pressure on the processor to process image parameters. Detection accuracy determines the edge detection capability and practicality of this method. The dataset was trained on the standard YOLOv5n, and the results were compared in the above three aspects. The mean accuracy at IoU = 0.5 (map@0.5) was selected as the detection accuracy. The comparison results are shown in Table 2.

[0067] Table 2 Comparative Experimental Data

[0068]

[0069] Compared to the standard YOLOv5n, the network model of this invention reduces the number of parameters by 53.0% and the number of floating-point operations by 36.6%, significantly reducing the pressure on the memory and processor requirements of edge deployment devices; some detection examples are as follows. Figure 7 As shown, the performance requirements are met for tasks such as multi-target detection and high occlusion detection.

[0070] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A lightweight method for detecting helmet wearing by electric vehicle riders, characterized in that, Includes the following steps: Step 1: Collect image data of electric vehicle riders, construct an image dataset of riders wearing helmets, and preprocess and label the images; Step 2: Construct an improved YOLOv5n network and train it on the training set to obtain the optimal weights; The improved YOLOv5n network includes: semantic information extraction via the backbone layer, which comprises 4 CBS convolutional modules, 2 Faster_3 modules, 2 Faster_6 modules, 1 Involution inner convolutional module, and 1 SPPF spatial pyramid pooling module; semantic fusion via the neck layer, which comprises 2 CBS convolutional modules, 2 upsampling modules, 4 Faster_False_3 modules, and 2 Involution inner convolutional modules; and finally, feature maps of three different scales are generated via Conv2d convolutional layers. The Faster_3 module consists of a CBS module and three Faster blocks. The input tensor is halved in channel dimension by passing through two CBS modules. The output tensor of the first CBS module is processed by the three Faster blocks for feature extraction. After concatenation with the output tensor of the second CBS module, it is passed through the third CBS module to obtain an output tensor with the same dimension as the original input tensor. The Faster_6 module consists of a CBS module and 6 Faster blocks. The input tensor is halved in channel dimension by passing through two CBS modules. The output tensor of the first CBS module is processed by the 6 Faster blocks for feature extraction. After being concatenated with the output tensor of the second CBS module, it is passed through a third CBS module to obtain an output tensor with the same dimension as the original input tensor. The Faster_False_3 module consists of a CBS module and three Faster_False blocks. The input tensor is halved in channel dimension by passing through two CBS modules. The output tensor of the first CBS module is processed by the three Faster_False blocks for feature extraction. After being concatenated with the output tensor of the second CBS module, it is then passed through the third CBS module to obtain an output tensor with the same dimension as the original input tensor. Step 3: Use the WIoU_v2 Loss bounding box regression loss function.

2. The lightweight electric vehicle rider helmet wearing detection method according to claim 1, characterized in that, Image preprocessing includes: removing blurry and ghosted images, and performing HSV color interference, Gaussian noise, and salt-and-pepper noise data enhancement on the remaining image data.

3. The lightweight electric vehicle rider helmet wearing detection method according to claim 1, characterized in that, The WIoU_v2 Loss bounding box regression loss function is as follows: ; ; ; ; in, For the true frame, For anchor frame; To predict the bounding box loss, An attention mechanism that uses the normalized distance from the center point as a metric; x , y ), ( , ) are the center coordinates of the predicted bounding box and the ground truth bounding box, respectively. , ( ) represents the width and height of the largest bounding box between the predicted bounding box and the ground truth bounding box. The WIoU_v2 Loss bounding box regression loss function is used. for The moving average; It is an adjustable factor; This indicates that the parameters are extracted from the calculation.

4. A lightweight electric vehicle rider helmet wearing detection system, characterized in that, include: Memory is used to store instructions that can be executed by the processor; A processor for executing instructions to implement the lightweight electric vehicle rider helmet wearing detection method as described in any one of claims 1-3.

5. A computer-readable medium storing computer program code, characterized in that, The computer program code, when executed by a processor, implements the lightweight electric vehicle rider helmet wearing detection method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Method for detecting helmet wearing of rider based on YOLOv5s

    CN116469062A

Cited By

  • A Cyclist Helmet Detection Method Based on VSC-YOLO

    CN122368929A