Vehicle detection algorithm and device based on Faster Net lightweight framework and feature fusion

By using the FasterNet lightweight framework and feature fusion method in the vehicle detection algorithm, the existing vehicle detection algorithm has solved the problem of high computational complexity and poor real-time performance in traffic road scenarios, and achieved higher detection accuracy and real-time performance.

CN120107543AActive Publication Date: 2025-06-06NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510004683.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-06-06
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

The existing vehicle detection algorithm based on deep learning has high computational complexity, poor real-time performance, high resource consumption in traffic road scenarios, and has low accuracy in multi-objective and small-objective detection in complex scenarios, making it difficult to meet the needs of intelligent transportation and autonomous driving.

Method used

The vehicle detection algorithm that integrates FasterNet lightweight framework and features is adopted, and the CSPDarkNet53 backbone network of YOLOv8 is replaced as a lightweight FasterNet network, and the F-C2f module is introduced into the neck network for multi-scale feature fusion, combining the Focal-EIoU loss function to improve detection accuracy and real-time.

Benefits of technology

The complexity and calculation cost of the model are reduced, the real-time and accuracy of vehicle detection are improved, especially in complex traffic scenarios, the detection effect of multiple and small targets is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107543A_ABST
    Figure CN120107543A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle detection algorithm and device based on a Faster Net lightweight framework and feature fusion, and relates to the field of deep learning and image processing, the algorithm comprises the steps: inputting an image, and carrying out the preprocessing of the image, the preprocessing mainly comprising LetterBox and normalization operation; inputting the preprocessed image into a backbone network based on a FasterNet network and a multi-head attention mechanism for feature extraction to obtain a preliminary extraction feature map; inputting the preliminarily extracted feature map into a neck network based on an F-C2f module for multi-scale feature fusion to obtain a mixed fusion feature map; inputting the mixed fusion feature map into a head network of a vehicle detection model, training by adopting a Focal-EIoU loss function, and stopping training when a training period reaches 100 to obtain a trained vehicle detection model; and inputting the test image into the trained vehicle detection model to obtain a vehicle detection result of the test image. The vehicle detection model provided by the algorithm can give consideration to the accuracy and real-time performance of vehicle detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of deep learning and image processing, and in particular to a vehicle detection algorithm and device based on the FasterNet lightweight framework and feature fusion. Background Art

[0002] With the rapid development of intelligent transportation systems and autonomous driving technologies, vehicle detection, as an important branch of target detection, is of great significance in improving traffic safety and optimizing traffic management.

[0003] Vehicle detection is to locate the vehicle target of interest in the image, give the category of the located vehicle target according to the feature information, and circle the located vehicle target with the detected border to complete the detection task. In recent years, with the vigorous development of Internet of Things technology and computer vision technology, vehicle detection algorithms based on deep learning have been favored by many researchers. Common vehicle detection algorithms based on deep learning include two-stage detection algorithms represented by the R-CNN (Regions with Convolutional Neural Network) series of networks, and single-stage detection algorithms represented by the SSD (Single Shot Multibox Detector) and YOLO (You Only Look Once) series. Among them, the detection result accuracy of the two-stage detection algorithm is higher, but the running speed is slower; while the single-stage detection algorithm improves the detection speed, but sacrifices a certain degree of accuracy. Therefore, the traditional vehicle detection algorithm cannot take into account the requirements of real-time and accuracy of vehicle detection in traffic road scenes. Summary of the invention

[0004] This application provides a vehicle detection algorithm and device based on the FasterNet lightweight framework and feature fusion. The technical solution is as follows:

[0005] According to one aspect of the present application, a vehicle detection algorithm based on the FasterNet lightweight framework and feature fusion is provided, the algorithm comprising:

[0006] Input an image and perform preprocessing on the image, wherein the preprocessing mainly includes LetterBox and normalization operations;

[0007] The preprocessed image is input into the backbone network of the vehicle detection model for feature extraction to obtain a preliminary extracted feature map, wherein the feature extraction network in the backbone network is a feature extraction network based on a lightweight FasterNet network and a multi-head attention mechanism;

[0008] Inputting the initially extracted feature map into the neck network for multi-scale feature fusion to obtain a mixed fusion feature map, wherein the neck network is a neck network based on the F-C2f module;

[0009] Input the hybrid fusion feature map into the head network of the vehicle detection model, use the Focal-EIoU loss function for training, and stop training when the training cycle reaches 100 to obtain a trained vehicle detection model;

[0010] The test image is input into the trained vehicle detection model to obtain the vehicle detection result of the test image.

[0011] According to another aspect of the present application, a vehicle detection device based on a FasterNet lightweight framework and feature fusion is provided, the device comprising:

[0012] A preprocessing module, used for inputting an image and performing preprocessing on the image, wherein the preprocessing mainly includes LetterBox and normalization operations;

[0013] A feature extraction module is used to input the preprocessed image into the backbone network of the vehicle detection model for feature extraction to obtain a preliminary extracted feature map. The feature extraction network in the backbone network is a feature extraction network based on a lightweight FasterNet network and a multi-head attention mechanism;

[0014] A feature fusion module, used for inputting the initially extracted feature map into a neck network based on an F-C2f module for multi-scale feature fusion to obtain a mixed fusion feature map, wherein the neck network is a neck network based on an F-C2f module;

[0015] A model training module is used to input the hybrid fusion feature map into the head network of the vehicle detection model, and use the Focal-EIoU loss function for training. When the training cycle reaches the set value, the model stops training. In the process of training the vehicle detection model in this application, a total of 100 training cycles (epochs) are set to ensure that the vehicle detection model is fully converged;

[0016] The model prediction module is used to input the test image into the trained vehicle detection model to obtain the vehicle detection result of the test image.

[0017] According to one aspect of the present application, an electronic device is provided, comprising: a processor and a memory storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the vehicle detection algorithm and device based on the FasterNet lightweight framework and feature fusion as described above.

[0018] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the vehicle detection algorithm and device based on the FasterNet lightweight framework and feature fusion as described above.

[0019] According to another aspect of the present application, a computer program product is provided, the computer program product comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned vehicle detection algorithm and apparatus based on the FasterNet lightweight framework and feature fusion.

[0020] The beneficial effects brought by the technical solution provided by the embodiment of the present application include at least:

[0021] Aiming at the problems of high computational complexity, poor real-time performance and high resource consumption of the current deep learning-based vehicle detection algorithm in traffic road scenes, as well as low detection accuracy of multiple targets and small targets in complex scenes, which makes it difficult to meet the requirements of vehicle detection in practical application scenarios such as intelligent transportation and autonomous driving, this application proposes a vehicle detection algorithm with FasterNet lightweight framework and feature fusion. The model uses FasterNet lightweight network to replace the CSPDarkNet53 backbone network of the original YOLOv8, and combines the idea of ​​partial convolution to reconstruct the C2f module in the neck of YOLOv8, realizing efficient fusion of multi-scale features and reducing the complexity and computational cost of the model. The Focal-EIoU loss function is proposed to replace the CIoU loss function in YOLOv8, accurately define the aspect ratio of the prediction box, alleviate the imbalance problem of positive and negative samples, and improve the detection accuracy of multiple targets and small targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Further details, features and advantages of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0023] Figure 1 A flow chart of a vehicle detection algorithm based on the FasterNet lightweight framework and feature fusion according to an exemplary embodiment of the present application is shown;

[0024] Figure 2 This is the structure diagram of the FasterNet network;

[0025] Figure 3 It is the convolution structure diagram of the PConv module;

[0026] Figure 4 It is the structural diagram of F-C2f module;

[0027] Figure 5 It is the structural diagram of the multi-head self-attention mechanism;

[0028] Figure 6 It is a parameter diagram of EIoU loss;

[0029] Figure 7 is a network structure diagram of an improved YOLOv8 model provided by an exemplary embodiment of the present application;

[0030] Figure 8 These are the vehicle detection experimental results of four methods;

[0031] Fig. 9 It is a structural schematic diagram of a vehicle detection device based on the FasterNet lightweight framework and feature fusion provided in an embodiment of the present application;

[0032] Fig.10 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present application is shown. DETAILED DESCRIPTION

[0033] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not intended to limit the scope of protection of the present application.

[0034] It should be understood that the various steps described in the algorithm implementation of the present application can be performed in different orders and / or in parallel. In addition, the algorithm implementation may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0035] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more". The names of the messages or information exchanged between multiple devices in the embodiments of this application are only for illustrative purposes, and are not used to limit the scope of these messages or information.

[0036] The following describes the solution of the present invention with reference to the accompanying drawings, and illustrates the technical solution provided by the embodiments of the present application in detail through specific embodiments and their application scenarios.

[0037] At present, vehicle detection algorithms based on deep learning include two-stage algorithms and single-stage algorithms. The two-stage vehicle detection algorithm inherits the sliding window idea of ​​the traditional detection algorithm. The entire algorithm is divided into two parts: candidate area selection and classification. The advantage of this algorithm is that the selection of multiple candidate boxes can fully extract the feature information of the vehicle target, the detection result has a high accuracy, and accurate positioning can be achieved at the same time. However, since it is divided into two steps, the running speed is slow and the real-time performance is low. It is suitable for scenes with high accuracy requirements. The single-stage vehicle detection algorithm omits the candidate area selection step, simplifies the network structure, and greatly improves the detection speed of the algorithm, but at the same time leads to insufficient feature extraction, sacrifices a certain degree of accuracy, and reduces the accuracy of the detection target positioning. It is suitable for scenes with high real-time requirements.

[0038] Therefore, in order to address the problems of high computational complexity, poor real-time performance, and high resource consumption of the current deep learning-based vehicle detection algorithm in traffic road scenarios, as well as low accuracy in detecting multiple targets and small targets in complex scenarios, which makes it difficult to meet the requirements of vehicle detection in practical application scenarios such as intelligent transportation and autonomous driving, this application proposes a vehicle detection algorithm and device with a FasterNet lightweight framework and feature fusion, which improves the real-time performance and accuracy of vehicle detection in complex scenarios such as urban roads, highways, and severe weather, and reduces resource consumption.

[0039] Please refer to Figure 1, which shows a flow chart of a vehicle detection algorithm based on the FasterNet lightweight framework and feature fusion according to an exemplary embodiment of the present application. The algorithm is applied to electronic equipment as an example for explanation. Figure 1 As shown, the algorithm includes:

[0040] Step 101: input an image and perform preprocessing on the image, wherein the preprocessing mainly includes LetterBox and normalization operations.

[0041] The vehicle detection algorithm based on deep learning essentially requires pre-training of the vehicle detection model. After the training completion conditions are met, the trained vehicle detection model can be used to detect vehicles on the input image. Before training the vehicle detection model, it is first necessary to select a data set as the model training data. In the vehicle detection application scenario, this application selects the UA-DETRAC data set as the model training data. The images in the data set are vehicle images of various road types (such as highways, urban main roads, secondary roads) and various weather conditions (such as normal weather and bad weather) in the city taken by a Cannon EOS 550D camera at 24 different locations; moreover, the vehicle images in the data set have three different degrees of occlusion, namely, completely visible, partially blocked by other vehicles, and partially blocked by the scene; the vehicle types contained in the images in the data set include the four categories of "car", "bus", "van", and "other (other types of vehicles)" commonly seen in road traffic. Each vehicle type includes a variety of vehicles from small to large, from private to public, and the data set is diverse. Since there are a large number of images in the dataset and the differences between adjacent images are very small, this application randomly extracts the UA-DETRAC dataset and selects one image out of every ten images as the final dataset, resulting in a training set of 8232 images and a test set of 5625 images.

[0042] Optionally, in order to increase the diversity of the data set, the images in the training set can also be processed by the Mosaic data enhancement technology. Specifically, four images are randomly selected from the training set, and different random scaling ratios between 0.5 and 2.0 are applied to each image to enhance the model's detection ability for vehicles of different sizes. According to the 640*640 image size required for the vehicle detection model input, a random cropping operation is performed on the scaled image to ensure that the cropped area is close to or equal to the required size and contains key target information (e.g., vehicle). If the cropped image block size is not large enough to fill the preset canvas (such as a 2*2 grid layout), it is adjusted by re-cropping, adjusting the scaling ratio, or changing the cropping strategy to ensure that the image blocks can be seamlessly combined. The four cropped and adjusted image blocks are arranged in a random order on the new canvas. If the cropped image block size is not large enough to fill the entire canvas, black (grayscale value 0) or white (grayscale value 255) is used as the fill color to ensure the consistency of the Mosaic image size. Visually, these two colors will not cause a strong conflict with most natural scenes, and will not introduce additional color information to interfere with model training.

[0043] Mosaic technology increases the diversity of the training data set by stitching four different images together, which not only improves the generalization ability of the model, but also can include small targets that are originally difficult to detect in the process of random scaling and cropping, thereby enhancing the model's detection effect on small targets. In addition, Mosaic data enhancement technology effectively reduces the risk of overfitting the model to the training data by increasing the amount of training data.

[0044] Moreover, considering that the vehicle detection model has a fixed-size image input requirement, illustratively, the target size of the vehicle detection model input of the present application is 640*640. The size of the original image selected from the UA-DETRAC dataset may not meet this requirement, so the original image needs to be preprocessed. The preprocessing method mainly includes LetterBox and normalization operations. The training efficiency and model performance of the model are improved through image preprocessing. Among them, LetterBox is an image preprocessing method, which is mainly used to adjust the size of the image while maintaining the original aspect ratio of the image. This method usually involves scaling the image to the longest side of the target size, while padding the short sides to ensure that the entire image meets the size requirements required by the model.

[0045] Specifically, the process of performing LetterBox processing on the original image may include the following steps:

[0046] (1) Calculate the scaling ratio based on the size of the original image and the target size (the target size is the image size required for the vehicle detection model input). Exemplarily, the scaling ratio calculation formula is shown in formula (1):

[0047]

[0048] It is assumed that the size of the original image is (H, W), the target size is (new_H, new_W), and r represents the scaling ratio.

[0049] (2) Calculate the height and width of the scaled image based on the scaling ratio and the size of the original image. Exemplarily, the calculation formulas for the height and width of the scaled image are shown in formula (2) and formula (3):

[0050] new_unpad_H=round(H*r) (2)

[0051] new_unpad_W=round(W*r) (3)

[0052] Among them, new_unpad_H represents the height of the scaled image, H represents the height of the original image, r represents the scaling ratio, new_unpad_W represents the width of the scaled image, and W represents the width of the original image.

[0053] (3) Based on the scaled image height, the scaled image width and the target size, the difference between the target size and the scaled image size is calculated to obtain the fill height and fill width that still need to be filled. Exemplarily, the calculation formulas for the fill height and fill width are shown in formula (4) and formula (5).

[0054] dh=new_H-new_unpad_H (4)

[0055] dw=new_W-new_unpad_W (5)

[0056] Among them, dh represents the padding height that needs to be filled on the image after the original image is scaled, and dw represents the padding width that needs to be filled on the image after the original image is scaled.

[0057] (4) Based on the padding height and padding width, determine the top padding size, the bottom padding size, the left padding size, and the right padding size. Exemplarily, the calculation formulas for the padding size are shown in formula (6), formula (7), formula (8), and formula (9), respectively.

[0058]

[0059] bottom=dh-top (7)

[0060]

[0061] right=dw-left (9)

[0062] Among them, top represents the top padding size, bottom represents the bottom padding size, left represents the left padding size, and right represents the right padding size.

[0063] Based on formula (1) to formula (9), the process of scaling the original image to the target size is as follows: first, the scaling ratio is calculated, and the original image is scaled to the size (new_unpad_H, new_unpad_W) by using the image scaling algorithm combined with the scaling ratio; then, the area around the scaled image is padded with a padding size of (top, bottom, left, right), thereby achieving the purpose of adjusting the original image to the target size while maintaining the original aspect ratio of the image. Using LetterBox for image preprocessing reduces information loss and image deformation caused by resizing, which is crucial for maintaining the authenticity of the image content; and LetterBox reduces the negative effects of excessive padding, such as reduced model training efficiency, by minimizing the padding area.

[0064] Based on the above description, in a possible implementation, after obtaining the original image from the training data set, the original image is subjected to image preprocessing, namely, LetterBox and normalization processing, to obtain the preprocessed image required for the final training of the vehicle detection model.

[0065] Step 102, input the preprocessed image into the backbone network of the vehicle detection model for feature extraction to obtain a preliminary extracted feature map, wherein the feature extraction network in the backbone network is a feature extraction network based on a lightweight FasterNet network and a multi-head attention mechanism.

[0066] Among the existing vehicle detection algorithms, the YOLOv8 model has a good detection effect. The network structure of the YOLOv8 model mainly consists of three parts: the backbone network (Backbone), the neck network (Neck) and the head network (Head). The backbone network starts with the input image and gradually extracts the high-level features of the image through a series of convolutional layers and C2f modules. The backbone network of YOLOv8 refers to the structure of CSPDarkNet-53, but has made many improvements. The most notable one is the introduction of the C2f (Cross-convolution with 2filters) structure, which replaces the C3 structure in YOLOv5. The C2f structure enhances the model performance by optimizing the gradient flow. The neck network adopts a PAN-FPN structure similar to YOLOv5, namely PathAggregation Network (PANet), which realizes cross-scale transmission of information and improves the multi-scale detection capability of the model. The design of PANet makes the YOLOv8 model perform well in object detection of different scales, especially in small object detection. The output feature map of the backbone network is subjected to multi-scale feature extraction through the SPP (Spatial Pyramid Pooling) structure, and then feature fusion is performed through PANet. This structure enables the YOLOv8 model to detect targets of different scales more effectively. The head network of the YOLOv8 model uses a Decoupled Head structure similar to YOLOX, which is responsible for converting the fused feature map into the final detection result. The head network separates the regression branch and the classification branch, which helps to improve the convergence speed and detection effect of the model.

[0067] In view of the high computational complexity, poor real-time performance and high resource consumption of the vehicle detection algorithm based on deep learning in the current traffic road scene, this application has made some improvements to the original YOLOv8 model, so as to improve the detection speed and timeliness while ensuring the detection accuracy; firstly, the lightweight FasterNet network is used to replace the YOLOv8 backbone network CSPDarkNet53, and the MHSA layer is added at the end of the backbone network to fully extract vehicle features. The FasterNet network successfully achieves a balance between speed and accuracy by introducing the innovative technology of partial convolution (PConv). Figure 2The structural diagram of the FasterNet network is given. The FasterNet network mainly has 4 layers, among which the embedding layer and the fusion layer are composed of 4*4 convolution and 2*2 convolution respectively, which have the functions of spatial downsampling and expanding the number of channels. FasterNetBlock is the main feature extraction module. Each FasterNet Block has a PConv layer (3*3), followed by two point-wise convolution (PWConv) layers (1*1). Batch normalization and ReLU activation function are used between these two 1*1 convolution layers to accelerate the training process and enhance the generalization ability of the model. Finally, residual connection is performed.

[0068] Among them, the PConv module uses the Conv operation to extract features from some channels of the feature map with an input dimension of H*W*C, while other channels remain unchanged. For the continuity of memory access, the first or last continuous C p When C p When the number of floating-point operations per second (FLOPS) of PConv is 1 / 4, compared with conventional convolution, the number of floating-point operations per second (FLOPS) of PConv is only 1 / 16 of that of conventional convolution. The processed channels are concatenated with the unprocessed channels, and the output is identically mapped to the original feature map dimension, so it has the same number of channels. PConv effectively solves the problem of model overfitting and excessive computation caused by feature redundancy by reducing redundant calculations and storage accesses. Figure 3 The convolution structure diagram of the PConv module is given. For a feature map with an input dimension of H*W*C, only the first C p The convolution operation is performed on consecutive channels, and the convolution kernel used in the convolution operation is k*k*C p , get the feature map after convolution, and keep the same CC p The channels are concatenated to obtain the feature map after convolution processing by the PConv module.

[0069] The FasterNet feature extraction network built on PConv can not only effectively capture and refine key feature information in the spatial dimension, but also make full use of the computing power of the hardware equipment to achieve a significant increase in detection speed. This optimization strategy not only improves the model's ability to process large data sets, but also shortens the processing time, providing strong technical support for vehicle detection in real-time application scenarios. Therefore, based on the original YOLOv8 model, the use of a lightweight FasterNet network to replace the CSPDarkNet53 in the YOLOv8 backbone network can enable the improved vehicle detection model to have improved detection real-time performance.

[0070] Compared with the traditional self-attention mechanism, MHSA extends the standard self-attention process by introducing the concept of multiple "heads". MHSA focuses on different positions in the input data and can better capture long-distance dependencies and structures in the data. Figure 5 The structural diagram of the multi-head self-attention mechanism is given. The multi-head self-attention mechanism divides the input sequence into multiple independent attention heads by linear transformation, and each attention head calculates the attention weight separately. These attention heads independently learn different relations and feature representations, and finally integrate the outputs of these attention heads by connection (concatenation) or weighted summation to form the final multi-head self-attention output, thereby obtaining a richer representation. Single dot product attention mechanism, additive attention mechanism, and self-attention mechanism cannot capture all information from multiple angles and granularity. Compared with the three single-head attention mechanisms of dot product attention mechanism, additive attention mechanism, and self-attention mechanism, in the multi-head self-attention mechanism, different attention heads can learn different relations and patterns, so as to better capture diverse information. Multiple attention heads reduce the overall number of parameters and reduce the risk of overfitting by sharing parameters.

[0071] The number of heads of the multi-head self-attention mechanism used in the original Transformer model is 8, while in some tasks where large models require fine-grained features, such as complex semantic understanding in natural language processing or fine feature capture in image recognition, the number of heads of the multi-head self-attention mechanism is generally 12. In resource-constrained environments, or in application scenarios where the model is lightweight and low-latency, in order to improve model performance without significantly increasing the computational burden, the number of attention heads is usually reduced to 4. This strategy not only helps to reduce the number of model parameters and memory usage, but also significantly improves the operating efficiency of the model and meets real-time requirements. Therefore, the multi-head self-attention mechanism layer in this application has 4 attention heads to ensure higher resource utilization and faster response speed while maintaining model performance.

[0072] This application adds a multi-head self-attention mechanism layer to the last layer of the backbone network to perform global information integration on the features extracted by the backbone network. This application uses 4 attention heads, each of which is an independent self-attention mechanism that needs to learn a set of weights. Multiple self-attention heads are connected to form an output as the calculation result of the multi-head self-attention mechanism, thereby improving the expression ability of the model. The calculation formulas are shown in formulas (10), (11), (12), (13) and (14).

[0073] Q=[q 1 ,q 2 ...q h ]=WQ ·[x 1 ,x 2 ...x h ] (10)

[0074] K=[k 1 ,k 2 ...k h ]=W K ·[x 1 ,x 2 ...x h ] (11)

[0075] V=[v 1 ,v 2 ...v h ]=W V ·[x 1 ,x 2 ...x h ] (12)

[0076]

[0077] M h (Q,K,V)=C(h 1 ,h 2 ...h n )W O (14)

[0078] Among them, Q, K, V, are the input query vector, key vector and value vector respectively, h 1 ,h 2 …,h n is the output of each self-attention layer, is the attention weight, S is the Softmax function, is the dot product scaling factor, h is the number of heads, and W O , W K , W V is a linear transformation, and C represents the concat function connecting multiple attention heads.

[0079] Step 103, inputting the initially extracted feature map into the neck network based on the F-C2f module for multi-scale feature fusion to obtain a mixed fusion feature map, wherein the neck network is a neck network based on the F-C2f module.

[0080] In the neck network of the original YOLOv8 model, the C2f module is a complex structure composed of a standard convolution Conv and multiple Bottelneck blocks. The multiple Bottelneck blocks perform a large number of skip-layer connections and Split operations, resulting in a complex network structure and a large amount of calculation, which cannot meet the real-time requirements of vehicle detection. Therefore, in addition to improving the backbone network of the original YOLOv8 model, the C2f module in the neck network is also improved, deleting the Split operation and redundant skip-layer connections in the original C2f module, reducing a large number of 3*3 ordinary convolutions in C2f, and replacing all Bottlenecks in the original network C2f module with Fasterblock, forming a new F-C2f module. While reducing the number of model parameters, the F-C2f module enables the model to fully and effectively utilize the information of all channels and maintain the diversity of features extracted by the model. Figure 4 The structural diagram of the F-C2f module is given. The F-C2f module consists of CBS, Split operation, connection layer and F-Bottleneck, which is mainly composed of convolutional layer and Fasterblock layer in FasterNet network.

[0081] The embodiment of the present application utilizes the idea of ​​partial convolution to improve the C2f module, specifically by deleting the Split operation and redundant skip-layer connections in the original C2f structure, reducing a large number of 3*3 ordinary convolutions in C2f, and replacing all Bottleneck layers in the original network C2f module with Fasterblock to form a new F-C2f module, thereby obtaining the vehicle detection model provided by the present application.

[0082] Step 104, input the hybrid fusion feature map into the head network of the vehicle detection model, and use the Focal-EIoU loss function for training. When the training cycle reaches 100, stop the training and obtain the trained vehicle detection model.

[0083] In order to further improve the target detection effect of the model, this application also improves the loss function of the original YOLOv8 model, adopts the Focal-EIoU loss function as the loss function of the vehicle detection model, improves the position accuracy of the vehicle detection model in the bounding box regression stage, and accelerates the convergence speed of the model. When detecting complex traffic scenes, it can effectively solve the category imbalance problem and ensure the accuracy of the bounding box, thereby improving the performance and robustness of the vehicle detection algorithm. For example, the calculation formula of the Focal-EIoU loss function is shown in formula (15):

[0084] L (Focal-EIoU) =IoU γ L EIoU (15)

[0085] Among them, IoU is the intersection over union ratio; γ is a hyperparameter that controls the degree of suppression of negative samples. When the background is extremely unbalanced, increasing the γ coefficient can improve the performance of the loss function. EIoU Represents EIoU loss. By introducing a loss penalty for the aspect ratio, the EIoU loss effectively solves the shortcomings of the CIoU (Complete-IoU) loss function when considering the simultaneous increase or decrease of width and height, ensuring that the width and height difference between the real box and the predicted box is minimized, and achieving rapid convergence of the loss function. The EIoU loss function not only improves the accuracy and training efficiency of vehicle detection, but also enhances the robustness of the model for complex scenes and various vehicle detections. The EIoU loss includes the IoU loss L IoU , distance loss L dis and aspect ratio loss L asp , the calculation formula of EIoU loss is shown in formula (16):

[0086] L EIoU =L IoU +L dis +L asp (16)

[0087] Among them, L IoU It represents the IoU loss, which is usually used to measure the similarity between the predicted bounding box and the true bounding box. The calculation formula of IoU loss can be shown as formula (17):

[0088]

[0089] In formula (17), IOU is defined as the ratio of the intersection area to the union area of ​​the predicted box and the labeled box (or true box), which ranges from [0, 1].

[0090] L dis represents the distance loss, which uses the square of the Euclidean distance between the center point of the predicted box and the labeled box, and normalizes it by dividing it by the square of the diagonal length of the minimum circumscribed rectangle of the two to enhance the accuracy of the model in locating the predicted box position; L asp Represents the aspect ratio loss, which focuses on minimizing the deviation between the predicted box and the labeled box in the width and height dimensions, so that the model can learn a more accurate bounding box shape. The calculation formula of the distance loss can be shown in formula (18), and the calculation formula of the aspect ratio loss can be shown in formula (19).

[0091]

[0092] Among them, ω c and h c are the minimum bounding box width and height covering the prediction box and the annotation box, b and bgt are the center points of the prediction box and the annotation box, ω, h, ω gt With h gt are the width and height of the prediction box and the annotation box respectively, ρ 2 represents the distance square function. Please refer to Figure 6 , Figure 6 A parameter diagram of EIoU loss is given. By decomposing the loss function into IoU loss, distance loss, and aspect ratio loss, L EIoU It can more comprehensively evaluate the geometric differences between the predicted box and the labeled box, and directly optimize the size of the predicted box to make it closer to the labeled box, thereby achieving faster convergence speed and more accurate positioning effect.

[0093] In summary, as shown in formulas (15) to (19), the parameters required for calculating the loss function of the vehicle detection model provided by the present application include: the first center point, first width and first height of the prediction box (the prediction box is obtained by predicting after the hybrid fusion feature network is input into the head network of the vehicle detection model); the second center point, second width and second height of the annotation box (the annotation box is the annotation data of the image after preprocessing or the real box); the third width and third height of the minimum external frame corresponding to the prediction box and the annotation box; the intersection area and union area of ​​the prediction box and the annotation box, etc. After obtaining the above parameters, the first center point, the second center point, the first width, the second width, the first height, the second height, the third width, the third height, the intersection area and the union area can be brought into the above loss function calculation formula to determine the prediction regression loss of the vehicle detection model.

[0094] In combination with formula (15) to formula (19), the specific process of determining the predicted regression loss based on the above parameters may also include the following steps, that is, step 104 may also include steps 104A to 104E.

[0095] Step 104A, determining the ratio of the intersection area to the union area as the intersection-to-union ratio loss.

[0096] First, the intersection-over-union loss, i.e., the IOU in formula (15), is determined based on the ratio of the intersection area and the union area of ​​the predicted box and the labeled box.

[0097] Step 104B: determining the similarity loss between the predicted box and the labeled box based on the intersection area, the union area and a preset value.

[0098] The preset value is 1. Combined with formula (17), the similarity loss between the predicted box and the labeled box can be calculated based on the intersection area, the union area and the preset value.

[0099] Step 104C: determine the distance loss between the predicted box and the labeled box based on the first center point, the second center point, the third width, and the third height.

[0100] Combined with formula (18), based on the first center point of the prediction box, the second center point of the annotation box, the third width and the third height of the minimum bounding box corresponding to the prediction box and the annotation box, the distance loss between the prediction box and the annotation box can be calculated.

[0101] Step 104D: determine the aspect ratio loss between the predicted box and the labeled box based on the first width, the second width, the first height, the second height, the third width, and the third height.

[0102] Combined with formula (19), based on the first width of the prediction box, the second width of the annotation box, the first height of the prediction box, the second height of the annotation box, the third width and third height of the minimum bounding box corresponding to the prediction box and the annotation box, the aspect ratio loss between the prediction box and the annotation box is calculated.

[0103] Step 104E, determining the prediction regression loss of the vehicle detection model based on the intersection-over-union loss, the similarity loss, the distance loss, and the aspect ratio loss.

[0104] Then, the prediction regression loss of the vehicle detection model is determined based on the intersection-over-union loss, similarity loss, distance loss and aspect ratio loss.

[0105] As shown in formula (15) and formula (16), there is also a certain operational relationship between the intersection-over-union loss, similarity loss, distance loss, aspect ratio loss and the final prediction regression loss. Optionally, step 104E may also include steps 104E1 to 104E3.

[0106] Step 104E1, determining the first part of the loss based on the intersection-over-union loss and preset hyperparameters.

[0107] Step 104E2, the sum of the similarity loss, the distance loss and the aspect ratio loss is determined as the second part of the loss.

[0108] Step 104E3, the product of the first part of the loss and the second part of the loss is determined as the prediction regression loss of the vehicle detection model.

[0109] First, the first part of the loss is determined according to the intersection-over-union loss and the preset hyperparameters. Then the sum of the similarity loss, distance loss and aspect ratio loss is determined as the second part of the loss (i.e., EIoU loss). Then, the product of the first part of the loss and the second part of the loss is finally determined as the prediction regression loss of the vehicle detection model.

[0110] The Focal-EIOU loss function introduced in the embodiment of the present application separates high-quality anchor boxes from low-quality anchor boxes from the perspective of gradient, optimizes the sample imbalance problem in the bounding box regression task, makes the regression process focus on high-quality anchor boxes, and improves the regression accuracy.

[0111] When calculating the loss function, it involves calculating the position of the prediction box. For example, the process of calculating the position of the prediction box can be as follows:

[0112] (1) Assuming that the coordinate offset of the center point of the bounding box output by the model is (dx, dy) and the coordinates of the reference point are (cx, cy), the calculation formula for the coordinates of the center point of the bounding box can be shown in formula (20) and formula (21):

[0113]

[0114] Among them, anchor_width and anchor_heigth represent the width and height of the prediction box respectively, and scale_factor is the scaling ratio, which is used to convert the offset output by the model into absolute coordinates relative to the input image.

[0115] (2) Assume that the border width and height scaling factors of the model output are (dw, dh), and the width and height of the preset box are (anchor width , anchor_height), the actual height and width of the predicted box can be calculated as shown in formula (22) and formula (23):

[0116] bw=anchor_width·exp(dw) (22)

[0117] bh=anchor_height·exp(dh) (23)

[0118] (3) Based on the known coordinates of the center point (bx, by), width (bw) and height (bh) of the prediction box, the coordinates of the four vertices of the prediction box can be calculated. The coordinates of the upper left corner of the prediction box are (bx-bw / 2, by-bh / 2), and the coordinates of the lower right corner are (bx+bw / 2, by+bh / 2).

[0119] In addition to outputting the prediction box, the vehicle detection model also outputs the category to which the vehicle target belongs. In one possible implementation, the Cls branch of the vehicle detection model is used to predict the category of the detection target (i.e., the vehicle). After the image data passes through the network, the Cls branch will output a vector whose dimension is (1, nc, 8400), where nc represents the number of categories and 8400 represents the total number of grid cells predicted on feature maps of three different scales. Each element of the vector is processed by the Sigmoid function, which maps each element to the interval (0, 1), indicating the confidence of the corresponding category. The Sigmoid function calculation formula can be shown in formula (24).

[0120]

[0121] This processing method can limit the output to the probability range, maintain the independence between category predictions, and introduce nonlinear relationships. The output of the Sigmoid function is not only easy to interpret as a probability value, but also convenient for gradient calculation and nonlinear activation, which improves the generalization ability of the model.

[0122] After obtaining the preliminary prediction results of the model, it is necessary to use the non-maximum suppression (NMS) module to process the output prediction box. The main purpose of NMS processing is to filter out highly repeated detection boxes that may point to the same target, and only retain the detection boxes with the highest confidence. The NMS module filters out low-confidence prediction results according to the preset confidence threshold, and converts the remaining prediction results from XYWH format to XYXY format to more intuitively represent the coordinates of the bounding box. The converted prediction results are processed using the NMS algorithm, and the overlapping prediction boxes are removed by calculating the intersection-over-union ratio between the bounding boxes, retaining the optimal prediction results for each vehicle target.

[0123] After obtaining the NMS processing result, the NMS processing result needs to be processed by the Scale_boxes module. The Scale_boxes module scales the predicted result after NMS processing and maps it back to the size of the original input image. During the processing, considering the gray edge filled by the LetterBox, the prediction box is offset corrected to ensure the accuracy of the bounding box. The corrected prediction box is cropped to prevent it from exceeding the actual size of the image. The steps of cropping the prediction box include: determining the actual size of the original input image, that is, the width W and height H of the image; calculating the cropping range to ensure that the upper left corner coordinates (x 1 ,y 1 ) is not less than (0, 0), the lower right corner coordinate (x 2 ,y 2) is not greater than (W, H) to limit it to the actual size of the image; each corrected prediction box is checked, and if its coordinates exceed the image boundary, it is adjusted according to the image size, such as 1 If it is less than 0, set it to 0. 2 If it is greater than W, set it to W. The same goes for y. 1 and 2 After these adjustments, the predicted box is the cropped result.

[0124] Step 105: input the test image into the trained vehicle detection model to obtain a vehicle detection result of the test image.

[0125] Exemplarily, the model training environment of the present application is: running under Windows 10 operating system, CPU 1GHz, memory 16G, hard disk 1T, graphics card NVIDIA GTX1070Ti, programming language is Python3.8, and model training is performed in the deep learning framework Pytorch1.2.1 version. In the process of training the vehicle detection model in the present application, a total of 100 training cycles (epochs) are set to ensure that the vehicle detection model fully converges. The number of images processed at a single time (batch size) is set to 8 to balance the computational efficiency and memory usage; the number of threads (works) used by the CPU during data loading is 8; the learning rate is 0.01; the loss function selects Focal-EIoU Loss; in terms of optimization algorithm, considering the needs of large data volume processing and the gradient update characteristics during training, the present application adopts stochastic gradient descent (Stochastic Gradient Descent, SGD) to train and optimize the neural network (i.e., the vehicle detection model), and SGD updates the model parameters by samples one by one or small batches of samples instead of using the entire data set. This method significantly improves the computational efficiency. The training parameters of the vehicle detection model are shown in Table 1.

[0126] Table 1. Training parameters of vehicle detection model

[0127]

[0128]

[0129] In summary, the embodiment of the present application provides a vehicle detection algorithm based on the FasterNet lightweight framework and multi-head self-attention mechanism. The lightweight FasterNet network is used to replace the YOLOv8 backbone network CSPDarkNet53, and the C2f module is reconstructed to reduce the complexity of the model, realize the lightweight of the model, and improve the speed of vehicle detection. By introducing the multi-head self-attention mechanism, the model's ability to understand complex spaces is enhanced. The Focal-EIoU Loss loss function is used to enhance the detection accuracy of multiple vehicle targets and small vehicle targets in complex traffic road scenes, and improve the accuracy of vehicle detection.

[0130] In order to verify the vehicle detection performance of the vehicle detection model (vehicle detection algorithm) proposed in this application in traffic scenarios, this application uses the evaluation indicators commonly used in target detection to comprehensively evaluate the vehicle detection model. Six objective evaluation indicators are mainly used, namely precision, recall, mean average precision (mAP), billion floating point operations (GFLOPs), frame rate per second (FPS) and model parameters (Params).

[0131] Precision and recall directly reflect the vehicle detection capability of the model and are the basic indicators for evaluating model performance. mAP, as a comprehensive indicator of precision and recall, can more comprehensively reflect the overall performance of the model and is a widely used evaluation indicator in the field of target detection. The three indicators of GFLOPs, FPS and Params quantitatively evaluate the model from three aspects: computational complexity, processing speed and model size, providing an important reference for the practical application of the model. The six objective evaluation indicators used in this application comprehensively and effectively evaluate the performance of the vehicle detection algorithm from multiple dimensions, providing a reference for further optimization and practical application.

[0132] Precision refers to the proportion of samples that are actually positive among the samples predicted by the model to be positive. Recall refers to the ability of the model to find all positive samples. The larger the value of precision, the lower the false positive rate of the model. The larger the value of recall, the lower the missed detection rate. The calculation formulas of precision and recall are shown in formula (25) and formula (26) respectively. Among them, TP represents the number of samples detected correctly, FP represents the number of samples detected incorrectly, and FN represents the number of samples missed.

[0133]

[0134] The curve drawn with recall as the horizontal axis and precision as the vertical axis is the PR curve. The integrated average value below the curve is the AP value of the category, and its calculation formula is shown in formula (27).

[0135]

[0136] mAP is the sum of the average precision of all labels divided by the total number of all categories. The larger the value of the average precision, the better the overall performance of the model. Its calculation formula is shown in formula (28).

[0137]

[0138] mAP@0.5 (Mean Average Precision at Intersection over Union = 0.50) measures the performance of the object detection model by calculating the average of the average precision over all detection categories under the condition that the intersection over union threshold is set to 0.5. The intersection over union threshold is set to 0.5, which means that the prediction is considered a correct detection only when the intersection area of ​​the predicted bounding box and the true bounding box accounts for a ratio of 50% or more of their union area. The higher the mAP@0.5 value, the more accurately the model can detect vehicles during the detection process, while maintaining a low false detection rate (misdetecting non-vehicle objects as vehicles) and missed detection (failure to detect actual vehicles) rate.

[0139] GFLOPs refers to the theoretical billion floating-point operations per second required by a model for forward propagation or backward propagation, and is usually used as an indicator to evaluate the theoretical amount of computation required by a model during forward reasoning or reverse optimization. A lower GFLOPs value means that the model requires fewer floating-point operations per unit time when it is executed, reflecting its advantage in computational efficiency. This advantage is not only reflected in processing speed, but also in its friendliness to computing resources. Under the same hardware conditions, low GFLOPs models can be more easily deployed on various platforms, such as embedded systems or edge computing devices.

[0140] FPS refers to the number of still image frames that the model can efficiently process and continuously output detection results per second. This indicator is directly related to the real-time response capability of the model, that is, the speed at which the model processes data and immediately feedbacks the results. The higher the frame rate, the more image frames the model can process per second, the stronger the real-time response capability, and the model can be seamlessly compatible and efficiently integrated into various real-time systems to ensure the overall smoothness and immediacy of the system. The FPS calculation formula is shown in formula (29), where N(p) represents the total number of processed images and T(p) represents the time to process the image.

[0141]

[0142] The number of model parameters is a key indicator for measuring model complexity, reflecting the total number of all adjustable weight parameters in the model. The number of model parameters is directly related to the structural complexity of the model and its demand for computing resources. Models with fewer parameters have a more compact structure, which not only reduces the computational burden during training, but also reduces the hardware resource consumption required during prediction, thereby improving the efficiency of vehicle detection.

[0143] In order to verify the advancement and effectiveness of the algorithm proposed in this application, the algorithm (vehicle detection model) proposed in this patent is compared with the EfficientDet model, D-FINE model and YOLOv11 model. The performance comparison results are shown in Table 2.

[0144] Table 2 Performance comparison results of different methods

[0145]

[0146]

[0147] In terms of accuracy, the D-FINE model ranks first with 84.61%, and has the highest accuracy when detecting vehicles. YOLOv11 and EfficientDet follow closely behind, reaching 83.62% and 83.54% respectively, with high accuracy. Although the algorithm proposed in this application is slightly lower than the first three, it also shows a high accuracy with an accuracy of 81.57%.

[0148] In terms of recall rate, D-FINE continues to lead, reaching 85.28%, indicating its comprehensiveness and low missed detection rate when detecting vehicles. In contrast, YOLOv11 has the lowest recall rate of 80.59%, indicating that YOLOv11 has some missed detections in some complex scenarios. The recall rates of EfficientDet and the algorithm proposed in this application are stable at 82.31% and 81.37%, respectively.

[0149] mAP@0.5 is an important indicator to measure the overall performance of the object detection algorithm. D-FINE is 84.57%, maintaining its leading position. The mAP@0.5 of the algorithm proposed in this application reached 83.81%, second only to D-FINE.

[0150] The EfficientDet model maintains a low level in both GFLOPs and parameter count, indicating that it has a good balance between computational complexity and model size, and is suitable for scenarios with certain requirements for computing resources. The D-FINE model has the highest computational complexity and parameter count, indicating that while D-FINE pursues higher performance, it also increases computational complexity and model size. The YOLOv11 model has a similar number of parameters to EfficientDet, but has a higher GFLOPs, indicating that YOLOv11 has advantages in speed and accuracy, and is suitable for scenarios with high requirements for computing resources. The algorithm proposed in this application performs well in GFLOPs, which is slightly higher than EfficientDet, but much lower than D-FINE and YOLOv11, and has the lowest number of parameters. This shows that while achieving similar performance, the algorithm greatly reduces the model's computational complexity and model size, making it more practical on resource-constrained devices.

[0151] On the whole, the algorithm proposed in this application is slightly lower than the D-FINE model in terms of precision, recall and average precision. This is mainly because the D-FINE model introduces FDR to refine the distribution representation of the bounding box, and transfers positioning knowledge from the deep layer to the shallow layer through GO-LSD to achieve accurate positioning of the target. Compared with the YOLOv11 model, although the accuracy and average accuracy of the algorithm proposed in this application are slightly reduced, the recall rate index is improved, and the GFLOPs and parameter quantity indicators are significantly reduced, indicating that although the algorithm proposed in this application has some reductions in individual accuracy indicators, it has advantages in comprehensive performance, real-time processing capabilities and resource efficiency. After the vehicle detection model is trained, in actual traffic scenarios, the real-time image captured can be input into the vehicle detection model to output real-time detection results of the vehicle.

[0152] The EfficientDet model is superior to the algorithm proposed in this application in terms of accuracy, recall, and GFLOPs, but the EfficientDet model has a large number of parameters and a relatively long model training time, which does not meet the real-time requirements in practical applications. In general, the GFLOPs value of the algorithm proposed in this application is slightly higher than that of the EfficientDet model, but much lower than that of the D-FINE model and the YOLOv11 model, and the number of model parameters is greatly reduced to 4.78, which is much lower than the other three models, indicating that the algorithm proposed in this application has achieved model lightweight while maintaining high performance.

[0153] This application conducted relevant experiments on the proposed algorithm and selected multiple images from the UA-DETRAC dataset for vehicle detection testing. The following is an example of an image from the UA-DETRAC dataset. Figure 8The experimental results of vehicle detection for this image using the algorithm proposed in this application, the EfficientDet model, the D-FINE model, and the YOLOv11 model are given. Figure 8 It can be seen that all four methods can detect vehicles on the main roads of the city and mark the location of the vehicles and the detection probability.

[0154] Depend on Figure 8 It can be seen that the algorithm proposed in this application uses a red border to mark the location of the detected vehicle and gives the detection probability. The detected vehicle bounding box has a high accuracy rate, few false detections, and can accurately detect vehicles of different colors and types, improving the detection accuracy of multiple vehicle targets and small vehicle targets in complex traffic scenes. This is mainly due to the optimization of the algorithm in using the FasterNet backbone network to perform preliminary feature extraction on the preprocessed image, using the neck network based on the F-C2f module for multi-scale feature fusion, and using the Focal-EIoU loss function to calculate the bounding box prediction results.

[0155] The EfficientDet model uses a green border to mark the location of the detected vehicle and gives the detection probability. It performs well in vehicle detection, but there are repeated detections of the same vehicle, which may be related to multiple factors such as the model structure and feature fusion network.

[0156] The D-FINE model uses a purple border to mark the location of the detected vehicle and gives the detection probability. In the vehicle detection results, there are relatively many missed detections and false detections. This may be due to the limited generalization ability of the model, which makes the model unable to accurately detect vehicles in some scenarios.

[0157] The YOLOv11 model uses a blue border to mark the location of the detected vehicle and gives the detection probability. It can detect large vehicle targets well, but the effect is poor when detecting small vehicle targets in the distance. This may be because the model can only extract limited feature information when processing small vehicle targets, which increases the difficulty of the model to accurately detect small vehicle targets.

[0158] Please refer to Fig. 9 , which is a structural diagram of a vehicle detection device based on the FasterNet lightweight framework and feature fusion provided in an embodiment of the present application. For example, Fig. 9 As shown, the device 900 includes:

[0159] The preprocessing module 901 is used to input an image and perform preprocessing on the image, wherein the preprocessing mainly includes LetterBox and normalization operations;

[0160] Feature extraction module 902, inputs the preprocessed image into the backbone network of the vehicle detection model for feature extraction to obtain a preliminary extracted feature map, wherein the feature extraction network in the backbone network is a feature extraction network based on a lightweight FasterNet network and a multi-head attention mechanism;

[0161] The feature fusion module 903 inputs the initially extracted feature map into the neck network based on the F-C2f module for multi-scale feature fusion to obtain a mixed fusion feature map, wherein the neck network is a neck network based on the F-C2f module;

[0162] Model training module 904, inputs the hybrid fusion feature map into the head network of the vehicle detection model, uses the Focal-EIoU loss function for training, and stops training when the training cycle reaches the set value. In the process of training the vehicle detection model in this application, a total of 100 training cycles (epochs) are set to ensure that the vehicle detection model is fully converged;

[0163] The model testing module 905 is used to input the test image into the trained vehicle detection model to obtain the vehicle detection result of the test image.

[0164] Optionally, the backbone network includes the FasterNet network, and the multi-head self-attention mechanism layer has 4 attention heads.

[0165] Optionally, the model training module 904 is further used to:

[0166] Inputting the mixed fusion feature map into the head network of the vehicle detection model to obtain a prediction box;

[0167] Acquire a first center point, a first width and a first height of the prediction box, and a second center point, a second width and a second height of the annotation box of the preprocessed image;

[0168] Determine a third width and a third height of the minimum bounding box corresponding to the prediction box and the annotation box;

[0169] Determine the intersection area and the union area of ​​the prediction box and the annotation box;

[0170] Determining a prediction regression loss of the vehicle detection model based on the first center point, the second center point, the first width, the second width, the first height, the second height, the third width, the third height, the intersection area, and the union area;

[0171] The vehicle detection model is trained based on the prediction regression loss.

[0172] Optionally, the model training module 904 is further used to:

[0173] The ratio of the intersection area to the union area is determined as the intersection-to-union ratio loss;

[0174] Determining a similarity loss between the predicted box and the labeled box based on the intersection area, the union area and a preset value;

[0175] Determine a distance loss between the prediction box and the annotation box based on the first center point, the second center point, the third width, and the third height;

[0176] Determine an aspect ratio loss between the prediction box and the annotation box based on the first width, the second width, the first height, the second height, the third width, and the third height;

[0177] The prediction regression loss of the vehicle detection model is determined based on the intersection-over-union loss, the similarity loss, the distance loss, and the aspect ratio loss.

[0178] Optionally, the model training module 904 is further used to:

[0179] Determine the first part of the loss based on the intersection-over-union loss and the preset hyperparameters;

[0180] Determine the sum of the similarity loss, the distance loss and the aspect ratio loss as the second part of loss;

[0181] The product of the first part of the loss and the second part of the loss is determined as the predicted regression loss of the vehicle detection model.

[0182] Optionally, the device further comprises:

[0183] A second acquisition module is used to acquire a preprocessed target input image, wherein the target input image includes a target vehicle to be detected;

[0184] A target detection module, used to input the target input image into a trained vehicle detection model to obtain at least one initial detection frame output by the vehicle detection model;

[0185] The target processing module is used to perform non-maximum suppression processing on the initial detection frame to obtain a target detection frame corresponding to the target input image.

[0186] The exemplary embodiment of the present application also provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to enable the electronic device to execute the vehicle detection algorithm and device based on the FasterNet lightweight framework and feature fusion according to the embodiment of the present application when executed by the at least one processor. It should be noted that the electronic device is the device end shown in the above embodiment.

[0187] An exemplary embodiment of the present application also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute a vehicle detection algorithm and device based on a FasterNet lightweight framework and feature fusion according to an embodiment of the present application.

[0188] The exemplary embodiment of the present application also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to enable the computer to execute a vehicle detection algorithm and device based on the FasterNet lightweight framework and feature fusion according to the embodiment of the present application.

[0189] refer to Fig.10 The structured block diagram of the electronic device 1000 that can be used as the server or client of the present application will now be described, which is an example of the hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent the computer device of various forms of digital electronics, such as, laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only used as examples, and are not intended to limit the implementation of the present application described herein and / or required.

[0190] like Fig.10 As shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0191] A plurality of components in the electronic device 1000 are connected to the I / O interface 1005, including: an input unit 1006, an output unit 1007, a storage unit 1008, and a communication unit 1009. The input unit 1006 may be any type of device capable of inputting information to the electronic device 1000, and the input unit 1006 may receive input digital or character information, and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 1007 may be any type of device capable of presenting information, and may include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1008 may include but is not limited to a disk, an optical disk. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0192] Optionally, the electronic device 1000 is further provided with a single-channel EEG signal acquisition module (not shown in the figure) for acquiring EEG signals and transmitting the EEG signals to a signal processor of the electronic device 1000 for processing the EEG signals.

[0193] The computing unit 1001 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above. For example, in an embodiment, Figure 1 The method shown may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. In some embodiments, the computing unit 1001 may be configured to execute the computer program by any other suitable means (e.g., by means of firmware). Figure 1 The method shown.

[0194] The program code for implementing the algorithm of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0195] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0196] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0197] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0198] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0199] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.

Claims

1. A vehicle detection algorithm based on FasterNet lightweight framework and feature fusion, characterized in that: The algorithm includes: Input an image and perform preprocessing on the image, wherein the preprocessing mainly includes LetterBox and normalization operations; The preprocessed image is input into the backbone network of the vehicle detection model for feature extraction to obtain a preliminary extracted feature map, wherein the feature extraction network in the backbone network is a feature extraction network based on a lightweight FasterNet network and a multi-head attention mechanism; Inputting the initially extracted feature map into the neck network for multi-scale feature fusion to obtain a mixed fusion feature map, wherein the neck network is a neck network based on the F-C2f module; Input the hybrid fusion feature map into the head network of the vehicle detection model, use the Focal-EIoU loss function for training, and stop training when the training cycle reaches 100 to obtain a trained vehicle detection model; The test image is input into the trained vehicle detection model to obtain the vehicle detection result of the test image.

2. The algorithm according to claim 1, characterized in that The backbone network includes the FasterNet network, and the multi-head self-attention mechanism layer has 4 attention heads.

3. The algorithm according to claim 1, characterized in that The step of inputting the hybrid fusion feature map into the head network of the vehicle detection model and training using the Focal-EIoU loss function includes: Inputting the mixed fusion feature map into the head network of the vehicle detection model to obtain a prediction box; Acquire a first center point, a first width and a first height of the prediction box, and a second center point, a second width and a second height of the annotation box of the preprocessed image; Determine a third width and a third height of the minimum bounding box corresponding to the prediction box and the annotation box; Determine the intersection area and the union area of ​​the prediction box and the annotation box; Determining a prediction regression loss of the vehicle detection model based on the first center point, the second center point, the first width, the second width, the first height, the second height, the third width, the third height, the intersection area, and the union area; The vehicle detection model is trained based on the prediction regression loss.

4. The algorithm according to claim 3, characterized in that The step of determining the prediction regression loss of the vehicle detection model based on the first center point, the second center point, the first width, the second width, the first height, the second height, the third width, the third height, the intersection area, and the union area includes: The ratio of the intersection area to the union area is determined as the intersection-to-union ratio loss; Determining a similarity loss between the predicted box and the labeled box based on the intersection area, the union area and a preset value; Determine a distance loss between the prediction box and the annotation box based on the first center point, the second center point, the third width, and the third height; Determine an aspect ratio loss between the prediction box and the annotation box based on the first width, the second width, the first height, the second height, the third width, and the third height; The prediction regression loss of the vehicle detection model is determined based on the intersection-over-union loss, the similarity loss, the distance loss, and the aspect ratio loss.

5. The algorithm according to claim 4, characterized in that The step of determining the prediction regression loss of the vehicle detection model based on the intersection-over-union loss, the similarity loss, the distance loss, and the aspect ratio loss includes: Determine the first part of the loss based on the intersection-over-union loss and the preset hyperparameters; Determine the sum of the similarity loss, the distance loss and the aspect ratio loss as the second part of the loss; The product of the first part of the loss and the second part of the loss is determined as the predicted regression loss of the vehicle detection model.

6. The algorithm according to claim 1, characterized in that The algorithm also includes: Acquire a preprocessed target input image, wherein the target input image includes a target vehicle to be detected; Inputting the target input image into a trained vehicle detection model to obtain at least one initial detection frame output by the vehicle detection model; The initial detection frame is subjected to non-maximum suppression processing to obtain a target detection frame corresponding to the target input image.

7. A vehicle detection device based on FasterNet lightweight framework and feature fusion, characterized in that: The device comprises: A preprocessing module, used for inputting an image and performing preprocessing on the image, wherein the preprocessing mainly includes LetterBox and normalization operations; A feature extraction module is used to input the preprocessed image into the backbone network of the vehicle detection model for feature extraction to obtain a preliminary extracted feature map. The feature extraction network in the backbone network is a feature extraction network based on a lightweight FasterNet network and a multi-head attention mechanism; A feature fusion module, used for inputting the initially extracted feature map into the neck network based on the F-C2f module for multi-scale feature fusion to obtain a mixed fusion feature map, wherein the neck network is a neck network based on the F-C2f module; A model training module is used to input the hybrid fusion feature map into the head network of the vehicle detection model, and use the Focal-EIoU loss function for training. When the training cycle reaches the set value, the model stops training; in the process of training the vehicle detection model in this application, a total of 100 training cycles (epochs) are set to ensure that the vehicle detection model is fully converged; The model prediction module is used to input the test image into the trained vehicle detection model to obtain the vehicle detection result of the test image.

8. An electronic device comprising: processor; as well as Memory for storing programs, Wherein, the program includes instructions, which, when executed by the processor, enable the processor to execute the vehicle detection algorithm and device based on the FasterNet lightweight framework and feature fusion according to any one of claims 1-7.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the vehicle detection algorithm and device based on the FasterNet lightweight framework and feature fusion according to any one of claims 1-7.

Citation Information

Patent Citations

  • Vehicle target detection method based on improved YOLOv3 algorithm

    CN116682090A

  • Lightweight network structure and method for real-time target detection in view angle of near-field security unmanned aerial vehicle

    CN118691938A