RoadObstacle Det model for road obstacle detection and construction method and detection method of RoadObstacle Det model
By proposing a RoadObstacleDet model based on YOLO in road obstacle detection, combining a specific network architecture and data set expansion method, the problems of low detection accuracy and high computing resource consumption in the prior art are solved, and a more efficient and robust road obstacle detection effect is achieved.
Patent Information
- Application Number
- CN202510202171.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has problems in road obstacle detection with low detection accuracy, poor robustness, large computing resource consumption and poor response to severe weather conditions.
A RoadObstacleDet model based on YOLO object detection framework is proposed. Through the combination of backbone network, neck module and prediction head, combined with LowFormer module, internal convolution module and shuffle attention module, we can improve feature extraction and detection accuracy. At the same time, the dataset extension method is used to increase the training samples to improve the generalization ability and detection performance of the model.
Significantly improve obstacle detection on highways, improve detection applicability under different conditions, improve detection accuracy and model calculation efficiency, and ensure high performance in real-time detection applications.
Smart Images

Figure CN120047923A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road obstacle detection, and in particular to a RoadObstacleDet model for road obstacle detection, a construction method thereof, and a detection method thereof. Background Art
[0002] Roads are critical infrastructures that are vital to the socio-economic development of regions and cities, primarily due to their role in transportation and communication. As transportation systems continue to advance, improving the safety and efficiency of transportation infrastructure has become a top priority. However, many hazards can affect transportation infrastructure, leading to cascading failures. In recent years, the number and severity of natural disasters have increased due to the impact of climate change. These types of disasters include heat waves, heavy rains, river flooding, storms, landslides, droughts, wildfires and avalanches. Natural disasters can cause road obstructions in the form of debris (the fall of fragments of pavement material), collapse (sudden catastrophic damage to the road structure) and landslides (downhill or lateral movement due to instability of soil or rock beneath the road).
[0003] Autonomous driving systems rely on road obstacle detection to assess the location, tracking, distance, and speed of obstacles. Methods such as LiDAR and millimeter wave radar have high detection accuracy and strong robustness. The latest advances in sensor fusion usually combine LiDAR and cameras to achieve 3D object detection. LiDAR determines the distance of nearby objects by evaluating the time it takes for laser pulses to propagate, providing accurate 3D data at short distances. In contrast, cameras provide rich visual information. Advanced driver assistance systems (ADAS), such as collision avoidance and adaptive cruise control (ACC), have long used radar, which has significant advantages. Compared with LiDAR and cameras, radar can remain stable in adverse weather conditions and detect objects at longer distances (up to 200 meters). In addition, radar uses the Doppler effect to estimate the speed of detected objects, without the need for time data. Radar point clouds require less processing power to detect objects than LiDAR point clouds. This is because radar data is generally lower in resolution and less computationally demanding, but still provides valuable information about the speed, distance, and position of objects, especially in adverse weather conditions. However, both cameras and lidar are susceptible to bad weather, which greatly reduces their effectiveness and range. Limited radar datasets for autonomous driving research are also a major challenge. Finally, the high cost of lidar and radar is also a significant issue.
[0004] In addition, the rapid development of artificial intelligence has enabled deep learning technology to detect problems such as cracks, collapses, and collapses from various road images. Deep learning provides strong support for the automatic analysis of images related to cracks, collapses, and collapses. Convolutional neural networks (CNNs) are part of deep learning technology, which can automatically identify and classify structural defects in visual data, including cracks, collapses, and collapses. The use of this technology can not only improve the accuracy of road obstacle detection, but also greatly improve the processing speed, enabling autonomous vehicles to quickly analyze large amounts of data during actual operation. At present, road obstacle detection algorithms based on deep learning are mainly divided into two categories: one-stage methods and two-stage methods. R-CNN, Fast-RCNN, Faster-RCNN, and Mask R-CNN are two-stage algorithms. The YOLO series is classified as a single-stage object detection method. Single-stage detection algorithms are known for their fast detection speed and achieve an effective balance between accuracy and efficiency. Therefore, they are particularly suitable for real-time detection of autonomous vehicles. However, the limited computing power of edge computing devices poses a huge challenge to the deployment of most existing object detection models. Although some lightweight detection models can be implemented on hardware devices, they often exhibit low accuracy and cannot meet the needs of practical detection applications. Summary of the invention
[0005] In view of the problems existing in the prior art, the present invention proposes a RoadObstacleDet model for road obstacle detection. The RoadObstacleDet model of the present application is a new model based on the YOLO target detection framework.
[0006] Specifically, a RoadObstacleDet model for road obstacle detection of the present invention includes: a backbone network module for receiving an image to be detected and extracting features in the image, a neck module for receiving features output by the backbone network module and integrating these features, and a prediction head module for receiving output of the neck module and responsible for generating detection results;
[0007] In the backbone network module, the input image first passes through the CBS module to extract the activation feature map with lower resolution and obtain A0 of size (32, 128, 128). Then A0 is processed by a series of convolutional layers + C3Ghost modules to generate feature maps of different sizes.
[0008] Based on the above scheme, there are four convolutional layer + C3Ghost modules connected together in the backbone network module;
[0009] Among them: the CBS module outputs A0 of size (32, 128, 128) through the first convolutional layer + C3Ghost module, and outputs A2 of size (64, 64, 64); A2 enters the second convolutional layer + C3Ghost module, generating an output A of size (128, 32, 32); the third convolutional layer + C3Ghost module converts A into B of size (256, 16, 16); the fourth convolutional layer + C3Ghost module converts B into C of size (512, 8, 8); subsequently, A, B and C are used as inputs of the neck module.
[0010] Based on the above scheme, in the neck module, C is transmitted to the SPPF module to generate D of size (512, 8, 8); then, D is transmitted to the GhostConv module to generate E of the same size; subsequently, E is transmitted to the upsampling layer module to obtain F of size (512, 16, 16); then, F is processed by the convolutional layer module to obtain G;
[0011] At the same time, B and G are element-wise added to obtain H of size (256, 16, 16); H is passed to the C3Ghost module to obtain J of the same dimension; then, J is sent to the convolutional layer module and then passed through the upsampling layer module to obtain L of size (128, 32, 32); L is processed by the inner convolution module to obtain M of the same dimension;
[0012] At the same time, A is transferred to the LowFormer module to obtain I of the same size as A; I and M are added vector by vector to obtain N.
[0013] Based on the above scheme, in the neck module, N is used and O is obtained through the C3Ghost module, and then O is processed by the shuffle attention module to obtain R with a size of (128, 32, 32). Finally, R is sent to the detector 1 in the detector module;
[0014] O also passes through an inner convolution module to obtain U; at the same time, J first passes through a LowFormer module to obtain P, whose size is kept as (256, 16, 16); P and U are added element by element to obtain V, which is then processed by the shuffle attention module to obtain S of size (256, 16, 16), and finally S is sent to detector 2 in the detector module;
[0015] O also passes through an inner convolution module to obtain W; at the same time, E is used and enters the LowFormer module to obtain Q with a size of (512, 8, 8); Q and W are added element by element to obtain X, and the size of X remains unchanged (512, 8, 8); then, X is processed by the shuffle attention module to obtain T with a size of (512, 8, 8), and finally T is sent to the detector 3 in the detector module.
[0016] The present invention also provides a method for constructing the above model. During training, in order to increase the number of training samples and improve model performance, a data set expansion method is adopted. The data set expansion method is as follows:
[0017] First, road obstacles are segmented using annotated bounding boxes to create more diverse training data and enhance the generalization ability of the model;
[0018] Each segmented image containing a single type of obstacle is then distorted in various ways to simulate real-world noise conditions; these include Gaussian blurring, resizing, adjusting brightness and contrast, and / or adding salt and pepper noise.
[0019] The processed images are then copied and pasted, specifically merged into the empty road images, further expanding the dataset.
[0020] The present invention also provides a method for detecting road obstacles, using the above-mentioned model.
[0021] The present invention also provides a server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method for detecting road obstacles when executing the computer program.
[0022] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and is characterized in that when the computer program is executed by a processor, the steps of the above-mentioned method for detecting road obstacles are implemented.
[0023] Beneficial effects of the present invention:
[0024] The RoadObstacleDet model of the present invention can significantly improve obstacle detection on highways and shows stronger applicability in detecting road obstacles under different conditions.
[0025] In the model of the present invention, the LowFormer module is used to extract information from different representation subspaces at multiple locations. This module uses low-resolution processing technology to make the execution of the attention mechanism more computationally efficient. In addition, the addition of the inner convolution module enables the spatial and channel domain attributes of the neck layer to be transformed and fused. This enhances the receptive field information and improves the channel representation of the feature map. In addition, the "Shuffle Attention" function is added to the RoadObstacleDet model to group the channel dimensions into sub-features, thereby effectively suppressing noise and highlighting important areas.
[0026] During training, data augmentation significantly increased the number of training images. Distorted images of road obstacles were superimposed on images without obstacles using copy-paste techniques, which improved the model’s generalization and ability to identify obstacles in different situations. As a result, a large number of annotated training samples were created, ultimately improving the model’s performance and the accuracy of detecting road obstacles. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 The overall architecture diagram of the RoadObstacleDet model of the present invention;
[0028] Figure 2 The training loss of the model of the present invention on the customized Shandong highway data set in Example 3;
[0029] Figure 3 The precision and recall rate indicators of applying the model of the present invention to the customized Shandong highway data set in Example 3;
[0030] Figure 4 The mAP@0.5 and mAP@0.5:0.95 indicators of the RoadObstacleDet model evaluated on the customized Shandong highway dataset in Example 3;
[0031] Figure 5 The Shandong road obstacle image samples and the object detection results of the RoadObstacleDet model used in Example 3;
[0032] Figure 6 The training loss of the model of the present invention on the customized Guizhou highway data set in Example 3;
[0033] Figure 7 The precision and recall rate indicators of applying the model of the present invention to the customized Guizhou highway data set in Example 3;
[0034] Figure 8 The mAP@0.5 and mAP@0.5:0.95 indicators of the RoadObstacleDet model evaluated on the customized Guizhou highway dataset in Example 3;
[0035] Fig. 9 The Guizhou road obstacle image samples and the object detection results of the RoadObstacleDet model used in Example 3; DETAILED DESCRIPTION
[0036] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.
[0037] Example 1
[0038] The present invention provides a new RoadObstacleDet model for road obstacle detection. The structure of the model is as follows: Figure 1 As shown, it consists of three main parts: the backbone that extracts features, the neck that integrates these features, and the prediction head that is responsible for generating detection results.
[0039] The initial size of the input image (including road obstacles) is 3×H0×W0, where 3 represents the number of color channels, H0 represents the height, and W0 represents the width. In this study, the original image taken by the camera or drone is resized to 3×256×256. In the backbone module, the input image first passes through the CBS module to extract the activation feature map with lower resolution and obtain A0 with a size of (32,128,128). After that, A0 is processed by a series of convolutional layers + C3Ghost modules to generate feature maps of different sizes. After passing through the first convolutional layer + C3Ghost module, the output A2 has a size of (64,64,64). Then, A2 enters the next convolutional layer + C3Ghost module to produce A with a size of (128,32,32). The third convolutional layer + C3Ghost module converts A to B with a size of (256,16,16). Finally, the last convolutional layer + C3Ghost module transforms B into C, and the resulting size is (512, 8, 8). Subsequently, A, B, and C are used as input to the neck layer.
[0040] In the neck module, C is passed to the SPPF (Spatial PyramidPooling Fast) module to generate D of size (512,8,8). Then, D is passed to the GhostConv module to generate E of the same size. Subsequently, E is passed to the UpSampling layer to obtain F of size (512,16,16). Then, F is processed by the convolutional layer module to obtain G. At the same time, B and G are element-wise added to obtain H of size (256,16,16). H is passed to the C3Ghost module to obtain J of the same dimension. Then, J is passed to the convolutional layer module and then passed through the UpSampling module to obtain L of size (128,32,32). L is processed by the inner convolution module to obtain M of the same dimension. In the other branch, A is transferred to the LowFormer module to obtain I of the same size as A. I and M are vector-added to obtain N.
[0041] Three different branches are formed and input into three detectors after processing respectively:
[0042] In the first branch, N is used and passed through the C3Ghost module to obtain O, and then processed by ShuffleAttention to obtain R, and the dimension of all data is (128, 32, 32). R is sent to detector 1. O also passes through a content convolution module to obtain U.
[0043] In the second branch, J is used and first passes through a LowFormer module to obtain P, whose size is kept as (256, 16, 16). P and U are element-wise added to obtain V, which is then processed by shuffle attention to obtain S, whose size is (256, 16, 16). S is fed into detector 2. O also passes through a content convolution module to obtain W.
[0044] In the third branch, E is used and fed into the LowFormer module to obtain Q of size (512, 8, 8). Q and W are element-wise added to obtain X, and the size of X remains unchanged (512, 8, 8). Then, X is processed by shuffle attention to obtain T and fed into detector 3.
[0045] R has a size of (128, 32, 32) and is fed into detector 1; S has a size of (256, 16, 16) and is fed into detector 2; and T has a size of (512, 8, 8) and is fed into detector 3. The three prediction heads in RoadObstacleDet use different resolutions (80×80, 40×40, and 20×20) for the input data, greatly improving its detection capabilities in various application scenarios. This approach leverages the architecture of YOLO. The model classifies each detection box to assess the presence of the target object and performs a regression step to accurately determine its location and size. After this step is completed, non-maximum suppression (NMS) is used to filter out overlapping or redundant bounding boxes, leaving only the most representative bounding boxes. This process improves the overall detection accuracy and ensures that the model outputs the best bounding box for the target object.
[0046] In the present invention, CBS can extract features from the input feature map through a convolution layer and perform nonlinear transformation; CBS combines convolution layers, batch normalization and SiLU activation function (Sigmoid Linear Unit). It extracts features from the input feature map through a convolution layer and performs nonlinear transformation through a batch normalization layer and an activation function.
[0047] One disadvantage of the prior art models (such as the R-CNN model) is that it is difficult to handle objects of different sizes, and spatial information may be lost during the feature extraction process. To solve this problem, in the model of the present invention, we introduced the SPPF module. The SPPF module helps manage objects of different sizes by aggregating features of multiple spatial scales. This method can ensure that the spatial context is retained while extracting features, prevent the loss of key information, and thus improve the performance of the model. SPPF normalizes the feature maps from different region proposals, enabling the model to better manage objects of different scales. This improves the accuracy and robustness of the object detection process. SPPF optimizes the pooling operation, while reducing the computational load, while maintaining the model's detection performance for multi-scale targets. SPPF captures feature information of different scales through an efficient pooling strategy, including the following components: Feature map input: The SPPF layer accepts the feature map of the previous layer as input; Fast pooling operation: It usually uses several smaller pooling layers (such as three 3x3 maximum pooling layers) instead of a single large kernel pooling operation, thereby effectively reducing the computational requirements while maintaining the ability to capture multi-scale features. Feature merging and output: After efficient pooling, the SPPF layer connects the outputs of different pooling layers to form a fixed-length feature vector, which can be used as the input of subsequent fully connected layers or other network components.
[0048] The C3Ghost module is a specialized neural network module that integrates the advantages of the C3 module (Cross Stage PartialNetwork) and the Ghost module. When processing the input, it first passes through the convolutional layer and then through the Ghost bottleneck layer to produce additional feature maps through low-cost linear transformations. Subsequently, the original input passes through another convolutional layer and the result is concatenated with the additional feature map generated by the Ghost bottleneck. Finally, the combined output passes through another convolutional layer. It is worth noting that despite the reduced computational load, the C3Ghost module effectively utilizes the feature maps and retains key information during the conversion process, thereby maintaining high performance.
[0049] In the architecture of the LowFormer module, scalar dot product attention (SDA) is located between two depthwise convolutions (DWConv) and two pointwise convolutions (PWConv). The above depthwise convolutions and pointwise convolutions are responsible for processing the input and output projection of queries (Q), keys (K), and values (V), respectively. LowFormer has the following three main features:
[0050] 1) During the input projection stage, channel compression reduces the dimensions of the Q, K, and V channels by 50%. After SDA, the channel dimensions are restored to their original size through output projection.
[0051] 2) Lower resolution: High working resolution will greatly reduce the convolution speed. To alleviate this problem, the convolution in the penultimate stage downsamples and upsamples the resolution of the feature map before and after SDA, so that the attention operation is performed at half the resolution.
[0052] 3) Multilayer Perceptron (MLP) after Attention: Finally, a multilayer perceptron (MLP) is applied after layer normalization. Compared with the attention operation, this combination not only improves the model accuracy but also shows higher hardware efficiency
[33] .
[0053] Through the LowFormer module, the RoadObstacleDet model of the present invention can extract information from different representation subspaces at different locations. Its low resolution enables the attention operation to be performed in a more efficient way. In the case of road obstacles, it can help RoadObstacleDet effectively obtain point information from different geometric angles, thereby ultimately improving the accuracy of road obstacle detection.
[0054] After adding the inner convolution module to RoadObstacleDet, the attributes of both the spatial domain and the channel domain can be converted and incorporated into the neck layer. For different types of road obstacles, this design can enrich the receptive field information and increase the channel information of the feature map.
[0055] ShuffleAttention
[0056] Shuffle Attention (SA) divides the channel dimension into sub-features and uses shuffle units of each sub-feature to simultaneously generate channel and spatial attention. This paper introduces an attention mask for each module to suppress noise and emphasize relevant semantic areas. The present invention adds a shuffle attention module to RoadObstacleDet. The model can group the channel dimension into sub-features and apply "shuffle units" to each sub-feature to simultaneously generate channel and spatial attention. Each attention module also has an attention mask for suppressing noise and emphasizing the correct area. For road obstacle detection scenarios, this design can extract obstacle pixel information and suppress noise, thereby significantly improving the accuracy of the road obstacle detection model.
[0057] Example 2
[0058] Based on the model in Example 1, this application provides an embodiment of model training and verification, as follows:
[0059] 1. Dataset and Experimental Setup
[0060] 1.1 Dataset
[0061] To evaluate the performance of the proposed model, two custom datasets are used in this study. Due to the lack of existing public datasets on road obstacles such as collapsed edges, landslides, and collapses, we developed the Shandong Highway Obstacles Dataset and the Guizhou Highway Obstacles Dataset.
[0062] The customized Shandong highway obstacle dataset was collected and annotated in Shandong Province. The terrain of Shandong is dominated by low mountains, hills, gentle slopes and crisscrossing valleys, and mountains and hills account for a large part of the terrain. The province is famous for its well-developed highway system, with flat and wide roads, providing people with a comfortable driving experience.
[0063] Similarly, a customized Guizhou highway obstacle dataset was collected and annotated. Guizhou Province has complex terrain and diverse landforms, dominated by mountains and hills. The ubiquitous karst landforms in the region often lead to collapse, landslides and other related phenomena, posing considerable challenges to the province's highways. In addition, due to terrain restrictions, bridges and tunnels are common on Guizhou's mountainous roads.
[0064] The resolution of the highway obstacle images is 3264 × 2448 pixels and then saved to the local computer memory for further analysis and model training.
[0065] For both custom datasets, a total of about 2,000 images were collected, each of which may depict various road obstacles such as cracks, collapses, and collapses. These obstacles were carefully annotated using LabelMe, and the annotation results were saved in JSON format, including obstacle type and bounding box location. In the preprocessing stage, all images were resized to 256×256 pixels, and the corresponding bounding box locations were also adjusted proportionally. To reduce image noise, a median filter was used to replace pixels with the median value within a predefined window size.
[0066] 1.2 Dataset Expansion
[0067] To increase the number of training samples and improve model performance, we used a dataset expansion method. First, the road obstacles were segmented using annotated bounding boxes to create more diverse training data and enhance the generalization ability of the model. Then, each segmented image containing a single type of obstacle was distorted in various ways to simulate real-world noise conditions. These deformations included Gaussian blurring, resizing, adjusting brightness and contrast, and adding salt and pepper noise. Then, a copy-and-paste operation was performed. In this operation, the distorted road obstacle images were smoothly merged into the empty road images, further expanding the dataset and enhancing the robustness of the model. This process further expanded the dataset and created more realistic and diverse training samples, thereby improving the robustness of the model. This involved randomly selecting 5 to 6 segmented and distorted obstacle images and then pasting them one by one onto the empty road image, ensuring that there was no overlap and that they were at least 5 pixels away from the image boundary. This enhancement process effectively generated additional road obstacle images and enriched the training dataset.
[0068] Using the copy-paste method, we successfully generated an additional 2,000 images for the customized Shandong Highway Obstacles Dataset and Guizhou Highway Obstacles Dataset. By combining these newly generated images with the original annotated images, we doubled the number of available annotated images. Then, the images in each dataset were divided into three subsets - training set, test set, and validation set - in a ratio of 3:1:1. Table 1 provides a summary of the detailed information of the datasets.
[0069] Table 1 Two custom datasets
[0070]
[0071] 1.3 Experimental setup
[0072] To evaluate the performance of the RoadObstacleDet model, we conducted multiple experiments using a customized Shandong highway obstacle dataset and a Guizhou highway obstacle dataset. The experimental results were compared with other existing road obstacle detection models. The experiments were conducted on a computer running the Ubuntu operating system, which was equipped with an Intel Core i7-8700 CPU, an NVIDIA TM GeForce GTX 3080 GPU and 32GB memory. The RoadObstacleDet model is implemented using PyTorch, OpenCV, and Python.
[0073] The RoadObstacleDet model is trained using the back-propagation learning algorithm. During the training process, the Adam optimizer is used to adjust the weights of the model. The learning rate is set to 0.001 and the weight decay is set to 0.0005. The entire training process lasted 100 times. Like other object detection models, the loss function of the RoadObstacleDet model consists of two key parts: classification loss and regression box loss. The regression loss includes L1 loss and generalized intersection over union (GIoU) loss. This loss function achieves the best two-end match between the predicted object and the ground truth object. After the match is completed, the loss associated with the bounding box is further optimized to improve the accuracy of the model in locating road obstacles.
[0074] Assumptions is the set of ground truth objects, and is the set of N predicted objects. To find a two-way match between the ground truth set and the predicted set, we determine the permutation σ∈S of the N elements N , this combination can minimize the cost.
[0075]
[0076] in, is the pairwise matching cost between the ground truth yi and the corresponding predicted object (indexed by σ(i)). The Hungarian algorithm is used to efficiently compute the optimal assignment.
[0077] The bounding box loss Lbox is defined as follows:
[0078]
[0079] in, ar is a hyperparameter, b i is the ground truth bounding box, To predict the bounding box. Training will produce the best model in binary format. During the prediction process, load the model and input the unknown image, and the output result is the detected road obstacles and the corresponding type.
[0080] In addition, we conducted experiments to evaluate the performance of the RoadObstacleDet model compared to the current leading object detection models (Faster-RCNN, RetinaNet, YOLOv5s, YOLOv6n, YOLOv7t, YOLOv8n, YOLOv8s). Both of the aforementioned datasets were used to train and test these models.
[0081] 1.4 Evaluation Metrics
[0082] Mean Average Precision (mAP) is a commonly used metric to evaluate the performance of visual object detectors. In this study, we used three different variants of mAP, namely mAP@0.5 and mAP@0.5:0.95. In all these variants, we set an Intersection over Union (IoU) threshold to decide whether a detection is a true positive based on how much it overlaps with the ground truth. In order to strike a balance between missed detection rate and false positive rate, Average Precision (AP) is calculated as the area under the precision-recall curve. The average precision for each class is calculated separately, and the final mAP is obtained by averaging the average precision values of all classes. mAP@0.5 uses an IoU threshold of 0.5, while mAP@0.5:0.95 is calculated by averaging the mAP values from 0.5 to 0.95 IoU thresholds (with a step size of 0.05).
[0083] At the same time, the recall value is used to evaluate the proportion of positive samples that are correctly detected. The recall value R = 100% means that there are no missed targets.
[0084] 2. Results and Discussion
[0085] 2.1 Evaluation of Shandong Highway Augmented Dataset
[0086] Using the experimental setup outlined above, we evaluate the performance of the RoadObstacleDet model and compare its results with those of other existing models on the enhanced Shandong highway dataset. All models were trained on the full training dataset, while the test dataset was used to evaluate their accuracy. The results are summarized in Table 2. Table 2 shows the recognition performance of various models on the enhanced Shandong highway dataset. As shown in the table, the RoadObstacleDet model of the present invention outperforms all other models. Specifically, it has an mAP@0.5 of 88.2%, an mAP@0.5:0.95 of 52.5%, a precision of 90.3%, and a recall of 82.4%. Compared with the best performing alternative model, the RoadObstacleDet model of the present invention improves by 5.3 percentage points, 5.1 percentage points, 2.7 percentage points, and 5.5 percentage points in mAP@0.5, mAP@0.5:0.95, precision, and recall, respectively. In addition, the model achieves an FPS of up to 134.6, ensuring its applicability in real-time detection applications. The experimental results of RoadObstacleDet are further introduced below.
[0087] Figure 2 The training loss of the RoadObstacleDet model of the present invention on the customized Shandong highway dataset is shown, which is divided into box loss (box_loss), object loss (obj_loss) and class loss (cls_loss). As shown in the figure, all three losses decrease rapidly in the first epochs. Specifically, in the first few epochs, the loss drops significantly. After about 10 epochs, the rate of decline slows down and becomes more stable. However, after about 20 epochs, the loss starts to drop rapidly again. When the model runs to 300 epochs, the rate of loss decline slows down, but a significant decline can still be observed. At around 350 epochs, the loss curve flattens out, indicating that the loss has reached a minimum and the model has converged. This shows that the trained model runs effectively.
[0088] Figure 3 The precision and recall indicators of applying the RoadObstacleDet model of the present invention to the customized Shandong highway data set are shown. As the training duration increases, both the precision and recall values increase, especially after 10 epochs. The curves of precision and recall show certain ups and downs in the first 100 epochs, among which the fluctuation of precision is more obvious. Afterwards, when reaching 380 epochs, they become stable, and when reaching 400 epochs, they tend to be consistent. Finally, the precision is close to 0.90 and the recall is close to 0.82.
[0089] Figure 4 The mAP@0.5 and mAP@0.5:0.95 metrics for the RoadObstacleDet model evaluated on the customized Shandong highway dataset are shown. Both mAP@0.5 and mAP@0.5:0.95 show continued improvement as training progresses, especially after the 10th epoch. During the first 100 epochs, the mAP curves show some fluctuations, with the values intermittently increasing or decreasing. However, after this, the curves become more stable, and by about 300 epochs, the curves begin to converge. Ultimately, mAP@0.5 reaches about 0.88, while mAP@0.5:0.95 stabilizes at about 0.52, indicating strong model performance.
[0090] Figure 5 The image samples of road obstacles and the object detection results of the RoadObstacleDet model (Shandong Expressway Augmented Dataset) are shown. The images are divided into three columns, with the top row showing the original images and the bottom row showing the detection results. In the first set of images, the model effectively identified the gaps in the road surface. In the second set of images, the collapse was identified and the model's detection bounding box showed high accuracy with a confidence level greater than 95%. This demonstrates the effectiveness and accuracy of the RoadObstacleDet model in identifying and locating road obstacles. In the third set, although the collapse is tiny, it is still easy to detect. This is because the attributes derived from the spatial domain and the channel domain can be transformed and then integrated into the neck layer. This integration process can enrich the receptive field information and enhance the channel information in the feature map.
[0091] Table 2 Recognition results of different models on Shandong highway expansion dataset
[0092]
[0093] 2.2 Evaluation of Guizhou Highway Augmented Dataset
[0094] Using the above experimental configuration, we obtain the results of the RoadObstacleDet model and several existing models on the enhanced Shandong highway dataset. These models are trained using the entire training dataset, while the test dataset is used to evaluate their accuracy. The results are summarized in Table 3.
[0095] Table 3 shows the recognition performance of different models on the enhanced Shandong highway dataset. As shown in the figure, the RoadObstacleDet model achieves excellent performance. Specifically, it has an mAP@0.5 of 85.2%, an mAP@0.5:0.95 of 48.9%, a precision of 88.7%, and a recall of 77.4%. Compared with the second-best performing model, the RoadObstacleDet model improves by 4.5, 3.5, 1.8, and 4.1 in mAP@0.5, mAP@0.5:0.95, precision, and recall, respectively. It also maintains an astonishingly high FPS of 137%, which is critical for real-time detection.
[0096] Compared with the traditional YOLO model, the RoadObstacleDet model of the present invention has higher accuracy. This improvement can be attributed to the integration of LowFormer, which helps to extract information from different representation subspaces at different locations. The low-resolution property of LowFormer helps to perform attention operations more effectively. In addition, the integration of inner convolution also helps to transform and combine the attributes of the spatial domain and channel domain within the neck layer, thereby enhancing the receptive field and enriching the channel information of the feature map. In addition, after adding "ShuffleAttention" to RoadObstacleDet, the model can group the channel dimension into sub-features, suppress noise, and emphasize the correct areas. These modifications together improve the detection accuracy of the RoadObstacleDet model.
[0097] Figure 6 The training loss of the RoadObstacleDet model on the customized Guizhou highway dataset is depicted, including the box loss (box_loss), object loss (obj_loss), and class loss (cls_loss). As can be seen from the figure, during the initial epochs, the box loss, object loss, and class loss decrease rapidly. After about 10 epochs, the rate of decrease stabilizes. After more than 10 epochs, the loss decreases rapidly. After 300 epochs, although the loss decreases not as fast, a significant decrease can still be seen. After about 300 epochs, the loss stabilizes and the loss curve flattens, which indicates that the loss has reached a minimum and the trained model shows good performance.
[0098] Figure 7The precision and recall metrics of the RoadObstacleDet model on the customized Guizhou highway dataset are shown. As the training time increases, the precision and recall values also increase, especially after 10 epochs. The precision and recall curves fluctuate within the first 100 epochs, with the fluctuation of precision being more obvious. Subsequently, they stabilize around 380 epochs and converge at 400 epochs. Finally, the precision is close to 0.88 and the recall is close to 0.77.
[0099] Figure 8 The mAP@0.5 and mAP@0.5:0.95 indicators of the RoadObstacleDet model on the customized Guizhou highway dataset are shown. As the training duration increases, the values of mAP@0.5 and mAP@0.5:0.95 also increase, especially after 10 epochs. The mAP curve fluctuates within the first 100 epochs and enters a stable and convergent stage after 300 epochs. In the end, mAP@0.5 reaches about 0.85, while mAP@0.5:0.95 reaches about 0.48.
[0100] Fig. 9 Example images of road obstacles are shown, along with the corresponding object detections using the RoadObstacleDet model. The images are divided into three columns, each representing a different set of images. The original images are shown in the top row, and the corresponding detections are shown in the bottom row. The first set of images clearly shows the cracking, while the second set of images highlights the collapse, with accurate bounding box localization and confidence levels above 95%. In the third set, the collapse is subtle but its detection is still clear. This is due to the integration of spatial and channel domain attributes into the neck layer, which enriches the receptive field and enhances the channel information of the feature map.
[0101] Table 3 Recognition results of different models on Guizhou highway expansion dataset
[0102]
[0103]
[0104] Data availability
[0105] The data used to support the findings of this study (https: / / github.com / wanhaifengytu / WiseRoad / tree / main / data) have been deposited in a GitHub repository (https: / / github.com / wanhaifengytu / WiseRoad).
[0106] It should be noted that, in the absence of conflict, the features in the embodiments of this application may be combined with each other.
[0107] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A RoadObstacleDet model for road obstacle detection, characterized in that: The model includes: a backbone network module for receiving an image to be detected and extracting features from the image, a neck module for receiving features output by the backbone network module and integrating these features, and a prediction head module for receiving output by the neck module and responsible for generating detection results; In the backbone network module, the input image of size (3, 256, 256) first passes through the CBS module to extract the activation feature map with lower resolution and obtain A0 of size (32, 128, 128). A0 is then processed by a series of convolutional layers + C3Ghost modules to generate feature maps of different sizes.
2. The RoadObstacleDet model for road obstacle detection according to claim 1, characterized in that: In the backbone network module, there are four convolutional layers + C3Ghost modules connected together; Among them: the CBS module outputs A0 of size (32, 128, 128) through the first convolutional layer + C3Ghost module, and outputs A2 of size (64, 64, 64); A2 enters the second convolutional layer + C3Ghost module, generating an output A of size (128, 32, 32); the third convolutional layer + C3Ghost module converts A into B of size (256, 16, 16); the fourth convolutional layer + C3Ghost module converts B into C of size (512, 8, 8); subsequently, A, B and C are used as inputs of the neck module respectively.
3. The RoadObstacleDet model for road obstacle detection according to claim 2, characterized in that: In the neck module, C is transmitted to the SPPF module to generate D of size (512, 8, 8); then, D is transmitted to the GhostConv module to generate E of the same size; subsequently, E is transmitted to the upsampling layer module to obtain F of size (512, 16, 16); then, F is processed by the convolutional layer module to obtain G; At the same time, B and G are element-wise added to obtain H of size (256, 16, 16); H is passed to the C3Ghost module to obtain J of the same dimension; then, J is sent to the convolutional layer module and then passed through the upsampling layer module to obtain L of size (128, 32, 32); L is processed by the inner convolution module to obtain M of the same dimension; At the same time, A is transferred to the LowFormer module to obtain I of the same size as A; I and M are added vector by vector to obtain N.
4. The RoadObstacleDet model for road obstacle detection according to claim 3, characterized in that: In the neck module, Use N and get O through the C3Ghost module, then O is processed by the shuffle attention module to get R of size (128, 32, 32), and finally, R is sent to detector 1 in the detector module; O also passes through an inner convolution module to obtain U; at the same time, J first passes through a LowFormer module to obtain P, whose size is kept as (256, 16, 16); P and U are added element by element to obtain V, which is then processed by the shuffle attention module to obtain S of size (256, 16, 16), and finally S is sent to detector 2 in the detector module; O also passes through an inner convolution module to obtain W; at the same time, E is used and enters the LowFormer module to obtain Q with a size of (512, 8, 8); Q and W are added element by element to obtain X, and the size of X remains unchanged (512, 8, 8); then, X is processed by the shuffle attention module to obtain T with a size of (512, 8, 8), and finally T is sent to the detector 3 in the detector module.
5. A method for constructing the model according to any one of claims 1 to 4, characterized in that: During training, the dataset expansion method was used to increase the number of training samples and improve model performance.
6. The construction method according to claim 5, characterized in that: The dataset expansion method is as follows: First, road obstacles are segmented using annotated bounding boxes to create more diverse training data and enhance the generalization ability of the model; Each segmented image containing a single type of obstacle is then distorted in various ways to simulate real-world noise conditions; these include Gaussian blurring, resizing, adjusting brightness and contrast, and / or adding salt and pepper noise. The processed images are then copied and pasted, specifically merged into the empty road images, further expanding the dataset.
7. A method for detecting road obstacles, characterized in that: A model constructed using the model described in any one of claims 1 to 4 or the method of claim 5 or 6.
8. A server comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for detecting road obstacles according to claim 7 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting road obstacles according to claim 7 are implemented.
Citation Information
Cited By
Obstacle detection method, device, equipment, medium and product
CN120783319A
Method and system for rapidly detecting abnormal articles on expressway
CN121259766A