Method for detecting thrown objects on expressway based on small target image recognition

By adopting a small target sprinkler detection model based on YoloV5s in the highway sprinkler detection, combining data enhancement, CA coordination attention mechanism and Varifocal Loss loss function, the problems of difficulty in extracting sprinkler characteristics and low detection accuracy are solved, and the accurate identification and efficient detection of small target sprinklers on the highway are achieved.

CN120125875APending Publication Date: 2025-06-10ZHEJIANG TRANSPORTATION GROUP TECHNICAL RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510119808.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the detection of highway sprinklers, the problems of difficulty in extracting sprinklers, low detection accuracy, and excessive dependence of deep learning models on data.

Method used

The small-target sprinkler detection model based on YoloV5s is adopted, and the model is optimized to improve the detection accuracy of small-target sprinklers through data enhancement, CA coordination attention mechanism and Varifocal Loss loss function.

Benefits of technology

It realizes accurate identification of small target spills on highways, reduces detection errors, solves data dependence problems, and improves detection efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125875A_ABST
    Figure CN120125875A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent traffic and computer vision, in particular to an expressway thrown object detection method based on small target image recognition, and aims at solving the problems that in the prior art, thrown object features are difficult to extract, the detection precision is low, and a deep learning model excessively depends on data. According to the technical scheme, the expressway throwing object detection method based on small target image recognition comprises the following steps of S1, collecting a small target throwing image of an expressway; s2, marking the position and the type of a small target throwing object on the image; s3, performing data enhancement on the image and dividing the image into a training set, a verification set and a test set; s4, establishing a small target thrown object detection model based on YoloV5s; s5, training the optimization model by using the training set and the verification set, and evaluating the model by using the test set; and S6, deploying the model on site, and detecting the small target spilled objects on the expressway in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation and computer vision technology, and in particular to a method for detecting spilled objects on a highway based on small target image recognition. Background Art

[0002] With the expansion of highway construction and the increase in traffic volume, accidents caused by spilled objects have become a major safety hazard. There are many types of objects spilled on highways, ranging from small objects such as nails and plastic bags to large objects such as tires and steel coils, and most of these objects are caused by cargo or parts falling from trucks. Take Zhejiang Province as an example. The province's more than 5,000 kilometers of highways require a large number of personnel and cleaning vehicles every day to clean up tons of spilled objects. If these spilled objects can be detected and cleaned up in time, the probability of traffic accidents can undoubtedly be significantly reduced.

[0003] Deep learning object detection is an effective method for identifying spilled objects on highways, showing great potential, but its greedy demand for data has become a major bottleneck. Building an accurate and generalizable model often requires a large amount of labeled data to support the training process. Although researchers continue to explore innovative paths such as transfer learning and small sample learning to try to alleviate the dilemma of data scarcity, for now, when these methods are applied to highway spilled object detection scenarios, the detection effect is still unsatisfactory, and there is still a large gap from the standard that can be stably and reliably put into practical use. The adaptability of transfer learning between different scenarios and the risk of model overfitting faced by small sample learning all restrict the improvement of detection accuracy and reliability.

[0004] Currently, YOLOV5s, as an efficient object detection algorithm, has attracted widespread attention due to its low computational overhead and high detection accuracy. However, the scattered objects on the highway are usually small, and traditional object detection algorithms often have difficulty in identifying these small objects. Therefore, it is urgent to develop a new method for detecting scattered objects, focusing on the direction of "small target image recognition". Summary of the invention

[0005] The purpose of the present invention is to overcome the deficiencies in the above-mentioned background technology and provide a highway spilled object detection method based on small target image recognition to solve the problems of difficulty in extracting features of spilled objects, low detection accuracy and excessive reliance of deep learning models on data.

[0006] The technical solution of the present invention is:

[0007] A method for detecting spilled objects on a highway based on small target image recognition comprises the following steps:

[0008] Step S1: Data collection

[0009] Collect images of small target spills on highways;

[0010] Step S2: Data annotation

[0011] Annotate the positions and types of small target spills on the images;

[0012] Step S3: Data processing

[0013] Perform data augmentation on the images and divide them into a training set, a validation set, and a test set;

[0014] Step S4: Model establishment

[0015] Establish a small target spill detection model based on YoloV5s;

[0016] Step S5: Training and evaluation

[0017] Use the training set and the validation set to train and optimize the model, and use the test set to evaluate the model;

[0018] Step S6: Field deployment

[0019] Deploy the model in the field to detect highway spills in real time.

[0020] In the above step 2, the types of small target spills include metal, wood, tire, cloth, stone, plastic, hazardous substances, roadblocks, oil stains, animals, and others.

[0021] In the above step 3, data augmentation includes random cropping, flipping, rotation, and adjustment of brightness and contrast.

[0022] In the above step 3, the ratio of the training set, the validation set, and the test set is 3:1:1.

[0023] In the above step 4, the small target spill detection model includes an input end, a Backbone network, a Neck network, and a detect network.

[0024] In the above step 4, the Backbone network includes a Focus layer, a first convolutional layer, a first C3 residual convolutional layer, a second convolutional layer, a convolutional expansion layer, a second C3 residual convolutional layer, a third convolutional layer, a third C3 residual convolutional layer, a fourth convolutional layer, a spatial pyramid pooling layer, and a fourth C3 residual structure convolutional layer connected in sequence;

[0025] The Neck network includes a fifth convolutional layer, a first upsampling layer, a first splicing layer, a fifth C3 residual convolutional layer, a sixth convolutional layer, a second upsampling layer, a second splicing layer, a sixth C3 residual structure convolutional layer, a first CA channel attention pooling layer, a seventh convolutional layer, a third splicing layer, a seventh C3 residual structure convolutional layer, a second CA channel attention pooling layer, an eighth convolutional layer, a fourth splicing layer, an eighth C3 residual structure convolutional layer, and a third CA channel attention pooling layer, which are connected in sequence;

[0026] The Detect network classifies and predicts boundaries of targets in the feature map based on a candidate box of a preset size. The Detect network includes a first detection layer, a second detection layer, and a third detection layer.

[0027] In step 4,

[0028] The output of the second C3 residual convolution layer is connected to the input of the second concatenation layer; the output of the third C3 residual convolution layer is connected to the input of the first concatenation layer; the output of the fourth C3 residual convolution layer is connected to the input of the fifth convolution layer;

[0029] The output of the fifth convolutional layer is connected to the input of the fourth splicing layer; the output of the sixth convolutional layer is connected to the input of the third splicing layer;

[0030] The output of the first CA channel attention pooling layer is connected to the input of the first detection layer; the output of the second CA channel attention pooling layer is connected to the input of the second detection layer; the output of the third CA channel attention pooling layer is connected to the input of the third detection layer.

[0031] In step 5, Varifocal Loss is used as the loss function in the training process.

[0032]

[0033] p is the predicted IoU-aware classification score and q is the objectness score.

[0034] The beneficial effects of the present invention are:

[0035] The small target scattered object detection model of the present invention uses cutting-edge deep learning and optimized YoloV5s algorithm. Through the convolution kernel, it deeply explores the details of small scattered objects and strengthens the small target learning with a special loss function, accurately identifies various types of small target scattered objects, and reduces detection errors. At the same time, the data set covers high-speed multi-scenes and is professionally processed and enhanced to perfectly integrate the small target scattered object detection model with the existing monitoring system, solving the problems of difficulty in extracting scattered object features, low detection accuracy and excessive reliance on data by deep learning models in the prior art, realizing real-time monitoring, rapid alarm and feedback, effectively improving the safety and traffic efficiency of highways and ensuring smooth traffic. Brief Description of the Drawings

[0036] Figure 1 is the flow chart of the present invention.

[0037] Figure 2 is the architecture diagram of the small target litter detection model of the present invention.

[0038] Figure 3 is the architecture diagram of the CA channel attention pooling layer of the present invention.

[0039] Figure 4 is the schematic diagram of data annotation of the present invention.

[0040] Figure 5 is the recognition effect diagram of the small target litter detection model of the present invention for the plastic objects left on the highway.

[0041] Figure 6 is the recognition effect diagram of the small target litter detection model of the present invention for the cloth objects left on the highway.

[0042] Figure 7 is the recognition effect diagram of the small target litter detection model of the present invention for other objects left on the highway. Detailed Description of the Invention

[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.

[0044] As Figure 1 shown, a highway litter detection method based on small target image recognition includes the following steps:

[0045] Step S1: Data collection

[0046] Collect small target litter images on the highway.

[0047] Install high-definition cameras at various key sections of the highway and locations where small target litter is likely to appear.

[0048] Use the high-definition cameras to collect small target litter images on the highway.

[0049] The following factors need to be considered during collection: various meteorological conditions under different weather conditions, different time periods (covering early morning, dawn, morning, noon, afternoon, evening and night), different traffic flows, and different lighting conditions.

[0050] To ensure the collection of a rich variety of small target spill images, it is also necessary to collect a large number of normal road background images without spills. The small target spill images and normal road background images are integrated according to a ratio of 9:1 to construct a dataset that comprehensively reflects the actual scenario.

[0051] Step S2: Data annotation

[0052] Annotate the positions and types of small target spills on the images.

[0053] As Figure 4 shown, for each image containing small target spills, experienced annotators use professional annotation tools to frame the accurate positions of the small target spills in the image and mark the types of small target spills according to the pre-set category system. The annotation information will serve as a supervision signal during model training to guide the model to learn the visual feature patterns of small target spills and their position distribution rules in the image.

[0054] The types of small target spills include metal, wood, tire, fabric, stone, plastic, hazardous substances, roadblocks, oil stains, animals, and others.

[0055] Step S3: Data processing

[0056] Perform data augmentation on the images and divide them into a training set, a validation set, and a test set;

[0057] Data augmentation includes random cropping, flipping, rotation, and brightness and contrast adjustment.

[0058] Random cropping: Randomly select local areas containing small target spills from the image for cropping to expand sample diversity, while ensuring the integrity of the small targets after cropping and the accuracy of the annotation information.

[0059] Flipping: Perform a flipping operation on the image along the horizontal or vertical direction to simulate the shapes of small target spills from different perspectives and increase the model's adaptability to the orientation changes of small target spills.

[0060] Rotation: Rotate the image within a certain angle range so that the model can learn the features of rotated small target spills in non-standard postures and improve robustness.

[0061] Brightness and contrast adjustment: Simulate different light intensity scenarios and randomly change the brightness and contrast of the image to ensure that the model can accurately identify small target spills under complex lighting conditions.

[0062] Divide the images of each category into a training set, a validation set, and a test set according to a ratio of 3:1:1.

[0063] Collect data of Hangzhou Ring Expressway from February 1st to December 20th, 2023, a total of 16,218 images, including 14,597 images with small target spillage and 1,621 normal road surface background images without spillage.

[0064] Data augmentation: Horizontally flip the images; Rotate 180 degrees clockwise; Randomly adjust the brightness between [0.5, 1.5]; Randomly adjust the contrast between [0.5, 1.5].

[0065] Use 9,730 of them as the training set, 3,244 as the validation set, and 3,244 as the test set.

[0066] Step S4: Establish a model

[0067] Establish a small target spillage detection model based on YoloV5s.

[0068] The small target spillage detection model uses object detection technology based on deep learning. Select the YoloV5s algorithm as the basis and improve it according to the requirements of small target detection of highway spillage. Add 1 3*3 convolutional kernel on the original network to improve the detection granularity of small targets, so as to cope with the problem of feature extraction of small target spillage with a size less than 30*30 cm, and introduce the CA coordination attention mechanism to enable the network to better capture the positional relationship of the target in space.

[0069] The small target spillage detection model specializes in detecting small targets with an actual size less than 30*30 cm.

[0070] As Figure 2 shown, the small target spillage detection model includes four parts: the input end, the Backbone network (main network), the Neck network (neck network), and the detect network (detection network).

[0071] The input end performs preprocessing operations on the highway images after processing the collected data, uniformly scales the images to the size required by the model, which is 640x640, and normalizes the pixel values so that their range meets the requirements of model training, laying a foundation for subsequent feature extraction.

[0072] The Backbone network includes a Focus layer, a first convolutional layer, a first C3 residual convolutional layer, a second convolutional layer, a convolutional expansion layer, a second C3 residual convolutional layer, a third convolutional layer, a third C3 residual convolutional layer, a fourth convolutional layer, a spatial pyramid pooling layer, and a fourth C3 residual structure convolutional layer connected in sequence.

[0073] The Neck network includes a fifth convolutional layer, a first upsampling layer, a first splicing layer, a fifth C3 residual convolutional layer, a sixth convolutional layer, a second upsampling layer, a second splicing layer, a sixth C3 residual structure convolutional layer, a first CA channel attention pooling layer, a seventh convolutional layer, a third splicing layer, a seventh C3 residual structure convolutional layer, a second CA channel attention pooling layer, an eighth convolutional layer, a fourth splicing layer, an eighth C3 residual structure convolutional layer, and a third CA channel attention pooling layer, which are connected in sequence.

[0074] The Detect network classifies the targets in the feature map and predicts the boundaries based on candidate boxes of a preset size. The Detect network includes a first detection layer, a second detection layer, and a third detection layer.

[0075] The output of the second C3 residual convolutional layer is connected to the input of the second splicing layer. The output of the third C3 residual convolutional layer is connected to the input of the first splicing layer. The output of the fourth C3 residual convolutional layer is connected to the input of the fifth convolutional layer.

[0076] The output of the fifth convolutional layer is connected to the input of the fourth splicing layer. The output of the sixth convolutional layer is connected to the input of the third splicing layer.

[0077] The output of the first CA channel attention pooling layer is connected to the input of the first detection layer. The output of the second CA channel attention pooling layer is connected to the input of the second detection layer. The output of the third CA channel attention pooling layer is connected to the input of the third detection layer.

[0078] The small target spill detection model has the following advantages:

[0079] 1. Receptive field adaptation: Reasonably stack 1 convolutional expansion layer, that is, a 3*3 convolutional kernel, at the key positions of the Backbone network. According to the actual size distribution of the spills on the highway, adjust the size of the model's receptive field to ensure sufficient extraction of the feature information of small targets, realize the enhancement of feature extraction, and cope with the problem of feature extraction of spills with a size less than 30*30 cm.

[0080] 2. Introduce the CA coordinated attention mechanism into the Neck network ( Figure 3 as shown in the existing methods), which helps to process and utilize the feature information more finely. Embed the position information of the image into the CA coordinated attention mechanism, so that the network can better capture the position relationship of the targets in space. Especially when dealing with the highway environment with strong spatial features, it can act finely on the area where the small targets are located, focus on the detailed features of small spills such as the sharpness of the edge contour and the uniqueness of the surface texture, make the small targets more distinguishable in the feature space, and at the same time avoid introducing too much computational burden.

[0081] 3. The backbone network is composed of multiple convolutional layers and a pooling layer arranged in an orderly manner. The convolutional layer selects a lightweight convolutional module (3*3) that is sensitive to small targets, which can effectively extract the deep features of the image while reducing the occupancy of computing resources, and initially capture the regional features where the spillage may exist.

[0082] 4. The Neck network fuses feature maps of different scales through upsampling and lateral connections, combines the deep semantic features output by the Backbone network with the shallow detail features, enabling the model to grasp both the overall scene information of the image and accurately focus on the subtle features of small target spillages, enhancing the expressiveness of small targets at the feature level.

[0083] Step S5: Training and evaluation

[0084] Use the training set and validation set to train and optimize the model, and use the test set to evaluate the model.

[0085] S5.1. Fine-tuning of model training:

[0086] Loss function customization: Use Varifocal Loss as the loss function during the training process. Since small targets have a small pixel proportion and blurred features in the image, through this weighted method, the model is made to pay more attention to the detection errors of small targets during training, prioritize focusing on learning the unique feature representations of small targets, reduce the missed detection and false detection rates of small targets, and then perform hyperparameter optimization by comprehensively considering the learning rate, batch processing size, and number of iterations.

[0087] The Varifocal Loss function is as follows:

[0088]

[0089] Where: p is the predicted IoU-aware classification score (IACS), and q is the target score.

[0090] During the training process, for positive samples, q is set to the IoU value between the predicted box and the ground truth box (gt box); for negative samples, the value of q is zero, which applies to all classes. In this way, the training process will focus more on candidate detection samples with higher IACS values. α is an adjustment ratio factor used to control the weights of positive and negative sample losses, and its value ranges from 0 to 1. By reasonably setting α, the excessive attention to negative samples can be reduced. At the same time, the value of γ is usually set greater than 1 to increase the loss weight of difficult samples.

[0091] Since small targets occupy a small proportion of pixels in the image and have blurred features, through this weighting method, the model pays more attention to the detection errors of small targets during training, focuses on learning the unique feature representations of small targets first, reduces the missed detection and false detection rates of small targets, and then performs hyperparameter optimization by comprehensively considering the learning rate, batch size, and number of iterations.

[0092] The learning rate refers to finding the learning rate value that enables the model to converge quickly and not get stuck in local optima during training by starting from an initially large learning rate and gradually decreasing it in an exponential or stepwise decay manner through multiple rounds of experiments.

[0093] The batch size refers to trying different batch sizes based on the hardware computing resources, balancing the computational amount per iteration and the stability of model parameter updates, and finding the optimal number of batch samples for model training.

[0094] The number of iterations refers to gradually increasing the number of iterations, observing the changes in the performance metrics of the model on the validation set, avoiding overfitting, and determining a suitable iteration upper limit that can fully learn the data features without causing model performance degradation.

[0095] This patent involves setting α to 0.8 and γ to 2.5 in the model for small target detection in complex scenarios, with a learning rate of 0.001, a batch size of 32, and 150 rounds of iterations.

[0096] S5.2, Model Evaluation: Establishment of the Evaluation Index System:

[0097] The mean average precision (mAP) is used to measure the overall detection accuracy of the model, reflecting the comprehensive accuracy of the model in detecting different types of spills.

[0098] The recall rate is used to evaluate the ability of the model to find all real spills, that is, the proportion of actual existing spills detected by the model, focusing on the recall of small target spills to ensure that key small targets are not missed.

[0099] The precision rate reflects the accuracy of the model's detection results, that is, the proportion of samples detected as spills that are truly spills, preventing false alarms from interfering with actual operations.

[0100] During the evaluation process, analyze each index corresponding to small target spills separately to deeply understand the detection effect of the model on small target spills. Run the trained model on a test set independent of the training set and validation set, and use the model after 150 epochs for test verification. The final actual test results (shown in Table 1) are:

[0101] Model Data volume Accuracy rate Recall rate Map Time consumption per single image Environment YoloV5s 2218 0.683 0.659 0.66 0.012s Single card The present invention 2218 0.736 0.721 0.71 0.012s Single card

[0102] Table 1

[0103] The experimental results in Table 1 show that under the condition that 2,218 spillage pictures obtained in the actual scenario are used as the test data set, the accuracy rate of the small target spillage detection model of the present invention reaches 0.736, the recall rate reaches 0.721, and the mAP reaches 0.71, all exceeding the original YoloV5s model, while the average time consumption is the same. That is, without reducing the detection efficiency (the average time consumption is the same), it has a better detection effect than the original model.

[0104] Step S6: On-site deployment

[0105] Deploy the model on-site to detect highway spillage in real time.

[0106] Integrate the optimized model into the intelligent monitoring system already deployed along the highway to ensure seamless docking of the model with hardware facilities such as camera image acquisition devices, data transmission links, and monitoring center servers. Input the high-definition video image stream collected by the camera into the model in real time, and the model quickly analyzes each frame of the image.

[0107] Once a spillage is identified (the identification results are as shown in Figure 5 , Figure 6 , Figure 7 ), immediately trigger an alarm signal. The deployed intelligent monitoring system packages key data such as the location information of the spillage and the type judgment result and sends it to the monitoring center. The staff in the monitoring center, based on the information fed back by the model, quickly dispatch road administration cleaning vehicles to the location of the spillage for timely cleaning to ensure the safe and smooth traffic on the highway.

[0108] The preferred embodiments of the present invention are given in the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in this specification. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure content of the present invention more thorough and comprehensive.

Claims

1. A method for detecting spilled objects on a highway based on small target image recognition, comprising the following steps: Step S1: Data collection Collect images of small target scattering on highways; Step S2: Data labeling Mark the location and type of small target spills on the image; Step S3: Data processing Perform data augmentation on the images and divide them into training set, validation set and test set; Step S4: Building the model Establish a small target spill detection model based on YoloV5s; Step S5: Training evaluation Use the training set and validation set to train and optimize the model, and use the test set to evaluate the model; Step S6: Field deployment The model is deployed in the field to detect spilled objects on highways in real time.

2. A highway spilled object detection method based on small target image recognition according to claim 1, characterized in that: In step 2, the types of small target spilled objects include metal, wood, tires, cloth, stones, plastics, hazardous chemicals, roadblocks, oil, animals and others.

3. The method for detecting spilled objects on highways based on small target image recognition according to claim 1, characterized in that: In step 3, data enhancement includes random cropping, flipping, rotation, and brightness contrast adjustment.

4. The method for detecting spilled objects on highways based on small target image recognition according to claim 1, characterized in that: In step 3, the ratio of the training set, the validation set, and the test set is 3:1:

1.

5. The method for detecting spilled objects on highways based on small target image recognition according to claim 1, characterized in that: In step 4, the small target spilled object detection model includes an input end, a Backbone network, a Neck network and a detect network.

6. A highway spilled object detection method based on small target image recognition according to claim 5, characterized in that: In step 4, the Backbone network includes a Focus layer, a first convolutional layer, a first C3 residual convolutional layer, a second convolutional layer, a convolutional expansion layer, a second C3 residual convolutional layer, a third convolutional layer, a third C3 residual convolutional layer, a fourth convolutional layer, a spatial pyramid pooling layer, and a fourth C3 residual structure convolutional layer connected in sequence; The Neck network includes a fifth convolutional layer, a first upsampling layer, a first splicing layer, a fifth C3 residual convolutional layer, a sixth convolutional layer, a second upsampling layer, a second splicing layer, a sixth C3 residual structure convolutional layer, a first CA channel attention pooling layer, a seventh convolutional layer, a third splicing layer, a seventh C3 residual structure convolutional layer, a second CA channel attention pooling layer, an eighth convolutional layer, a fourth splicing layer, an eighth C3 residual structure convolutional layer, and a third CA channel attention pooling layer, which are connected in sequence; The Detect network classifies and predicts boundaries of targets in feature maps based on candidate boxes of preset sizes. The Detect network includes a first detection layer, a second detection layer, and a third detection layer.

7. A highway spilled object detection method based on small target image recognition according to claim 6, characterized in that: In step 4, The output of the second C3 residual convolution layer is connected to the input of the second concatenation layer; the output of the third C3 residual convolution layer is connected to the input of the first concatenation layer; the output of the fourth C3 residual convolution layer is connected to the input of the fifth convolution layer; The output of the fifth convolutional layer is connected to the input of the fourth splicing layer; the output of the sixth convolutional layer is connected to the input of the third splicing layer; The output of the first CA channel attention pooling layer is connected to the input of the first detection layer; the output of the second CA channel attention pooling layer is connected to the input of the second detection layer; the output of the third CA channel attention pooling layer is connected to the input of the third detection layer.

8. The method for detecting spilled objects on highways based on small target image recognition according to claim 1, characterized in that: In step 5, Varifocal Loss is used as the loss function in the training process. p is the predicted IoU-aware classification score and q is the objectness score.

Citation Information

Cited By

  • Road surface scattering detection method and device based on deep learning, electronic equipment and program product

    CN121121507A

  • Method and device for detecting road surface spilling based on deep learning, electronic equipment and program product

    CN121121507B