Road detection method and system based on deep learning
By constructing a deep learning-based road detection method, using SIoU loss function and multi-scale feature pyramid structure, the problems of low road detection accuracy and poor adaptability in the existing technology are solved, and high-precision and efficient road detection effects are achieved.
Patent Information
- Application Number
- CN202411946780.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
The existing road detection methods have low accuracy and are difficult to adapt to different road environments and complex scenarios. The effect needs to be improved especially when dealing with variable weather conditions, shading and narrow or irregular roads.
The road detection method based on deep learning is adopted to build a convolutional neural network model and optimize the model using SIoU loss function, combining the multi-scale feature pyramid structure in the feature extraction and object detection stages to improve the detection accuracy and efficiency of the model.
The detection accuracy and efficiency are significantly improved, especially when dealing with different scales and shape targets, the model shows high detection accuracy and stability, achieving efficient implementation of the technology and value transformation.
Smart Images

Figure CN119992470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a road detection method and system based on deep learning. Background Art
[0002] Road detection methods and systems based on deep learning refer to a technical system that uses deep learning technology to accurately identify and segment road areas from road images or videos. With the rapid development of deep learning, road detection based on deep learning has made significant progress in both accuracy and efficiency. Deep learning models can automatically learn and extract complex features in road images, thereby greatly improving the accuracy of detection and segmentation. Traditional methods usually rely on manually designed features and rules to identify roads. Although the calculation speed is fast, the accuracy is low and it is difficult to adapt to different road environments and complex scenes. Deep learning algorithms use deep neural networks to learn road images end-to-end. They have the advantages of high accuracy and strong generalization ability and have become the mainstream method for road detection. However, due to the complexity of actual road scenes, existing methods still face certain difficulties in dealing with challenges such as changing weather conditions and obstructions, especially when dealing with narrow or irregular roads. Summary of the invention
[0003] In view of the above-mentioned problems, the present invention is proposed.
[0004] Therefore, the technical problem solved by the present invention is that the existing road detection methods have low accuracy and are difficult to adapt to different road environments and complex scenes. They still face certain difficulties when dealing with challenges such as changeable weather conditions and obstructions, and the effect of dealing with narrow or irregular roads needs to be improved.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: a road detection method based on deep learning, comprising collecting and processing image data of roads; constructing a deep learning model for road detection based on a convolutional neural network model of deep learning; optimizing the model by calculating the difference between prediction and actual results using the SIoU loss function; and training and adjusting the collected road data set through the network model.
[0006] As a preferred solution of the road detection method based on deep learning described in the present invention, a vehicle equipped with a camera sensor drives on the road to collect image data of the road.
[0007] Obtain road image data from various real-life scenarios. The collected data should cover different weather, time, road types, and traffic conditions.
[0008] As a preferred solution of the road detection method based on deep learning described in the present invention, wherein: a data set is prepared, the collected data is processed, and the image is annotated using a tool to prepare the data set;
[0009] The collected raw image data needs to be preprocessed, including image cropping, scaling, and color adjustment operations;
[0010] Remove noise from the image, standardize the image, and mark the road and non-road areas in the image. The marking result is a pixel-level mask corresponding to each image. Each pixel is assigned a label value according to its category to distinguish the road from the non-road area. After the marking is completed, the data set is sorted and divided into training set, validation set and test set, with 80% of the data used for training, 10% of the data used for validation, and 10% of the data used for testing.
[0011] As a preferred solution of the road detection method based on deep learning described in the present invention, wherein: based on the convolutional neural network model of deep learning, a deep learning model for road detection is constructed, and the model is divided into two stages: feature extraction and detection;
[0012] Based on the YOLOv4 structure, the backbone network and neck network of the network model are constructed to extract features. The backbone network is composed of CSPDarkNet, which introduces a cross-stage partial connection mechanism.
[0013] By applying the pooling operation on the feature map, pooled features are generated and the pooled features are concatenated to form a feature vector or feature map;
[0014] Pooling operations are performed on the input feature map using pooling windows of different sizes. The feature maps generated by pooling capture the features of different levels and regions in the input feature map. All pooled feature maps of different scales are concatenated through the channel dimension to generate a feature map containing multi-scale information.
[0015] As a preferred solution of the road detection method based on deep learning described in the present invention, the neck network is composed of a Feature Pyramid Network, and the FPN uses feature maps of different levels in the convolutional neural network to construct a pyramid structure containing multi-scale features through top-down upsampling and lateral connections;
[0016] The head network of the constructed network is used for detection. The head network consists of three layers of upsampling networks. The feature map of the last layer after feature extraction is enlarged. After the feature map is enlarged to the original image size, the feature map passes through the fully connected layer to output the segmented road information in the image.
[0017] As a preferred solution of the road detection method based on deep learning described in the present invention, wherein: the difference between the prediction and the actual result is calculated using the SIoU loss function to adjust the optimization model during training;
[0018] Introducing scale invariance, SIoU evaluates the degree of overlap between the predicted box and the true box. SIOU is expressed as:
[0019]
[0020]
[0021]
[0022] Among them, Δ represents distance loss, Ω represents shape loss, and IoU represents intersection over union loss.
[0023] As a preferred solution of the road detection method based on deep learning described in the present invention, wherein: the network model is trained with the collected road data set and continuously adjusted;
[0024] Use the built network model to systematically train the collected diverse road data sets;
[0025] Continuously adjust and optimize network parameters through repeated iterations and verification.
[0026] Another object of the present invention is to provide a road detection system based on deep learning, which can solve the problem of low accuracy of current road detection methods by calculating the difference optimization model between prediction and actual results using the SIoU loss function.
[0027] As a preferred solution of the deep learning-based road detection system described in the present invention, it includes: an image data acquisition and annotation module, a deep learning model construction and training module, and a model training and optimization module; the image data acquisition and annotation module is used to collect road image data, process the collected data, and use annotation tools to create a data set; the deep learning model construction and training module is used to build a road detection model including feature extraction and target detection stages based on a deep learning convolutional neural network model, and use the SIoU loss function to optimize the model performance; the model training and optimization module is used to use the prepared data set and the constructed network model for training, and improve the detection accuracy of the model through repeated optimization.
[0028] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a road detection method based on deep learning.
[0029] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a road detection method based on deep learning.
[0030] Beneficial effects of the present invention: The deep learning-based road detection method provided by the present invention not only effectively reduces the computational burden, but also significantly improves the detection accuracy and efficiency. In particular, the introduced SIoU loss function further enhances the model's ability to detect targets of different scales and shapes by more accurately evaluating the degree of overlap between the predicted box and the true box. Ultimately, the model, which has been systematically trained and continuously optimized, demonstrated excellent road detection performance in actual scenarios, achieving efficient implementation of technology and value conversion. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0032] Figure 1 A flowchart for collecting and producing data sets for a road detection method based on deep learning provided in the first embodiment of the present invention.
[0033] Figure 2 A flowchart for constructing a road detection method based on deep learning provided for the first embodiment of the present invention.
[0034] Figure 3 A network model structure diagram of a road detection method based on deep learning provided for the first embodiment of the present invention.
[0035] Figure 4 A schematic diagram of detection results of a road detection method based on deep learning provided in the first embodiment of the present invention.
[0036] Figure 5 An overall flow chart of a road detection system based on deep learning provided for the third embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0038] Example 1, reference Figure 1 - Figure 4 , is an embodiment of the present invention, and provides a road detection method based on deep learning, comprising:
[0039] S1: Collect and process road image data.
[0040] Furthermore, a vehicle equipped with a camera sensor is driven on the road to collect image data of the road. Image data of the road is obtained from various real scenes, which usually comes from the sensor of the vehicle camera. The collected data should cover different weather, time, road type and traffic conditions to ensure the generalization ability of the model.
[0041] It should be noted that the dataset is made, the collected data is processed, and the image is annotated using tools to make the dataset. The collected raw image data needs to be preprocessed, including image cropping, scaling, color adjustment and other operations to ensure the consistency and quality of the data. In addition, it is necessary to remove the noise in the image and standardize the image. In order to semantically segment and annotate the image, a professional annotation software LabelMe is usually used to annotate roads and other areas in the image. Use the annotation tool to annotate the collected images. The annotation result is a pixel-level mask corresponding to each image. Each pixel is assigned a label value according to its category to distinguish the road from other areas. After the annotation is completed, the dataset is sorted and divided into a training set, a validation set and a test set. 80% of the data is used for training, 10% of the data is used for validation, and 10% of the data is used for testing.
[0042] S2: Build a deep learning model for road detection based on the deep learning convolutional neural network model.
[0043] Furthermore, based on the convolutional neural network model of deep learning, a deep learning model for road detection is constructed. The model is divided into two stages: feature extraction and detection.
[0044] It should be noted that based on the YOLOv4 structure, the backbone network and neck network of the network model are constructed to extract features. The backbone network is composed of CSPDarkNet (Cross Stage Partial DarkNet). CSPDarkNet effectively reduces the computational burden of the network by introducing the cross-stage partial connection mechanism, and reduces the number of parameters and computational costs while maintaining a high detection accuracy. In terms of structural design, the features between some network layers are split and fused, which enhances the gradient flow and optimizes the learning ability of the network, making the model more efficient when processing complex tasks. The core idea of SPP (Spatial Pyramid Pooling) is to generate pooled features of different scales by applying multiple pooling operations of different sizes on the feature map. Then, these pooled features are spliced together to form a feature vector or feature map. This operation not only retains global and local information, but also enables the network to better adapt to various sizes and proportions of the input image. The SPP module uses pooling windows of different sizes to perform pooling operations on the input feature map. The result of pooling is multiple feature maps of different scales, which capture the features of different levels and different regions in the input feature map. All the pooled feature maps of different scales are concatenated through the channel dimension to generate a feature map containing multi-scale information. This concatenated feature map contains both global information (large-scale pooling) and detail information (small-scale pooling). The neck network is composed of Feature Pyramid Network (FPN), which is an architecture for enhancing the multi-scale detection capability of the target detection model. FPN uses feature maps of different levels in the convolutional neural network to construct a pyramid structure containing multi-scale features through top-down upsampling and lateral connections. In this way, it can capture high-resolution detail information and high-level semantic information at the same time, significantly improving the model's detection performance for targets of different sizes, especially when dealing with small targets and complex backgrounds. The head network of the constructed network is used for detection. The head network consists of three layers of upsampling networks. The last layer of feature maps after feature extraction is enlarged. The feature maps after feature extraction contain highly abstract semantic information. After the feature maps are enlarged to the original image size, the feature maps are output through the fully connected layer to obtain the segmented road information in the image.
[0045] S3: Use SIoU loss function to calculate the difference between prediction and actual results to optimize the model.
[0046] Furthermore, the SIoU loss function is used to calculate the difference between the prediction and the actual result, which is used to adjust the optimization model during training.
[0047] It should be noted that SIoU is an improved Intersection over Union (IoU) calculation method, which specifically considers targets of different scales and aspect ratios. In target detection tasks, targets of different scales and aspect ratios have different detection difficulties, and the traditional CIoU calculation method may ignore these differences, resulting in poor performance of the model when dealing with targets of different sizes and shapes. By introducing scale invariance, SIoU can better evaluate the degree of overlap between the predicted box and the true box, thereby improving the robustness of target detection. In contrast, although CIoU also considers factors such as shape similarity and center point distance, SIoU may have more advantages when dealing with targets with large scale changes. Therefore, SIoU helps to improve the model's detection performance for targets of different scales and shapes in application, so that it can maintain high detection accuracy and stability in various complex scenarios. The formula of SIOU is as follows:
[0048]
[0049]
[0050]
[0051] Among them, Δ is the distance loss, Ω is the shape loss, and IoU is the intersection over union loss. This loss function redefines the distance loss while taking into account the angle loss to make a more comprehensive evaluation of the position, size, and orientation of the object.
[0052] S4: The collected road dataset is trained and adjusted through the network model.
[0053] Furthermore, the collected road data set is trained with the built network model, and it is continuously adjusted until better detection results are achieved and applied in actual scenarios.
[0054] It should be noted that the network model was used to systematically train the collected diverse road data sets. In this process, the network parameters were constantly adjusted and optimized, and repeated iterations and verifications were performed to ensure that the model could accurately identify road feature elements. When the model's detection results achieved good accuracy and stability, they were successfully applied to actual scenarios, thus realizing the implementation of technology and the transformation of value.
[0055] Example 2, an embodiment of the present invention, provides a road detection method based on deep learning. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0056] First, the deep learning framework used for model training is Pytorch, the programming language is Python, the training graphics card is NVIDIA RTX309024GB, the training round epoch is 300, and the batch size is 16. Finally, the training result of this method reaches 92.3% mIoU. The figure shows the original road image and the actual detection result. It can be clearly seen that the model accurately marks the drivable area. The car transmits the image information in front of the car to the car computing terminal through the on-board camera, and the road information ahead is identified through the proposed algorithm, so as to safely avoid obstacles.
[0057] The proposed road detection method based on deep learning not only effectively reduces the computational burden, but also significantly improves the detection accuracy and efficiency. In particular, the SIoU loss function introduced further enhances the model's ability to detect objects of different scales and shapes by more accurately evaluating the degree of overlap between the predicted box and the true box. Ultimately, the model, which has been systematically trained and continuously optimized, has demonstrated excellent road detection performance in actual scenarios, achieving efficient implementation of technology and value conversion.
[0058] Example 3, reference Figure 5 , is an embodiment of the present invention, and provides a road detection system based on deep learning, including an image data acquisition and annotation module, a deep learning model construction and training module, and a model training and optimization module.
[0059] The image data acquisition and annotation module is used to collect road image data, process the collected data, and use annotation tools to create data sets. The deep learning model construction and training module is used to build a road detection model based on a deep learning convolutional neural network model, which includes feature extraction and target detection stages, and uses the SIoU loss function to optimize the model performance. The model training and optimization module is used to train with the prepared data set and the constructed network model, and improve the detection accuracy of the model through repeated optimization.
[0060] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0061] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0062] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0063] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc. It should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and are not limited. Although the present invention is described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.
[0064] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A road detection method based on deep learning, characterized in that: include: Collect and process image data of roads; Construct a deep learning model for road detection based on a deep learning convolutional neural network model; The SIoU loss function is used to calculate the difference between the prediction and the actual results to optimize the model; The collected road dataset is trained and adjusted through the network model.
2. The road detection method based on deep learning as claimed in claim 1, characterized in that: The collecting and processing of the image data of the road includes driving a vehicle equipped with a camera sensor on the road to collect the image data of the road. Obtain road image data from various real-life scenarios. The collected data should cover different weather, time, road types, and traffic conditions.
3. The road detection method based on deep learning as claimed in claim 2, characterized in that: The collecting and processing of the image data of the road includes making a data set, processing the collected data, and using a tool to annotate the image to make the data set; The collected raw image data needs to be preprocessed, including image cropping, scaling, and color adjustment operations; Remove noise from the image, standardize the image, and mark the road and non-road areas in the image. The marking result is a pixel-level mask corresponding to each image. Each pixel is assigned a label value according to its category to distinguish the road from the non-road area. After the marking is completed, the data set is sorted and divided into training set, validation set and test set, with 80% of the data used for training, 10% of the data used for validation, and 10% of the data used for testing.
4. The road detection method based on deep learning as claimed in claim 3, characterized in that: The deep learning-based convolutional neural network model constructs a deep learning model for road detection, including a convolutional neural network model based on deep learning, to construct a deep learning model for road detection, the model is divided into two stages: feature extraction and detection; Based on the YOLOv4 structure, the backbone network and neck network of the network model are constructed to extract features. The backbone network is composed of CSPDarkNet, which introduces a cross-stage partial connection mechanism. By applying the pooling operation on the feature map, pooled features are generated and the pooled features are concatenated to form a feature vector or feature map; Pooling operations are performed on the input feature map using pooling windows of different sizes. The feature maps generated by pooling capture the features of different levels and regions in the input feature map. All pooled feature maps of different scales are concatenated through the channel dimension to generate a feature map containing multi-scale information.
5. The road detection method based on deep learning as claimed in claim 4, characterized in that: The deep learning model for road detection based on the deep learning convolutional neural network model includes a neck network composed of a Feature Pyramid Network, and the FPN uses feature maps of different levels in the convolutional neural network to construct a pyramid structure containing multi-scale features through top-down upsampling and lateral connections; The head network of the constructed network is used for detection. The head network consists of three layers of upsampling networks. The feature map of the last layer after feature extraction is enlarged. After the feature map is enlarged to the original image size, the feature map passes through the fully connected layer to output the segmented road information in the image.
6. The road detection method based on deep learning as claimed in claim 5, characterized in that: The method of using the SIoU loss function to calculate the difference between the prediction and the actual result to optimize the model includes using the SIoU loss function to calculate the difference between the prediction and the actual result to adjust the optimization model during training; Introducing scale invariance, SIoU evaluates the degree of overlap between the predicted box and the true box. SIOU is expressed as: Among them, Δ represents distance loss, Ω represents shape loss, and IoU represents intersection over union loss.
7. The road detection method based on deep learning as claimed in claim 6, characterized in that: The road data set collected by the network model training and adjusting includes the road data set collected by the network model training, and continuously adjusting; Use the built network model to systematically train the collected diverse road data sets; Continuously adjust and optimize network parameters through repeated iterations and verification.
8. A system using the deep learning-based road detection method according to any one of claims 1 to 7, characterized in that: Including image data acquisition and annotation module, deep learning model construction and training module, model training and optimization module; The image data acquisition and annotation module is used to acquire road image data, process the acquired data, and create a data set using an annotation tool; The deep learning model construction and training module is used to construct a road detection model including feature extraction and target detection stages based on a deep learning convolutional neural network model, and uses a SIoU loss function to optimize model performance; The model training and optimization module is used to perform training using the prepared data set and the constructed network model, and improve the detection accuracy of the model through repeated optimization.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the road detection method based on deep learning described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the deep learning-based road detection method according to any one of claims 1 to 7 are implemented.