Unstructured Road Recognition Network Training Method, Application Method and Storage Medium
By introducing attention modules and depth separable convolution modules into the unstructured road recognition network, the problems of low accuracy and poor real-time performance of unstructured road recognition are solved, and fast and accurate road segmentation is achieved.
Patent Information
- Application Number
- CN202210085609.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-01-25
AI Technical Summary
The non-structured road identification method in the prior art has the problem of low recognition accuracy, poor real-time performance and susceptible to noise interference, making it difficult to achieve accurate and efficient road segmentation.
The unstructured road recognition network training method is adopted, and by introducing attention modules and depth separation convolution modules, the loss function is constructed, the network is trained to convergence, and the road segmentation is performed using a complete training network.
It realizes rapid and accurate segmentation and identification of unstructured roads, reduces network parameters, improves recognition performance and improves real-time performance.
Smart Images

Figure CN114627441B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and particularly to an unstructured road recognition network training method, an application method and a storage medium. Background Art
[0002] The unmanned driving technology is an important field of artificial intelligence. As a kind of unmanned platform, ground unmanned vehicles play increasingly important functions and tasks both in the civilian field and in the military field. An autonomous vehicle can use on-vehicle sensors to sense the surrounding environment of the vehicle, and control the steering and speed of the vehicle according to the road, vehicle position and obstacle information obtained by the sensing, so that the vehicle can drive safely and reliably on the road. Image Semantic Segmentation is a basic and extremely challenging task in the field of computer vision. Its goal is to estimate the class label of each pixel in the image, and it plays an increasingly important role in fields such as geographic information systems, unmanned driving, medical image analysis and robotics. For unmanned driving, image semantic segmentation can perform high-level processing on environmental information, thereby providing important road condition information for intelligent vehicles, making accurate judgments on road conditions, and ensuring the safety of autonomous vehicles.
[0003] In terms of road recognition, the roads on which vehicles travel can be divided into structured roads and unstructured roads. Structured roads generally refer to highways, urban arterial roads and other well-structured roads. Such roads have clear road marking lines, the road background environment is relatively simple, and the geometric features of the road are also relatively obvious. Therefore, the road detection problem for it can be simplified to the detection of lane lines or road boundaries. Unstructured roads generally refer to roads with low structure degree such as urban non-main roads and rural streets. Such roads have no lane lines and clear road boundaries. Coupled with the influence of shadows and water marks, etc., it is difficult to distinguish between road areas and non-road areas. The diverse road types, complex environmental backgrounds, as well as shadows, occlusions and changing weather, etc. are all the difficulties faced by unstructured road detection. For the pedestrian roads in areas such as communities, schools, scenic spots, and the countryside, because they generally have no obvious boundaries and the environment they are in is relatively complex, they should belong to unstructured roads, and there is relatively little research on such roads at present.
[0004] In the prior art, some scholars used an improved seed and Support Vector Machine (SVM) to propose an unstructured road detection and recognition method based on the combination of vision and 2D lidar detection. However, this method mainly targets forest environments, and the dataset needs to be expanded when applied in other scenarios. There are also problems with the existing unstructured road recognition methods, such as poor real-time performance in full-pixel domain calculation and classification processing, and being easily interfered by noise data. Therefore, an unstructured road recognition method based on SLIC (Simple linear iterative clustering) superpixel segmentation and improved region growth algorithm was proposed. However, there are deviations in the case of weak color and contrast. Therefore, the above existing methods have problems of poor recognition accuracy and weak real-time performance for unstructured roads. Therefore, how to perform accurate, efficient, and fast unstructured road recognition is an urgent problem to be solved. Summary of the Invention
[0005] In view of this, it is necessary to provide a training method, application method, and storage medium for an unstructured road recognition network to overcome the problems of inaccurate, inefficient, and slow recognition of unstructured roads in the prior art.
[0006] To solve the above technical problems, the present invention provides a training method for an unstructured road recognition network, including:
[0007] Obtain an image training sample set containing annotation information, where the annotation information includes the actual road classification label of each sample image pixel in the image training sample set;
[0008] Determine the value of the loss function of the unstructured road recognition network according to the actual road classification label, where the unstructured road recognition network includes a backbone network layer with an attention module added, and a pyramid pooling layer with an attention module and a depthwise separable convolution module added;
[0009] Adjust the parameters of the unstructured road recognition network according to the value of the loss function until the convergence condition is met, and determine the trained complete unstructured road recognition network.
[0010] Further, the determining the value of the loss function of the unstructured road recognition network according to the actual road classification label includes:
[0011] Input the image training sample set into the unstructured road recognition network, and determine the predicted road classification label corresponding to each sample image pixel;
[0012] Determine the loss function according to the error between the predicted road classification label and the actual road classification label.
[0013] Further, the network structure of the unstructured road recognition network includes an encoder and a decoder. The encoder includes an input layer, a depth convolutional neural network module, and an atrous spatial pyramid pooling module connected in sequence; the decoder includes a first decoding layer, a second decoding layer, a third decoding layer, and a decoding fusion layer.
[0014] Further, in the encoder, the depth convolutional neural network module includes a first convolutional block attention module, a first depth convolutional layer, a second depth convolutional layer, a third depth convolutional layer, a fourth depth convolutional layer, and the second convolutional block attention module connected in sequence, where:
[0015] The first convolutional block attention module is used to perform an attention mechanism operation that combines space and channels on the input image of the input layer to determine a first attention extraction map;
[0016] The first depth convolutional layer is used to perform a depthwise separable convolution operation on the first attention extraction map to determine a low-level feature map;
[0017] The second depth convolutional layer is used to perform a depthwise separable convolution operation on the low-level feature map to determine an intermediate-level feature map;
[0018] The third depth convolutional layer is used to perform a depthwise separable convolution operation on the intermediate-level feature map to determine a third depth convolutional feature map;
[0019] The fourth depth convolutional layer is used to perform a depthwise separable convolution operation on the third depth convolutional feature map to determine a fourth depth convolutional feature map;
[0020] The second convolutional block attention module is used to perform an attention mechanism operation that combines space and channels on the fourth depth convolutional feature map to determine a high-level feature map.
[0021] Further, in the encoder, the atrous spatial pyramid pooling module includes a first convolutional pooling layer to a fifth convolutional pooling layer, an encoding fusion layer, a third convolutional block attention module, and a convolutional output layer in parallel, where:
[0022] The first convolutional pooling layer to the fifth convolutional pooling layer are used to perform convolutional pooling operations on the high-level feature map respectively to determine a first pooling feature map to a fifth pooling feature map;
[0023] The encoding fusion layer is used to fuse the first pooling feature map to the fifth pooling feature map to determine a fused feature map;
[0024] The third convolutional block attention module is used to perform an attention mechanism operation that combines space and channels on the fused feature map to determine a third attention extraction map;
[0025] The convolutional output layer is used to perform a convolutional operation on the third attention extraction map to determine a convolutional output map.
[0026] Furthermore, in the decoder:
[0027] The first decoding layer is used to perform a depthwise separable convolutional operation on the low-level feature layer to determine a first decoded feature map;
[0028] The second decoding layer is used to perform a depthwise separable convolutional operation and a downsampling operation on the intermediate-level feature layer to determine a second decoded feature map;
[0029] The third decoding layer is used to perform a downsampling operation on the convolutional output map to determine a third decoded feature map;
[0030] The decoding fusion layer is used to fuse the first decoded feature map, the second decoded feature map, and the third decoded feature map and then perform a depthwise separable convolutional operation to determine the final decoded output map.
[0031] Furthermore, the loss function is represented by the following formula:
[0032]
[0033] where, represents the loss function, N represents the number of samples of the sample image pixels, represents the loss error of the i-th sample image pixel, represents the actual road classification label of the i-th sample image pixel, represents the probability that the i-th sample image pixel is predicted as an unstructured road.
[0034] The present invention also provides a method for applying an unstructured road recognition network, including:
[0035] Obtain a road image to be measured;
[0036] Input the road image to be measured into the trained unstructured road recognition network to determine a predicted road classification label, where the trained unstructured road recognition network is determined according to the unstructured road recognition network training method described above;
[0037] Determine a road segmentation map according to the predicted road classification label.
[0038] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method for training an unstructured road recognition network and / or the above-mentioned method for applying an unstructured road recognition network are implemented.
[0039] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for training an unstructured road recognition network and / or the above-mentioned method for applying an unstructured road recognition network are implemented.
[0040] Compared with the prior art, the beneficial effects of the present invention include: in the method for training an unstructured road recognition network, first, an image training sample set is constructed by using the actual road classification labels of each sample image pixel, and the image training sample set is effectively obtained; then, a corresponding loss function is constructed through the actual road classification labels to train the unstructured road recognition network, effectively mining the corresponding association between the sample image pixels and the actual road classification labels, adopting an attention module and a depthwise separable convolution module to improve the network recognition performance and realize the lightweight of the network; finally, the loss function is used to train the unstructured road recognition network until convergence to obtain a trained complete unstructured road recognition network. Subsequently, by using this unstructured road recognition network, the segmentation recognition result of the unstructured road can be quickly obtained. In the method for applying an unstructured road recognition network, first, a road image to be measured is effectively obtained; then, the above-mentioned trained complete unstructured road recognition network is used to effectively recognize the road image to be measured, and each pixel is recognized separately to output the corresponding road segmentation map. In summary, by introducing an attention module and a depthwise separable convolution module, the present invention improves the backbone network and the pooling network, fully extracts its multi-scale feature information, improves the network performance, reduces the network parameters, realizes the lightweight of the network, and achieves the purpose of quickly and accurately recognizing unstructured roads. Description of the Drawings
[0041] Figure 1 It is a schematic flowchart of an embodiment of the method for training an unstructured road recognition network provided by the present invention;
[0042] Figure 2 It is a schematic structural diagram of an embodiment of the attention module provided by the present invention;
[0043] Figure 3 It is a schematic structural diagram of an embodiment of the depthwise separable convolution module provided by the present invention;
[0044] Figure 4 Provided by the present invention Figure 1 It is a schematic flowchart of an embodiment of step S102 in
[0045] Figure 5 Schematic diagram of a structural embodiment of the unstructured road recognition network provided by the present invention;
[0046] Figure 6 Schematic flowchart of a method embodiment for applying the unstructured road recognition network provided by the present invention;
[0047] Figure 7 Experimental data of the training process provided by the present invention Figure 1 Schematic diagram of an embodiment;
[0048] Figure 8 Schematic comparison diagram of a segmentation result embodiment provided by the present invention;
[0049] Figure 9 Schematic diagram of a structural embodiment of the unstructured road recognition network training device provided by the present invention;
[0050] Figure 10 Schematic diagram of a structural embodiment of the unstructured road recognition network application device provided by the present invention;
[0051] Figure 11 Schematic diagram of a structural embodiment of the electronic device provided by the present invention. Detailed implementation manners
[0052] The following will specifically describe the preferred embodiments of the present invention in conjunction with the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, rather than to limit the scope of the present invention.
[0053] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0054] In the description of the present invention, referring to "embodiment" means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the described embodiments may be combined with other embodiments.
[0055] The present invention provides a method for training an unstructured road recognition network, an application method, and a storage medium, introducing an attention module and a depthwise separable convolution module to reduce network parameters, providing a new idea for further improving the accuracy and efficiency of unstructured road recognition.
[0056] Before describing the embodiments, the relevant terms are defined as follows:
[0057] Unstructured road: Actual roads can generally be divided into two categories: structured roads and unstructured roads. Structured roads generally refer to highways, urban arterial roads and other well-structured roads. Such roads have clear road marking lines, a relatively simple road background environment, and obvious geometric features. Therefore, the road detection problem for it can be simplified to the detection of lane lines or road boundaries. Unstructured roads generally refer to roads with a low degree of structure such as urban non-main roads and rural streets. Such roads do not have lane lines and clear road boundaries. Coupled with the influence of shadows and water stains, etc., it is difficult to distinguish between road areas and non-road areas. The diverse road types, complex environmental backgrounds, as well as shadows, water stains and changing weather, etc. are all the difficulties faced by unstructured road detection and are also the main research directions of current road recognition technology.
[0058] Attention mechanism: The attention mechanism is the attention to the weight distribution of the input. The attention mechanism was first used in the encoder-decoder. The attention mechanism obtains the input variable of the next layer by taking a weighted average of the hidden states of all time steps of the encoder.
[0059] Depthwise separable convolution: In a convolutional neural network, the spatial dimension and the channel (depth) dimension of the feature map can be decoupled. The standard convolution calculation uses a weight matrix to achieve the joint mapping of spatial and channel dimension features, but at the cost of high computational complexity, high memory overhead, and a large number of weight coefficients. Conceptually, depthwise separable convolution maps the spatial and channel dimensions separately and combines the results, reducing the number of weight coefficients while basically retaining the representation learning ability of the convolutional kernel. Considering the difference in the number of input and output channels, the number of weights of depthwise separable convolution is about 10% to 25% of the number of weights of standard convolution. Some convolutional neural networks built using depthwise separable convolution, such as Xception, perform better in the image recognition task of the ImageNet dataset than Inception v3 with the same hidden layer weights but using standard convolution and Inception modules. Therefore, depthwise separable convolution is also considered to improve the utilization efficiency of convolutional kernel parameters.
[0060] Based on the description of the above technical terms, in the prior art, neural networks are often directly used to identify unstructured roads. However, there are too many network parameters, resulting in disadvantages such as low accuracy and poor timeliness. Traditional semantic segmentation extracts low-level semantics of images, such as size, texture, color, etc. In complex environments, there are obvious defects in robustness and accuracy. In recent years, with the rapid development of deep learning, breakthroughs have been made in the field of semantic segmentation. In 2015, Long et al. creatively proposed the Fully Convolutional Network (FCN) based on the deep convolutional neural network, marking a leap-forward progress of deep learning in the field of semantic segmentation and having milestone significance. Compared with traditional semantic segmentation methods, semantic segmentation methods based on deep learning can obtain more and higher-level semantic information to express the information in images. The Deeplab series of architectures was first proposed by Google. The early DeepLab v1, DeepLab v2, and DeepLab v3 adopted a cascaded architecture. With the emergence of semantic segmentation architectures such as U-Net and SegNet, the encoder-decoder structure has become the mainstream, and one of the most representative is DeepLab v3+. Therefore, the present invention aims to propose an efficient and accurate training method and application method for an unstructured road recognition network based on the DeepLab v3+ model.
[0061] The following will separately elaborate on specific embodiments in detail:
[0062] An embodiment of the present invention provides a training method for an unstructured road recognition network, combined with Figure 1 seen as Figure 1 is a schematic flowchart of an embodiment of the training method for the unstructured road recognition network provided by the present invention, including steps S101 to S103, where:
[0063] In step S101, an image training sample set containing annotation information is obtained, where the annotation information includes the actual road classification label of each sample image pixel in the image training sample set;
[0064] In step S102, the value of the loss function of the unstructured road recognition network is determined according to the actual road classification label, where the unstructured road recognition network includes a backbone network layer with an attention module added, and a pyramid pooling layer with an attention module and a depthwise separable convolution module added;
[0065] In step S103, the parameters of the unstructured road recognition network are adjusted according to the value of the loss function until the convergence condition is met, and a trained complete unstructured road recognition network is determined.
[0066] In an embodiment of the present invention, in the method for training an unstructured road recognition network, first, an image training sample set is constructed by using the actual road classification labels of each sample image pixel, and the image training sample set is effectively obtained; then, a corresponding loss function is constructed through the actual road classification labels to train the unstructured road recognition network, effectively mining the corresponding association between the sample image pixels and the actual road classification labels, and an attention module and a depthwise separable convolution module are adopted to improve the network recognition performance and realize the lightweight of the network; finally, the loss function is used to train the unstructured road recognition network until convergence, and a trained complete unstructured road recognition network is obtained. Subsequently, by using this unstructured road recognition network, the segmentation and recognition results of the unstructured road can be quickly obtained.
[0067] As a preferred embodiment, in combination with Figure 2 it can be seen that Figure 2 FIG. is a schematic structural diagram of an embodiment of the attention module provided by the present invention. The attention module is used to combine spatial and channel attention, including a spatial attention module and a channel attention module, and its specific structure is shown in Figure 2 .
[0068] Among them, it can be seen from Figure 2 that the input feature map first passes through the channel attention module to model the dependence relationship between each channel in the image, so as to selectively enhance the channel information of the mutually dependent features and further improve the feature expression ability of the network. Each channel of the feature represents a dedicated detector. Therefore, channel attention focuses on what kind of features are meaningful, and the calculation process is shown in the following formula:
[0069]
[0070] In the formula, MC(F) represents the channel attention map, F is the input feature map, σ represents the sigmoid activation function, MLP (Multi-Layer Perceptron) is a multi-layer perceptron, and MaxPool and AvgPool respectively represent the global maximum pooling layer and the global average pooling layer.
[0071] Among them, the input feature map is first subjected to global max pooling MaxPool and global average pooling AvgPool operations to summarize the spatial information of the feature map, generating two different spatial context descriptors, representing the average pooling feature and the max pooling feature respectively. Then, these two feature descriptors are forward propagated to the shared network MLP, which consists of a multi-layer perceptron MLP with a hidden layer. Finally, the output features are subjected to pixel-wise summation operation, and then passed through a sigmoid activation operation to generate a channel attention feature map. Then, through the spatial attention module, the dependence of each pixel in the image on other pixels is modeled, and the spatial position information is selectively strengthened. The spatial attention module pays more attention to the more important regions in the image, and at the same time can reduce the interference of surrounding redundant information, avoid affecting valuable information, and increase the representation ability. The spatial attention calculation process is shown as follows:
[0072]
[0073] In the formula, MS(F) represents the spatial attention map, represents a convolutional kernel with a size of 7×7.
[0074] It should be noted that spatial attention is a supplement to channel attention. For the feature map processed by the channel attention module, average pooling and max pooling operations are applied along the channel axis, and they are concatenated to generate an effective feature descriptor. Behind the concatenated feature descriptor, a convolutional layer with a relatively large convolutional kernel is used to synthesize the features around each point, thereby generating a spatial attention feature map, representing the weights of the spatial positions of the input feature map, that is, which regions need more attention and where the redundant information can reduce the attention and its weight can be reduced.
[0075] It should be further noted that both max pooling and average pooling are included in the two attention models. Average pooling can effectively encode the global feature attributes of this point and obtain the feature attributes of this point. At the same time, applying max pooling can retain some information of relatively unique features, which can compensate for the neglect of unique features by average pooling due to the averaging operation of global information on the channel. Compared with only using one of them, the expression ability of the network can be greatly improved. Combining average pooling and max pooling can obtain a more refined feature map.
[0076] As a preferred embodiment, combined with Figure 3 to see, Figure 3 is a schematic structural diagram of an embodiment of the depthwise separable convolution module provided by the present invention. The depthwise separable convolution module is used to reduce the number of network parameters, and its specific structure is shown in Figure 3 .
[0077] In the embodiments of the present invention, the depthwise separable convolution
[24] can be divided into depthwise convolution and pointwise convolution. The processes of conventional convolution and depthwise separable convolution are as Figure 3 shown. Essentially, the depthwise separable convolution is the decomposition of the 3D convolution kernel (decomposition on the depth channel). Although only a very small change is made to the conventional convolution, it significantly reduces the number of parameters and is beneficial to the lightweight of the network.
[0078] As a preferred embodiment, in combination with Figure 4 viewed as Figure 4 provided by the present invention Figure 1 is a schematic flowchart of an embodiment of step S102 in []. In step S102, it specifically includes steps S201 to S202, where:
[0079] In step S201, the image training sample set is input into the unstructured road recognition network to determine the predicted road classification label corresponding to each sample image pixel;
[0080] In step S202, according to the error between the predicted road classification label and the actual road classification label, the loss function is determined.
[0081] In the embodiments of the present invention, by using the predicted road classification label and the actual road classification label, the loss function is effectively constructed to complete the convergence training of the network.
[0082] As a preferred embodiment, in combination with Figure 5 viewed as Figure 5 is a schematic structural diagram of an embodiment of the unstructured road recognition network provided by the present invention. The network structure of the unstructured road recognition network includes an encoder and a decoder. The encoder includes an input layer, a depth convolutional neural network module, and an atrous spatial pyramid pooling module connected in sequence; the decoder includes a first decoding layer, a second decoding layer, a third decoding layer, and a decoding fusion layer.
[0083] In the embodiments of the present invention, the structures of the encoder and the decoder are set to ensure the lightweight of the network and the diversity of recognition features.
[0084] In a specific embodiment of the present invention, the classic Resnet101 is adopted as the backbone network and certain improvements are made to it. A CBAM module (i.e., the attention module) is added before the first layer and after the last layer of Resnet101 to make full use of the detailed information of the feature image, thereby reducing the phenomena of misclassification and missed classification, and increasing the diversity of features. In addition, in the original model, only the feature map with a size of 1 / 4 in the backbone network is used as the low-level feature for subsequent processing, ignoring the rich semantic information in other feature maps generated during the process. Therefore, in the present invention, the feature map with a size of 1 / 8 generated in the backbone network is used as the intermediate feature map to make full use of the semantic features; a CBAM attention mechanism is added to the ASPP module (i.e., the atrous spatial pyramid pooling module) to extract the deep features of the image, and the ordinary convolutional layer in the ASPP module is replaced with a depthwise separable convolutional layer to reduce the number of parameters and the amount of computation and speed up the training speed. The backbone network adopted by the encoder is Resnet101 integrated with the attention mechanism.
[0085] As a preferred embodiment, in the encoder, the deep convolutional neural network module includes a first convolutional block attention module, a first depth convolutional layer, a second depth convolutional layer, a third depth convolutional layer, a fourth depth convolutional layer, and the second convolutional block attention module connected in sequence, where:
[0086] The first convolutional block attention module is used to perform an attention mechanism operation that combines space and channels on the input image of the input layer to determine a first attention extraction map;
[0087] The first depth convolutional layer is used to perform a depthwise separable convolution operation on the first attention extraction map to determine a low-level feature map;
[0088] The second depth convolutional layer is used to perform a depthwise separable convolution operation on the low-level feature map to determine an intermediate feature map;
[0089] The third depth convolutional layer is used to perform a depthwise separable convolution operation on the intermediate feature map to determine a third depth convolutional feature map;
[0090] The fourth depth convolutional layer is used to perform a depthwise separable convolution operation on the third depth convolutional feature map to determine a fourth depth convolutional feature map;
[0091] The second convolutional block attention module is used to perform an attention mechanism operation that combines space and channels on the fourth depth convolutional feature map to determine a high-level feature map.
[0092] In the embodiment of the present invention, multiple hierarchical structures of an encoder are set up to complete feature recognition of an input image, obtaining multiple feature images. The CBAM module is utilized to fully utilize the detailed information of the feature images, thereby reducing misclassification and missed classification phenomena, increasing the diversity of features, and multiple deep convolutional layers are set up to fully utilize semantic features, reducing the number of parameters and the amount of computation.
[0093] As a preferred embodiment, in the encoder, the atrous spatial pyramid pooling (ASPP) module includes a first convolutional pooling layer to a fifth convolutional pooling layer in parallel, an encoding fusion layer, a third convolutional block attention module, and a convolutional output layer, where:
[0094] The first convolutional pooling layer to the fifth convolutional pooling layer are used to perform convolutional pooling operations on the high-level feature map respectively to determine a first pooled feature map to a fifth pooled feature map;
[0095] The encoding fusion layer is used to fuse the first pooled feature map to the fifth pooled feature map to determine a fused feature map;
[0096] The third convolutional block attention module is used to perform an attention mechanism operation that combines space and channels on the fused feature map to determine a third attention extraction map;
[0097] The convolutional output layer is used to perform a convolutional operation on the third attention extraction map to determine a convolutional output map.
[0098] In the embodiment of the present invention, the atrous spatial pyramid pooling module is set up to further extract multi-faceted feature information and improve the segmentation effect.
[0099] As a preferred embodiment, in the decoder:
[0100] The first decoding layer is used to perform a depthwise separable convolutional operation on the low-level feature layer to determine a first decoded feature map;
[0101] The second decoding layer is used to perform a depthwise separable convolutional operation and a downsampling operation on the intermediate-level feature layer to determine a second decoded feature map;
[0102] The third decoding layer is used to perform a downsampling operation on the convolutional output map to determine a third decoded feature map;
[0103] The decoding fusion layer is used to fuse the first decoded feature map, the second decoded feature map, and the third decoded feature map and then perform a depthwise separable convolutional operation to determine the final decoded output map.
[0104] In the embodiments of the present invention, the high-level and low-level feature maps and the intermediate feature maps with a size of 1 / 8 added in the encoder are finally fused to effectively restore the detailed information of the high-level features and improve the segmentation effect. The number of parameters in the original decoder is also relatively large, and the depthwise separable convolutional layer can be used to replace the ordinary convolutional layer to reduce the number of parameters.
[0105] As a preferred embodiment, the loss function is represented by the following formula:
[0106]
[0107] Wherein, represents the loss function, N represents the number of samples of the pixels of the sample image, represents the loss error of the i-th sample image pixel, represents the actual road classification label of the i-th sample image pixel, represents the probability that the i-th sample image pixel is predicted as an unstructured road.
[0108] In the embodiments of the present invention, a loss function is set to ensure the effective training and convergence of the unstructured road recognition network.
[0109] The embodiments of the present invention also provide a method for applying an unstructured road recognition network. In combination with Figure 6 seen, Figure 6 is a schematic flowchart of an embodiment of the method for applying the unstructured road recognition network provided by the present invention, including steps S601 to S603, wherein:
[0110] In step S601, an image of the road to be measured is obtained;
[0111] In step S602, the image of the road to be measured is input into the trained unstructured road recognition network to determine the predicted road classification label, wherein the trained unstructured road recognition network is determined according to the above-mentioned method for training the unstructured road recognition network;
[0112] In step S603, according to the predicted road classification label, a road segmentation map is determined.
[0113] In the embodiments of the present invention, first, an effective acquisition of the image of the road to be measured is performed; then, the above-mentioned trained unstructured road recognition network is used to effectively recognize the image of the road to be measured, and each pixel thereof is separately recognized, and then the corresponding road segmentation map can be output.
[0114] Next, in combination with a specific application scenario, the training process of the technical solution of the present invention is more clearly described. Among them, the unstructured road recognition network is named Improved DeepLab v3+, and the specific process is as follows:
[0115] First, preparation of the dataset:
[0116] The dataset used in this invention is the unstructured roads within a certain university, which are photographed under different lighting conditions and shooting angles. The image resolution is 3024×4032. To better utilize the image information, the data is divided into a training set, a validation set, and a test set in a ratio of 4:2:2, and the images are normalized to 512×512. The collected images are enhanced using the opencv method, and operations such as horizontal, vertical, and diagonal flipping, image translation, and scaling are performed on the images, thus greatly expanding the dataset. A total of 3211 images are obtained, which is beneficial for training a better network model.
[0117] Second, experimental platform and training details:
[0118] The program of this invention is implemented using the deep learning framework pytorch, and the machine configuration is shown in Table 1.
[0119] Table 1 Experimental machine configuration
[0120]
[0121] The original model adopted in this invention is Deeplabv3+, the backbone network uses Resnet101, the input image size (crop size) is 513×513, the initial learning rate is 0.007, the "poly" learning strategy is adopted, the momentum is 0.9, the weight decay rate is set to 0.0005 to prevent overfitting, the optimizer uses SGD, the number of training epochs is 100, the batch-size is 8, and pre-trained parameters are adopted. The above hyperparameters are set only once for comparison experiments. As shown in the following formula:
[0122]
[0123] In the formula, the power parameter controls the lowest value that the learning rate reaches in the saturation state, which is set to 0.9, new_lr represents the new learning rate, base_lr represents the initial learning rate of 0.007, iter represents the number of iterations, and max_iter is the total number of iterations.
[0124] Among them, the cross-entropy loss function is adopted. In the case of binary classification, there are only two cases for the result that the model finally needs to predict, and the obtained probabilities are p and 1-p. The expression is:
[0125]
[0126] In the formula, y iDenote the label of sample \(i\), where the positive class is \(1\) and the negative class is \(0\). p i Denote the probability that sample \(i\) is predicted as the positive class.
[0127] In the field of semantic segmentation of images, common performance evaluation metrics mainly include pixel accuracy (PA), mean intersection over union (mIoU), and frequency weighted intersection over union (FWIoU), etc. The accuracy evaluation metric mainly adopted in this invention is mIoU. The specific definition and expression of mIoU are shown in the following formula, which represents the result of summing and averaging the ratios of the intersections to the unions of the predicted values and the ground truth values for each class. It is the most commonly used evaluation metric in the current field of image semantic segmentation. As shown in the following formula:
[0128]
[0129] In the formula, \(N\) represents the number of columns of image pixels; \(T\) i represents the total number of pixels of the \(i\)-th class; \(X\) ii represents the total number of pixels with the actual class \(i\) and the predicted class \(i\); \(X\) ji represents the total number of pixels with the actual class \(i\) and the predicted class \(j\).
[0130] The training process of this invention is regarded as a binary classification problem. Initialize the hyperparameters and start training. Format the dataset collected in this invention according to the PASCAL VOC2012 dataset. The ratio of the training set, validation set, and test set is 4:2:2. Input the pictures in the training set, and after learning by the neural network, verify with the pictures in the validation set, evaluate the value of mIoU. The output results are only two types, namely the background class and the road area class. After obtaining the results of each round, adjust the learning rate through the learning strategy, and then conduct the next round of training until the training ends to obtain the model with the optimal mIoU finally. The accuracy of the model can be tested through the test set. Display the training process through the tensorboard visualization tool as Figure 7 , Figure 7 which is the experimental data of the training process provided by this invention Figure 1 The schematic diagram of the embodiment.
[0131] Among them, it can be seen from the figure that the mIoU value reaches 98.56%, the accuracy is 99.37%, the loss of the training set is 5.13, and the loss of the validation set is 0.21.
[0132] Third, result analysis:
[0133] First, in terms of network parameter quantity, model complexity and training time, the original Deeplabv3+ model has a large number of parameters and high model complexity, which greatly increases the difficulty of training. One of the goals of the present invention is to minimize the parameters and model complexity without greatly affecting the accuracy. Depth-wise separable convolution can greatly reduce the number of parameters in the training process and improve the model training efficiency. Table 2 compares the parameters, complexity and training time of the PSP, DeepLab v3, DeepLab v3+ and the improved DeepLab v3+ network of the present invention. The results show that the parameters of the improved model are reduced by 21.74%, FLOPs are reduced by 34.8%, and the training time is reduced by 15.31% compared with the original model.
[0134] Table 2
[0135]
[0136] Secondly, for model size, running time, speed and accuracy, PSP, DeepLab v3, DeepLab v3+ and the improved DeepLab v3+ network of the present invention are trained on the data set collected by the present invention, and the comparison of model size, running loading time, speed and mIoU value is shown in Table 3. It can be seen from the data in the table that the model obtained by the improved network training of the present invention is reduced by 22.32% in volume, while the running loading time, speed and mIoU value are improved. The effectiveness of the network model proposed by the present invention is further verified.
[0137] Table 3
[0138]
[0139] Finally, for the segmentation results, combined Figure 8 Come and see, Figure 8 The schematic diagram of the comparison of the segmentation results of an embodiment of the present invention is shown in Figure 1. The improved algorithm of the present invention is verified on an unstructured road dataset. A darker test picture is selected to test the robustness of the model trained under the condition of poor visualization effect. The comparison of the segmentation results is shown in Figure 1. Figure 8 As shown in the figure, the segmentation results of the original image on PSP, DeepLab v3, DeepLab v3+ and the improved DeepLab v3+ network of the present invention are shown. It can be seen from the figure that the improved network of the present invention has a better segmentation effect on unstructured roads, can smooth the edges of the road, and can achieve higher segmentation accuracy even in poor visibility. At the same time, the overall model is more lightweight and easy to transplant.
[0140] The embodiment of the present invention also provides an unstructured road recognition network training device, combined withFigure 9 As shown Figure 9 FIG. 3 is a schematic structural diagram of an embodiment of an unstructured road recognition network training device provided by the present invention. The unstructured road recognition network training device 900 includes:
[0141] A first acquisition unit 901, configured to acquire an image training sample set including annotation information, where the annotation information includes an actual road classification label of each sample image pixel in the image training sample set;
[0142] A first processing unit 902, configured to determine a value of a loss function of the unstructured road recognition network according to the actual road classification label, where the unstructured road recognition network includes a backbone network layer with an attention module added, and a pyramid pooling layer with an attention module and a depthwise separable convolution module added;
[0143] A training unit 903, configured to adjust parameters of the unstructured road recognition network according to the value of the loss function until a convergence condition is met, and determine a trained complete unstructured road recognition network.
[0144] For a more specific implementation manner of each unit of the unstructured road recognition network training device, reference may be made to the description of the above unstructured road recognition network training method, and it has a similar beneficial effect, which will not be elaborated here.
[0145] An embodiment of the present invention further provides an unstructured road recognition network application device. Combining Figure 10 As shown Figure 10 FIG. 4 is a schematic structural diagram of an embodiment of an unstructured road recognition network application device provided by the present invention. The unstructured road recognition network application device 1000 includes:
[0146] A second acquisition unit 1001, configured to acquire an image of a road to be measured;
[0147] A second processing unit 1002, configured to input the image of the road to be measured into the trained complete unstructured road recognition network to determine a predicted road classification label, where the trained complete unstructured road recognition network is determined according to the above unstructured road recognition network training method;
[0148] A segmentation unit 1003, configured to determine a road segmentation map according to the predicted road classification label.
[0149] For a more specific implementation manner of each unit of the unstructured road recognition network application device, reference may be made to the description of the above unstructured road recognition network application method, and it has a similar beneficial effect, which will not be elaborated here.
[0150] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-described method for training an unstructured road recognition network and / or the above-described method for applying an unstructured road recognition network are implemented.
[0151] Generally speaking, computer instructions for implementing the method of the present invention can be carried by any combination of one or more computer-readable storage media. A non-transitory computer-readable storage medium can include any computer-readable medium except for the signals propagating temporarily themselves.
[0152] The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the context of the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0153] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. In particular, the Python language suitable for neural network computing and platform frameworks based on TensorFlow, PyTorch, etc. can be used. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0154] An embodiment of the present invention also provides an electronic device, in combination with Figure 11 viewed Figure 11Schematic diagram of the structure of an embodiment of the electronic device provided by the present invention. The electronic device 1100 includes a processor 1101, a memory 1102, and a computer program stored on the memory 1102 and executable on the processor 1101. When the processor 1101 executes the program, it implements the above-described method for training an unstructured road recognition network, and / or the above-described method for training an unstructured road recognition network, and / or the above-described method for applying an unstructured road recognition network.
[0155] As a preferred embodiment, the above electronic device 1100 further includes a display 1103 for displaying the processor 1101 executing the above-described method for training an unstructured road recognition network, and / or the above-described method for applying an unstructured road recognition network.
[0156] Exemplarily, the computer program may be divided into one or more modules / units. One or more modules / units are stored in the memory 1102 and executed by the processor 1101 to complete the present invention. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1100. For example, the computer program may be divided into the first acquisition unit 901, the first processing unit 902, the training unit 903, the second acquisition unit 1001, the second processing unit 1002, and the segmentation unit 1003 in the above embodiment. The specific functions of each unit are as described above and will not be elaborated here one by one.
[0157] The electronic device 1100 may be a desktop computer, a notebook, a palm computer, a smart phone, or other devices with an adjustable camera module.
[0158] Among them, the processor 1101 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 1101 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0159] Among them, the memory 1102 can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electric Erasable Programmable Read-Only Memory (EEPROM), etc. Among them, the memory 1102 is used to store programs. After receiving an execution instruction, the processor 1101 executes the program. The method defined by the process disclosed in any embodiment of the foregoing embodiments of the present invention can be applied to the processor 1101 or implemented by the processor 1101.
[0160] Among them, the display 1103 can be an LCD display screen or an LED display screen. For example, the display screen on a mobile phone.
[0161] It can be understood that Figure 11 The structure shown is only a schematic structural diagram of the electronic device 1100, and the electronic device 1100 may further include more or fewer components than Figure 11 those shown. Figure 11 Each component shown in can be implemented by hardware, software, or a combination thereof.
[0162] According to the computer-readable storage medium and the electronic device provided in the foregoing embodiments of the present invention, the unstructured road recognition network training method and / or the unstructured road recognition network application method as described above can be implemented with reference to the content specifically described, and has beneficial effects similar to those of the unstructured road recognition network training method and / or the unstructured road recognition network application method as described above, which will not be elaborated here.
[0163] The present invention discloses a method for training an unstructured road recognition network, an application method, and a storage medium. In the method for training an unstructured road recognition network, first, an image training sample set is constructed by using the actual road classification labels of each pixel of the sample image, and the image training sample set is effectively obtained; then, a corresponding loss function is constructed through the actual road classification labels to train the unstructured road recognition network, effectively mining the corresponding association between the sample image pixels and the actual road classification labels, adopting an attention module and a depthwise separable convolution module to improve the network recognition performance and achieve the lightweight of the network; finally, the unstructured road recognition network is trained by using the loss function until convergence to obtain a trained and complete unstructured road recognition network. Subsequently, by using this unstructured road recognition network, the segmentation and recognition result of the unstructured road can be quickly obtained. In the application method of the unstructured road recognition network, first, the image of the road to be measured is effectively obtained; then, the above-trained and complete unstructured road recognition network is used to effectively recognize the image of the road to be measured, and each pixel is separately recognized to output the corresponding road segmentation map.
[0164] The technical solution of the present invention improves the backbone network and the pooling network by introducing an attention module and a depthwise separable convolution module, fully extracts its multi-scale feature information, improves the network performance, reduces the network parameters, realizes the lightweight of the network, and achieves the purpose of quickly and accurately recognizing unstructured roads.
[0165] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for training an unstructured road recognition network, characterized in that Including: Obtain an image training sample set containing annotation information, where the annotation information includes the actual road classification label of each sample image pixel in the image training sample set; Determine the value of the loss function of the unstructured road recognition network according to the actual road classification label, where the unstructured road recognition network includes a backbone network layer with an attention module added, and a pyramid pooling layer with an attention module and a depthwise separable convolution module added; Adjust the parameters of the unstructured road recognition network according to the value of the loss function until the convergence condition is met, and determine the trained complete unstructured road recognition network; The network structure of the unstructured road recognition network includes an encoder and a decoder. The encoder includes an input layer, a depth convolutional neural network module, and an atrous spatial pyramid pooling module connected in sequence; the decoder includes a first decoding layer, a second decoding layer, a third decoding layer, and a decoding fusion layer; In the encoder, the depth convolutional neural network module includes a first convolutional block attention module, a first depth convolutional layer, a second depth convolutional layer, a third depth convolutional layer, a fourth depth convolutional layer, and a second convolutional block attention module connected in sequence, where: The first convolutional block attention module is used to perform a spatial and channel combined attention mechanism operation on the input image of the input layer to determine a first attention extraction map; The first depth convolutional layer is used to perform a depthwise separable convolution operation on the first attention extraction map to determine a low-level feature map; The second depth convolutional layer is used to perform a depthwise separable convolution operation on the low-level feature map to determine an intermediate-level feature map; The third depth convolutional layer is used to perform a depthwise separable convolution operation on the intermediate-level feature map to determine a third depth convolutional feature map; The fourth depth convolutional layer is used to perform a depthwise separable convolution operation on the third depth convolutional feature map to determine a fourth depth convolutional feature map; The second convolutional block attention module is used to perform a spatial and channel combined attention mechanism operation on the fourth depth convolutional feature map to determine a high-level feature map; The loss function is represented by the following formula: Among them, represents the loss function, N represents the number of samples of the sample image pixels, represents the loss error of the i-th sample image pixel, represents the actual road classification label of the i-th sample image pixel, represents the probability that the i-th sample image pixel is predicted as an unstructured road.
2. The method for training an unstructured road recognition network according to claim 1, wherein The determining the value of the loss function of the unstructured road recognition network according to the actual road classification label includes: Input the image training sample set into the unstructured road recognition network to determine the predicted road classification label corresponding to each sample image pixel; Determine the loss function according to the error between the predicted road classification label and the actual road classification label.
3. The method for training an unstructured road recognition network according to claim 1, wherein In the encoder, the atrous spatial pyramid pooling module includes a first convolutional pooling layer to a fifth convolutional pooling layer, an encoding fusion layer, a third convolutional block attention module, and a convolutional output layer connected in parallel, where: The first convolutional pooling layer to the fifth convolutional pooling layer are used to perform convolutional pooling operations on the high-level feature map respectively to determine a first pooling feature map to a fifth pooling feature map; The encoding fusion layer is used to fuse the first pooling feature map to the fifth pooling feature map to determine a fused feature map; The third convolutional block attention module is used to perform an attention mechanism operation that combines space and channels on the fused feature map to determine a third attention extraction map; The convolutional output layer is used to perform a convolutional operation on the third attention extraction map to determine a convolutional output map.
4. The method for training an unstructured road recognition network according to claim 1, wherein In the decoder: The first decoding layer is used to perform a depthwise separable convolutional operation on the low-level feature layer to determine a first decoded feature map; The second decoding layer is used to perform a depthwise separable convolutional operation and a downsampling operation on the intermediate-level feature layer to determine a second decoded feature map; The third decoding layer is used to perform a downsampling operation on the convolutional output map to determine a third decoded feature map; The decoding fusion layer is used to fuse the first decoded feature map, the second decoded feature map, and the third decoded feature map and then perform a depthwise separable convolutional operation to determine a final decoded output map.
5. A method for applying an unstructured road recognition network, characterized in that, Comprising: Obtain a road image to be measured; Input the road image to be measured into a trained unstructured road recognition network to determine a predicted road classification label, where the trained unstructured road recognition network is determined according to the unstructured road recognition network training method described in any one of claims 1 to 4; Determine a road segmentation map according to the predicted road classification label.
6. An electronic device, comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the unstructured road recognition network training method described in any one of claims 1 to 4, and / or the unstructured road recognition network application method described in claim 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the unstructured road recognition network training method described in any one of claims 1 to 4, and / or the unstructured road recognition network application method described in claim 5.
Citation Information
Patent Citations
A substation patrol robot road scene recognition method based on depth learning
CN109446970A
Lane line multi-task learning detection method based on road segmentation
CN110414387A