Road surface type identification method based on improved lightweight neural network structure

By improving the lightweight neural network structure, introducing attention mechanism CBAM and hardswish activation functions, ShuffleNet V2 was transformed, which solved the lack of performance of traditional models in scenarios of real-time and limited computing resources, and achieved efficient and accurate road type recognition.

CN119942481AInactive Publication Date: 2025-05-06JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411876658.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional large-scale convolutional neural network models perform poorly in scenarios with high real-time requirements and cannot run on embedded devices with computing resource constraints, making it difficult to accurately identify road surface types.

Method used

The improved lightweight neural network structure is adopted, the attention mechanism CBAM and the nonlinear activation function hardswish are introduced to transform the basic unit of ShuffleNet V2 network to reduce the computing cost and memory access cost.

Benefits of technology

While maintaining a small amount of parameter calculation, it improves network performance, reduces computing costs, and makes road type identification more efficient and accurate, and is suitable for embedded devices with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942481A_ABST
    Figure CN119942481A_ABST
Patent Text Reader

Abstract

The invention relates to a pavement type identification method based on an improved lightweight neural network structure, and the method comprises the steps: collecting pictures of three typical pavements, and building a training set and a test set of different pavement types after data enhancement, so as to train a neural network model; a lightweight neural network ShuffleNet V2 backbone model is established; according to the method, an attention mechanism CBAM and a nonlinear activation function hardswitch are introduced to modify a basic unit of a ShuffleNet V2 network, an improved lightweight neural network structure is obtained, the model parameter quantity is effectively reduced, the model computing resource utilization efficiency is improved, the network performance is improved while the small parameter computing quantity is kept, and the method is suitable for large-scale popularization and application. And the activation function has a certain spatial modeling capability while the MAC is reduced. According to the invention, a neural network which is based on deep learning and can identify road types is constructed, and a network model which is small in calculation amount, high in calculation efficiency and better in performance is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of automobile suspension, and in particular relates to a road type recognition method based on an improved lightweight neural network structure. Background Art

[0002] With the advancement of suspension technology research, people have put forward higher requirements for the performance of automobile suspension. The safety of automobiles on various road surfaces and their riding comfort have gradually attracted attention, which has led to the rapid development of adaptive active suspension control technology based on different road surfaces. However, how to accurately identify different types of road surfaces (such as flexible road surfaces, rigid road surfaces, and special road surfaces) to optimize the active suspension control performance has become a new research problem.

[0003] At present, some scholars have proposed an image recognition method based on deep learning convolutional neural network (CNN). By building a deep learning model and training the model with training sets and test sets of different types of road pictures, a model with high recognition accuracy can be obtained. However, traditional large convolutional neural network models, such as AlexNet, GoogLeNet and ResNet, generally have many parameters and large computational complexity. Therefore, they usually perform poorly in scenarios with high real-time requirements. In addition, due to the large size of the model, it can usually only run on GPU devices with high computing power and cannot be ported to embedded devices with computing resource constraints. This type of problem is particularly obvious when the computing resources of vehicle intelligent control are limited.

[0004] Based on this, it is very necessary to use lightweight neural networks to improve traditional neural network image recognition methods to reduce the computational cost of identifying road types. At the same time, if the attention mechanism CBAM is introduced and the nonlinear activation function hardswish that can reduce the memory access cost MAC is used to transform the lightweight neural network structure, the computing cost can be further reduced on the original basis. This is a new attempt and will surely have a wide range of application value. Summary of the invention

[0005] The purpose of the present invention is to provide a road type recognition method based on an improved lightweight neural network structure, which uses the attention mechanism CBAM and the nonlinear activation function hardswish to transform the basic unit of the ShuffleNet V2 network to solve the problem of improving network performance while maintaining a small amount of parameter calculation.

[0006] The objective of the present invention is achieved through the following technical solutions:

[0007] A road type recognition method based on an improved lightweight neural network structure comprises the following steps:

[0008] Step A: collect pictures of three typical road surfaces, and build training sets and test sets of different road surface types after data enhancement to train the neural network model;

[0009] Step B, building a lightweight neural network ShuffleNet V2 backbone model;

[0010] Step C, introduce the attention mechanism CBAM and the nonlinear activation function hardswish to transform the basic unit of the ShuffleNet V2 network to obtain an improved lightweight neural network structure.

[0011] Further, step A comprises the following steps:

[0012] Step A1, based on different types of typical road surfaces, a road surface image is collected by a camera as an input of a lightweight neural network;

[0013] Step A2, by performing data enhancement for each typical road type, randomly adjusting the brightness of the road image, randomly rotating or shearing, and randomly translating a number of pixels of the image, to achieve data enhancement for the training set and the test set;

[0014] Step A3, divide the samples obtained after data enhancement for each road surface, randomly select data from them to form a training set for neural network training, and the remaining data to form a test set for verifying the recognition accuracy of the neural network; the three-channel color image obtained after the image is processed by PyCharm's RGB is used as the input of the lightweight neural network model, and the neural network is trained by supervised learning; after several rounds of training, the accuracy of the neural network model after training is judged by the basically converged cross entropy loss function curve and the accuracy curve of the verification set; the model calculation efficiency and model size are judged by the number of floating-point operations FLOPs, memory access cost MAC and model parameter quantity.

[0015] Furthermore, in step A1, asphalt pavement, cement pavement and gravel pavement are selected to correspond to flexible pavement, rigid pavement and special pavement respectively.

[0016] Furthermore, in step A3, the data volume ratio of the training set and the test set is 4:1, and the proportion of each road surface image to the total number of samples is the same.

[0017] Further, step B comprises the following steps:

[0018] Step B1, using ShuffleNet V2 as the neural network backbone model for transformation;

[0019] Step B2, the ShuffleNet V2 network designer designs the basic network unit of ShuffleNet V2 according to four lightweight network design standards;

[0020] Step B3, based on the four criteria, the basic units of the ShuffleNet V2 network are designed as follows: Figure 1 shown.

[0021] Further, in step B3, the design standard is:

[0022] Step B31, keep the number of channels of input and output in the convolutional layer as similar as possible to minimize MAC;

[0023] Step B32, reducing the use of group convolution operation GConv;

[0024] Step B33, reducing network branches in the network structure;

[0025] Step B34, reduce element-by-element operations.

[0026] Further, step C comprises the following steps:

[0027] Step C1, replace the MLP with a one-dimensional convolutional layer with adaptively changing size to construct the CBAM module; use the CBAM attention mechanism to transform the basic unit of the ShuffleNet V2 network;

[0028] Step C2, using hardswish function instead of ReLU function to transform the basic unit of ShuffleNet V2 network, and using it to build a neural network to obtain an improved lightweight neural network structure;

[0029] Step C3, training the lightweight neural network model with the constructed training set and test set of the three road type pictures to obtain a lightweight neural network, and using the improved lightweight neural network to perform image recognition of road type.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] 1. The present invention is based on an improved road type recognition method based on a lightweight neural network structure, and constructs a neural network based on deep learning that can recognize road types, providing an optimization method for vehicle intelligent control;

[0032] 2. The present invention uses the lightweight neural network ShuffleNet V2 to replace the traditional convolutional neural network, reducing the number of model parameters and improving the efficiency of model computing resource utilization;

[0033] 3. The present invention introduces the attention mechanism CBAM to transform the basic unit of the network, which improves the network performance while maintaining a small amount of parameter calculation;

[0034] 4. The present invention uses the hardswish activation function to replace the original activation function ReLU of the basic unit of the network, which reduces MAC while enabling the activation function to have a certain spatial modeling capability. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0036] Figure 1 is the network unit of ShuffleNet V2; Figure 1 a is the basic unit of ShuffleNet V2 network, Figure 1 b is the basic unit of the spatial downsampling ShuffleNet V2 network, DWConv is the depth convolution, GConv is the group convolution, and BN is the regularization process;

[0037] Figure 2 is the channel attention module and spatial attention module of CBAM, where Figure 2 a is the channel attention module, Figure 2 b is the spatial attention module;

[0038] Figure 3 It is the improved ShuffleNet V2 network basic unit, in which, Figure 3 a is the improved ShuffleNet V2 network basic unit, Figure 3 b is the basic unit of the ShuffleNet V2 network with improved spatial downsampling. DETAILED DESCRIPTION

[0039] The present invention will be further described below in conjunction with embodiments:

[0040] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0041] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0042] The present invention proposes a road type recognition method based on an improved lightweight neural network structure. Firstly, pictures of three typical roads are collected. After data enhancement, training sets and test sets of different road types are constructed, and the road surface grades of these roads can be determined according to experiments, so as to apply the recognition results to actual control situations in the future. Secondly, a representative ShuffleNet V2 network model with high development potential is adopted to establish a lightweight neural network backbone model. Finally, the attention mechanism CBAM and the new nonlinear activation function hardswish are introduced to transform the basic unit of the ShuffleNet V2 network, so as to obtain a network model with small computational complexity, high computational efficiency and better performance.

[0043] The present invention is based on an improved lightweight neural network structure and a road type recognition method, comprising the following steps:

[0044] Step 1: Collect pictures of three typical road types, and build training sets and test sets of different road types after data enhancement to train the neural network model.

[0045] Step 2: Build the lightweight neural network ShuffleNet V2 backbone model.

[0046] Step 3: Introduce the attention mechanism CBAM and the nonlinear activation function hardswish to transform the basic unit of the ShuffleNet V2 network to obtain an improved lightweight neural network structure.

[0047] Specifically, step 1 includes the following steps:

[0048] Step 11, according to different types of typical road surfaces, the road surface images collected by the camera can be used as the input of the lightweight neural network. The present invention selects asphalt road surface, cement road surface and gravel road surface to correspond to flexible road surface, rigid road surface and special road surface respectively.

[0049] Step 12, in order to improve the generalization of the neural network, that is, the neural network needs to have the ability to accurately predict data other than training data, an effective measure is to perform data enhancement on the training set and test set. When there is less training data, a larger number of training sets and test sets can be obtained by performing data enhancement on each typical road type. In addition, the main measures for data enhancement are to randomly adjust the brightness of the road surface image, randomly rotate or cut, and randomly translate a number of pixels of the image, which makes the training data more suitable for actual application. At the same time, it broadens the recognition range of the neural network.

[0050] Step 13, after data enhancement, several samples are obtained for each road surface. Then, these samples are divided, and data are randomly selected from them to form a training set for neural network training, and the remaining data form a test set to verify the recognition accuracy of the neural network. Among them, the data volume ratio of the training set and the test set is 4:1, and the proportion of each road surface image to the total number of samples is the same. Then, these pictures are processed by PyCharm's RGB to obtain three-channel color images as the input of the lightweight neural network model, and the neural network is trained using supervised learning. Finally, after several rounds of training, the accuracy of the neural network model after training is judged by the basically converged cross entropy loss function curve and the accuracy curve of the verification set; the model calculation efficiency and model size are judged by the number of floating-point operations FLOPs, memory access cost MAC and model parameter quantity.

[0051] Specifically, step 2 includes the following steps:

[0052] Step 21, typical lightweight neural networks include SqueezeNet, Xception, ShuffleNet series and MobileNet series. ShuffleNet V2 is an upgraded version of ShuffleNet V1 proposed by Ma. It is based on the standards of channel shuffling and four lightweight neural network designs, and its accuracy is better than ShuffleNetV1 and MobileNet V2 under the same computational complexity. Moreover, compared with MobileNet V3 (introducing SE attention mechanism), the basic structure of ShuffleNet V2 does not introduce the attention mechanism, and the nonlinear activation function still uses the traditional ReLU function, but its accuracy is not much different from that of MobileNet V3. Therefore, ShuffleNet V2 has better development potential, and the present invention also uses ShuffleNet V2 as the backbone model of the neural network for transformation.

[0053] Step 22, the ShuffleNet V2 network designer designs the basic network unit of ShuffleNet V2 according to four lightweight network design standards. The design standards are as follows:

[0054] 1. Keep the number of channels of the input and output of the convolutional layer as similar as possible to minimize MAC.

[0055] Taking one-dimensional convolution as an example, the number of floating-point operations FLOPs for one-dimensional convolution is:

[0056] B=hwc1c2 (1)

[0057] In formula (1), c1 is the number of input channels; c2 is the number of output channels; h and w are the sizes of the feature map.

[0058] Then the memory access cost MAC is:

[0059] MAC=hw(c1+c2)+c1c2 (2)

[0060] According to the mean inequality, we can get the inequality of memory access cost MAC with respect to the number of floating point operations FLOPs:

[0061]

[0062] In formula (3), B is the number of floating point operations FLOPs.

[0063] MAC has a lower bound determined by FLOPs, and the inequality is equal when c1 and c2 are equal.

[0064] 2. Reduce the use of group convolution operations GConv. Too much group convolution will increase MAC.

[0065] Grouped convolution operations are very common in modern neural network structures. They reduce computational complexity (FLOPs) by making dense convolutions between all channels sparse. Although it can increase network capacity under given FLOPs and allow the use of more channels, the increase in the number of channels will also increase MAC. Assuming that the number of groups in the grouped convolution is g, the relationship between MAC and FLOPs in the case of one-dimensional convolution is:

[0066]

[0067] In formula (4), g is the number of groups in the grouped convolution.

[0068] From formula (4), we know that when the number of floating-point operations B is constant, MAC increases with the increase of g, which indicates that the more group convolution is used, the lower the computational efficiency of the network.

[0069] 3. Reduce the network branches in the network structure. The more network branches there are, the lower the computing efficiency.

[0070] In the structure of the Inception network system, most of the basic network modules adopt a multi-branch structure. The more network branches are used, the more fragmented the network structure will be. Although it has been proven that this fragmented network structure can improve the recognition accuracy of the overall network, it will reduce the network parallelism and computing efficiency because this structure will introduce additional computing resource consumption, such as convolution kernel startup and synchronization consumption.

[0071] 4. Reduce element-by-element operations.

[0072] In the lightweight neural network model, element-by-element operations include ReLU, AddTensor, AddBias, etc. Although these operations have small FLOPs, their MAC is large. Therefore, too many of these operators are deleted in the network structure of ShuffleNet V2. However, in each basic unit of ShuffleNet V2, the nonlinear activation function ReLU is inevitably used at the end, which provides an idea for the present invention to use other nonlinear activation functions with lower MAC to transform ShuffleNet V2.

[0073] Step 23, based on the four criteria, the basic units of the ShuffleNet V2 network are designed as follows: Figure 1 shown.

[0074] Specifically, step 3 includes the following steps:

[0075] Step 31, reference to the attention mechanism.

[0076] In order to improve the performance of convolutional neural networks, current research mostly focuses on three factors of the network: depth, width, and cardinality. Studies have shown that networks with deeper depth, wider width, and larger cardinality have higher performance. However, the performance improvement of the model by these three factors is based on the expansion and stacking of the unit structure, which will increase the computational burden. Therefore, the present invention introduces an attention mechanism with less computational effort.

[0077] In the human visual system, an important mechanism is attention. Due to energy limitations, humans will not try Figure 1The entire scene received is processed at once, so humans use local attention mechanisms to selectively focus on the salient parts of the information. The main attention mechanisms in convolutional neural networks are SENet, ECANet, and CBAM. SENets use Squeeze and Excitation modules, respectively, to obtain the importance of each channel of the feature map using global average pooling and two fully connected layers. Then, this importance is used to assign weights to each feature channel, so that the network can focus on certain important channels with higher weights and suppress feature channels that have less effect on the current task. ECANet is an improved version of SENet. Its attention mechanism module uses one-dimensional convolution directly after the global average pooling layer, removing two fully connected layers, which avoids dimensionality reduction and effectively captures cross-channel interactions. However, in general, SENet and ECANet only focus on analysis in the channel domain, thus ignoring the problem of where to focus attention on feature maps in space (the two dimensions of image recognition convolutional neural networks are the spatial dimensions). CBAM introduces the spatial attention module, which is combined with the channel attention module to arrange the channels in the order of the channel priority space (experiments have shown that sequential arrangement is better than parallel arrangement). In addition, the global maximum pooling is combined with the global average pooling to construct the CBAM attention module. However, the channel attention module of the CBAM attention module uses MLP (multi-layer perceptron), and still uses two fully connected layers to build the relationship between channels.

[0078] Based on the inspiration of ECANet, the present invention replaces MLP with a one-dimensional convolutional layer with adaptively variable size to reduce the impact of dimensionality reduction. Thus, the constructed CBAM module is as follows: Figure 2 shown. Figure 2 In the figure, F is the input feature map; F' is the output feature map of the channel attention mechanism; MaxPool is the global maximum pooling; AvgPool is the global average pooling; sigmoid is the normalization function; Cat is the dimension stacking; Mc is the channel attention weight after normalization; Ms is the spatial attention weight after normalization.

[0079] Since the channel attention module only uses pooling function, one-dimensional convolution layer and normalization processing, its parameter calculation amount is very small, and the spatial attention module can complement the channel attention module, achieving the improvement of network performance while maintaining a small parameter calculation amount. Therefore, the present invention uses the CBAM attention mechanism to transform the basic unit of the ShuffleNet V2 network.

[0080] Step 32, reference to hardswish.

[0081] In convolutional neural networks, in order to enhance the nonlinear ability of the model, a nonlinear activation layer is usually added after a convolution layer. The activation function often uses the ReLU function, which has a fast calculation speed and speeds up the training of the network. However, from a spatial perspective, the ReLU function performs a single operation on all pixels and lacks spatial modeling capabilities. In addition, the ReLU function increases the memory access cost MAC, so it is necessary to use other nonlinear activation functions instead.

[0082] In the design of MobileNet V3, the nonlinear activation function hardswish is introduced to replace ReLU. It is an upgraded version of the swish activation function. The expression of hardswish is as follows:

[0083]

[0084] In formula (5), ReLU6 is used instead of a custom clipping constant.

[0085] First, from the deployment point of view, the optimization implementation of ReLU6 can be implemented on almost all software and hardware architectures; second, in practice, hardswish implements a segmentation function, and the product of the ReLU6 function and x can enable it to have a certain spatial modeling ability while keeping the MAC small. Therefore, the present invention uses the hardswish function to replace the ReLU function to transform the basic unit of the ShuffleNet V2 network.

[0086] Combining the CBAM attention mechanism and the hardswish activation function, the basic units of the modified ShuffleNet V2 network are as follows: Figure 3 shown.

[0087] The improved ShuffleNet V2 network basic unit is used to build a neural network, and an improved lightweight neural network structure can be obtained. Then, the lightweight neural network model is trained with the constructed training set and test set of three road type images to obtain a lightweight neural network with small parameters, high computational efficiency and high accuracy in road type image recognition. Finally, using the improved lightweight neural network for road type image recognition will be beneficial to active suspension control with limited computing resources.

[0088] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A road type recognition method based on an improved lightweight neural network structure, characterized in that: The following steps are involved: Step A: collect pictures of three typical road surfaces, and build training sets and test sets of different road surface types after data enhancement to train the neural network model; Step B, building a lightweight neural network ShuffleNet V2 backbone model; Step C, introduce the attention mechanism CBAM and the nonlinear activation function hardswish to transform the basic unit of the ShuffleNet V2 network to obtain an improved lightweight neural network structure.

2. The method for identifying road types based on an improved lightweight neural network structure according to claim 1, characterized in that: Step A comprises the following steps: Step A1, based on different types of typical road surfaces, a road surface image is collected by a camera as an input of a lightweight neural network; Step A2, by performing data enhancement for each typical road type, randomly adjusting the brightness of the road image, randomly rotating or shearing, and randomly translating a number of pixels of the image, to achieve data enhancement for the training set and the test set; Step A3, divide the samples obtained after data enhancement for each road surface, randomly select data from them to form a training set for neural network training, and the remaining data to form a test set for verifying the recognition accuracy of the neural network; the three-channel color image obtained after the image is processed by PyCharm's RGB is used as the input of the lightweight neural network model, and the neural network is trained by supervised learning; after several rounds of training, the accuracy of the neural network model after training is judged by the basically converged cross entropy loss function curve and the accuracy curve of the verification set; the model calculation efficiency and model size are judged by the number of floating-point operations FLOPs, memory access cost MAC and model parameter quantity.

3. The method for identifying road types based on an improved lightweight neural network structure according to claim 2, characterized in that: Step A1, selecting asphalt pavement, cement pavement and gravel pavement corresponding to flexible pavement, rigid pavement and special pavement respectively.

4. The method for identifying road types based on an improved lightweight neural network structure according to claim 2, characterized in that: In step A3, the data volume ratio of the training set and the test set is 4:1, and the proportion of each road surface image to the total number of samples is the same.

5. The road type identification method based on an improved lightweight neural network structure according to claim 1 is characterized in that: Step B comprises the following steps: Step B1, using ShuffleNet V2 as the neural network backbone model for transformation; Step B2, the ShuffleNet V2 network designer designs the basic network unit of ShuffleNet V2 according to four lightweight network design standards; Step B3, based on the four criteria, the basic unit of the ShuffleNet V2 network is designed as shown in Figure 1.

6. The method for identifying road types based on an improved lightweight neural network structure according to claim 1, characterized in that: Step B3, design criteria are: Step B31, keep the number of channels of input and output in the convolutional layer as similar as possible to minimize MAC; Step B32, reducing the use of group convolution operation GConv; Step B33, reducing network branches in the network structure; Step B34, reduce element-by-element operations.

7. The method for identifying road types based on an improved lightweight neural network structure according to claim 1, characterized in that: Step C comprises the following steps: Step C1, replace the MLP with a one-dimensional convolutional layer with adaptively changing size to construct the CBAM module; use the CBAM attention mechanism to transform the basic unit of the ShuffleNet V2 network; Step C2, using hardswish function instead of ReLU function to transform the basic unit of ShuffleNet V2 network, and using it to build a neural network to obtain an improved lightweight neural network structure; Step C3, training the lightweight neural network model with the constructed training set and test set of the three road type pictures to obtain a lightweight neural network, and using the improved lightweight neural network to perform image recognition of road type.