Risky road scene recognition method based on multi-stage attention deep learning

Through the multi-stage deep learning method of attention, deep convolutional neural network is used to identify risky road scenarios, solving the problem of insufficient recognition of risky road scenarios in the existing technology, and achieving efficient and low-cost risky road scenario recognition, which is suitable for the field of traffic information technology.

CN114049532BActive Publication Date: 2025-05-23RES INST OF HIGHWAY MINIST OF TRANSPORT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111309230.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-06
Publication Date
2025-05-23
Estimated Expiration
2041-11-06

AI Technical Summary

Technical Problem

The existing technology mainly warns driving speed and collisions during driving, and fails to effectively identify and warn of risky road scenarios, such as road bends, bridges, tunnels, ramps, adjacent water and adjacent cliffs, resulting in high frequency and harm of traffic accidents.

Method used

Using a method based on multi-stage attention deep learning, images are collected through vehicle cameras, and deep convolutional neural networks with multi-stage attention mechanisms are used to identify risky road scenarios, including convolutional networks and residual networks, and combined with channel and spatial attention mechanisms to extract and classify image features.

Benefits of technology

The recognition efficiency and accuracy of risky road scenarios are improved, and efficient identification of 6 common accident scenarios are achieved, with an average accuracy rate of 91.90%, strong adaptability and low cost, and only image acquisition equipment is required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114049532B_ABST
    Figure CN114049532B_ABST
Patent Text Reader

Abstract

The present invention discloses a risk road scene recognition method based on multi-stage attention deep learning, wherein the method includes proposing a risk road scene recognition model structure based on a multi-stage attention mechanism deep convolutional neural network, constructing a risk road scene data set, and utilizing transfer learning to construct a risk road scene recognition model based on a multi-stage attention mechanism deep convolutional neural network; collecting scene image data in the vehicle's forward direction in real time, and utilizing the constructed model to automatically calculate the category likelihood value of the risk road scene, and determining the category of the scene by obtaining the maximum likelihood value, which can effectively determine the type of risk scene that the vehicle passes through, thereby improving the risk prevention and control level of the driving vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of traffic information technology, and specifically relates to a risky road scene recognition method based on multi-stage attention deep learning. Background Art

[0002] With the popularization of driving recorders, image data in front of the vehicle can be effectively obtained through driving recorders during driving. By processing these image data through computer vision technology, risky road scenes can be effectively identified, early warning can be achieved, and the level of driving risk prevention and control can be improved.

[0003] At present, the risk warnings taken during driving are mainly for driving speed and collision warnings. However, according to the statistics of traffic accidents, in risky road scenes, such as highway curves, highway bridges, highway tunnels, highway ramps, highways near water and cliffs, etc., the frequency of traffic accidents is relatively higher, and the harm caused by traffic accidents is more serious. Therefore, it is of great significance to identify risky road scenes in advance. Summary of the invention

[0004] In view of the above, the present invention proposes a risky road scene recognition method based on multi-stage attention deep learning. By collecting images of the vehicle's travel direction at a certain frequency and identifying risky road scenes through the deep learning model proposed by the present invention, it is applied to driving risk prevention and control, thereby improving the ability to prevent risky roads.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A risky road scene recognition method based on multi-stage attention deep learning includes the following steps:

[0007] (1) When the car is driving, the camera collects the image in front of the vehicle in real time and uses the video decoder to convert the image into digital image information.

[0008] (2) Using the linear interpolation method, the digital image information collected in step (1) is transformed into a three-dimensional tensor of size 299×299×3,

[0009] (3) Establish a deep convolutional neural network with a multi-stage attention mechanism and use samples to train the network to obtain a risky road scene recognition model.

[0010] (4) The three-dimensional tensor obtained by the transformation in step (2) is input into the risk road scene recognition model trained in step (3) to obtain the recognition likelihood values ​​of six scene types: highway curves, highway bridges, highway tunnels, highway ramps, highways adjacent to water, and highways adjacent to cliffs.

[0011] (5) The category with the largest recognition likelihood value is taken as the recognized scene category.

[0012] Furthermore, step (1) includes: while the car is driving, the image in front of the car is collected once every 1 second by four cameras of the on-board driving recorder.

[0013] Furthermore, the deep convolutional neural network in step (3) is composed of an input layer, a convolutional network, a residual network and a fully connected layer connected in sequence, and an attention network is added to the convolutional network stage and the residual network respectively.

[0014] Furthermore, the convolutional network in the deep convolutional neural network in step (3) is composed of three convolutional blocks connected in sequence. Since the highway scene is characterized by an outdoor scene, it also contains a large area of ​​semantically consistent scene areas and detailed road scene risk features, such as a large area of ​​sky area and road surface area, as well as curve features and tunnel features with detailed features. Therefore, when designing the convolution kernel, it is necessary to consider that the size of the convolution kernel is moderate and has differences. The convolution kernel sizes of the three convolutional blocks are 3×3, 3×3, and 2×2, respectively, and the activation function is PReLU (ParametricRectified Linear Unit) function.

[0015]

[0016] Where c is the channel, x c is the feature of the image, a c is a parameter, and batch normalization is adopted in all three convolution blocks.

[0017] Furthermore, the attention network of the convolutional network stage in the deep convolutional neural network in step (3) is connected to the channel attention network and the spatial attention network in sequence after the three-layer convolution block, and then connected to the two-dimensional spatial attenuation layer SpatialDropout2D and the convolution layer with a convolution kernel size of 1×1 and an activation function of PReLU, wherein the output of the attention network is added to the output of the three convolution blocks, and the output result is sequentially input into the two-dimensional spatial attenuation layer SpatialDropout2D and the convolution layer with a convolution kernel size of 1×1 and an activation function of PReLU as filtering of the output features of the attention mechanism.

[0018] Furthermore, the residual network includes a 7-layer sequential network structure, wherein the residual network structure of the first 6 layers is: the input and output of each layer of the residual network are superimposed and input into the maximum pooling layer MaxPooling2D, and this is repeated 6 times to form the first 6 layers of the residual network structure, and the residual network structure of the 7th layer is to superimpose the output of the attention network and the output of the 7th layer of the residual network and then input into the global maximum pooling layer, and the output value is then input into the fully connected layer, and after generating a 6-dimensional vector, batch normalization is used to obtain features, and then a random decay algorithm is used to select 60% of the outputs of the neurons to form a feature vector, and finally a softmax function is used. As a classifier, x is the output feature of the fully connected layer, θ is the classifier parameter obtained through training, and the classification likelihood values ​​of the six classes are obtained by taking the values ​​of the six softmax functions. Then, the class with the maximum likelihood value is taken as the final classification result of the sample.

[0019] Furthermore, the deep convolutional neural network is trained by the transfer learning method. The model parameters are initialized by the parameters obtained by training the deep convolutional neural network InceptionResnet V2 network model on the ImageNet dataset. The model training is optimized by constructing a sample library with various road risk scenarios, and the cross entropy is used as the loss function Loss(θ). m=16 is the number of batch samples, which can be reset according to the needs of training and the training hardware conditions. k=6 is the number of scene categories. is the true classification result of sample i for scene category j, x (i) is the characteristic of sample i, is based on the classification parameter θ j The classifier of scene category j is obtained through multiple iterations and convergence to obtain the risk road scene recognition model parameters.

[0020] Furthermore, the attention mechanism in the deep convolutional neural network in step (3) consists of a channel attention mechanism and a spatial attention mechanism, and the output of the channel attention mechanism is the input of the spatial attention mechanism.

[0021] Furthermore, the specific process of training the deep convolutional neural network in step (3) is as follows: first, various parameters, learning rate, optimization algorithm and maximum number of iterations in the temporal convolutional neural network are initialized; then, the image samples are batched into a four-dimensional tensor and input into the convolutional neural network for training, and the cross entropy classification error function L between the output result of the convolutional neural network and the actual scene classification number is calculated, and then the parameters in the entire neural network are continuously updated through the optimization algorithm until the cross entropy classification error function L converges or reaches the maximum number of iterations, thereby completing the training and obtaining the prediction model.

[0022] Furthermore, the samples in step (3) are obtained in the following manner: road scene images are collected by a driving recorder, and then the collected images are manually labeled with six risky road scenes: highway curves, highway ramps, highways adjacent to water, adjacent to cliffs, highway tunnels, and highway bridges, and these labeled image data sets are used as samples.

[0023] Through the above technical solutions, the identification method proposed by the present invention can have the following advantages:

[0024] (1) The present invention proposes a risky road scene recognition model based on a deep convolutional neural network with a multi-stage attention mechanism. This recognition model has high recognition efficiency for road risk scenes.

[0025] (2) Based on the 2016 Traffic Accident Statistics Yearbook, the present invention analyzed six road risk scenarios where accidents are prone to occur, and constructed recognition models for the six road scenarios where accidents often occur, namely, highway curves, highway ramps, highways near water, highways near cliffs, highway tunnels, and highway bridges.

[0026] (3) When implementing the method proposed in the present invention, fewer sensors are used, the cost is low, and the deployment is simple. Only an image acquisition device is required, and risky road scene recognition can be achieved using a single-frame image.

[0027] (4) The proposed method is highly robust and can resist scale changes and illumination. It is tested on the road risk scenario dataset we constructed, and the average accuracy (AP) reaches 91.90%, and the average AUC under 6 scenarios is 0.972892. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A flow chart of the identification method proposed by the present invention;

[0029] Figure 2 It is a schematic diagram of the structure of a deep convolutional neural network based on a multi-stage attention mechanism;

[0030] Figure 3 Schematic diagram of the residual network structure;

[0031] Figure 4 Schematic diagram of the channel attention network structure;

[0032] Figure 5 Schematic diagram of the spatial attention network structure;

[0033] Figure 6 This is the ROC curve tested on road scene images. DETAILED DESCRIPTION

[0034] In order to describe the present invention more clearly and completely, the technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. However, those skilled in the art should know that the following embodiments are not the only limitations on the technical solution of the present invention, and any equivalent changes or modifications made under the spirit of the technical solution of the present invention should be deemed to belong to the protection scope of the present invention.

[0035] like Figure 1 ,The risk road scene recognition method based on a multi-stage attention mechanism deep ,convolutional neural network in the present invention includes two parts: a training module and a prediction module.

[0036] In the training module, a risky road scene image training set is first constructed, and then the deep convolutional neural network model with a multi-stage attention mechanism constructed according to the present invention is trained using this image training set, thereby obtaining a risky road scene recognition model. In the prediction module, a camera is first used to collect images of the vehicle's travel direction, and the collected images are input into the risky road scene recognition model trained in the training module to obtain the likelihood values ​​of different risky scene categories, and the category with the largest likelihood value is the identified risky scene category.

[0037] Specifically, the present invention includes the following parts:

[0038] (1) Building a network model framework

[0039] The core part of the technical framework of the present invention is that the attention mechanism is used in the convolutional network and residual network stages in the risky road scene recognition model. Adding the attention network in the convolutional stage can improve the significance of the effective features of the image, and adding attention in the residual network stage can also effectively improve the significance of the effective features of the residual network. The network model architecture is as follows Figure 2 shown.

[0040] The deep convolutional neural network with multi-stage attention mechanism is composed of an input layer, a convolutional network, a residual network and a fully connected layer, which are connected in sequence. The attention network is added to the convolutional network stage and the residual network respectively.

[0041] The convolution network in the deep convolutional neural network with a multi-stage attention mechanism is composed of three convolution blocks connected in sequence. Since the highway scene is characterized by an outdoor scene, it also contains a large area of ​​semantically consistent scene areas and detailed road scene risk features, such as a large area of ​​sky area and road pavement area, as well as curve features and tunnel features with detailed features. Therefore, when designing the convolution kernel, the size of the convolution kernel should be moderate and different. Therefore, the convolution kernel sizes of these three convolution blocks are 3×3, 3×3, and 2×2, respectively. In this way, the granularity of the convolution feature will not be too small due to the convolution kernel being too small, which will reduce the computational efficiency and be unfavorable for the feature extraction of large-area continuous scenes. At the same time, the convolution kernel will not be too large, resulting in the granularity of the convolution feature being too large to ignore the detailed features. The activation function is PReLU (ParametricRectified Linear Unit) function

[0042]

[0043] Where c is the channel, x c is the feature of the image, a c is a parameter, and batch normalization is adopted in all three convolution blocks.

[0044] Furthermore, the attention network in the convolutional network stage of the deep convolutional neural network of the attention mechanism in this stage is connected to the channel attention network and the spatial attention network in sequence after the three-layer convolution block, and then connected to the two-dimensional spatial attenuation layer SpatialDropout2D and the convolution layer with a convolution kernel size of 1×1 and an activation function of PReLU, wherein the output results are obtained by adding the output of the attention network and the output of the three convolution blocks, and then sequentially input into the two-dimensional spatial attenuation layer SpatialDropout2D and the convolution layer with a convolution kernel size of 1×1 and an activation function of PReLU as filtering of the output features of the attention mechanism.

[0045] Furthermore, the residual network stage includes 7 layers of sequentially connected residual networks, where the residual network structure of the first 6 layers is: after the residual network, the input of the residual network and the output of the residual network are superimposed, and then input into the maximum pooling layer MaxPooling2D. The residual network structure of the 7th layer is to superimpose the output of the attention network and the output of the 7th layer residual network and then input into the global maximum pooling layer. The output value is input into the fully connected layer. After the vector is generated, batch normalization is used to obtain the features, and then the random decay algorithm is used to select 60% of the outputs of the neurons to form the feature vector. Finally, 6 softmax functions are used. As a classifier, θ is the classifier parameter obtained through training. The classification likelihood value of each class is obtained by taking the maximum value of the softmax function, and then the category with the maximum likelihood value is taken as the final classification result of the sample.

[0046] The attention mechanism is implemented through the attention network, including the channel attention network (such as Figure 4 ), and the spatial attention network (as Figure 5 As shown in Figure 4 The number of neurons in the channel-shared fully connected layer 1 and the channel-shared fully connected layer 2 of the channel attention network in is consistent with the number of channels of the input feature. The calculation method of the channel attention network is shown in formula (1):

[0047] M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))(1)

[0048] In formula (1), σ is the activation function, AvgPool(F) is the average pooling layer function, F is the input feature, MaxPool(F) is the maximum pooling layer, and MLP is a multi-layer perceptron.

[0049] The calculation method of the spatial attention network is shown in formula (2):

[0050] M s (F) = σ(f 7×7 ([AvgPool(F);MaxPool(F)]))(2)

[0051] f 7×7 Represents a convolution operation with a filter size of 7×7.

[0052] The activation function used in the activation layer of the channel attention network and the spatial attention network is the hard sigmoid function, as shown in formula (3).

[0053]

[0054] Here, x is the input feature.

[0055] (2) Constructing a risky road scenario training dataset

[0056] First, the images of road scenes are collected through the driving recorder, and then the collected images are labeled with six risky road scenes: highway curves, highway ramps, highways near water, near cliffs, highway tunnels, and highway bridges through manual annotation, and these labeled image data sets are used as training data sets. Among them, the six road risk scenes of highway curves, highway ramps, highways near water, near cliffs, highway tunnels, and highway bridges are summarized by the inventors after analysis based on the 2016 traffic accident statistics annual report.

[0057] (3) Training a risky road scene recognition model based on a multi-stage attention mechanism deep convolutional neural network

[0058] The training method in the training module is based on transfer learning. The model parameters of the deep convolutional neural network InceptionResnet V2 that has been trained on the ImageNet dataset are used for parameter initialization to improve the convergence speed of model training. The number of batch input samples is 16, and the RAdam function is used as the optimization function. Using the constructed risky road scene dataset, the proposed risky road scene recognition model based on the multi-stage attention mechanism deep convolutional neural network is learned.

[0059] More specifically, the training of the deep convolutional neural network based on the multi-stage attention mechanism adopts the transfer learning method. The model parameters are initialized using the parameters obtained by training the deep convolutional neural network InceptionResnet V2 network model on the dataset ImageNet. The model training is optimized through the constructed sample library with various road risk scenarios, and the cross entropy is used as the loss function Loss(θ).

[0060]

[0061] m=16 is the number of batch samples, which can be reset according to the needs of training and the training hardware conditions. k=6 is the number of scene categories. is the true classification result of sample i for scene category j, x (i) is the characteristic of sample i, is based on the classification parameter θ j classifier for scene category j.

[0062] The risky road scene recognition model parameters are obtained through multiple iterations and convergence.

[0063] The specific process of training the convolutional neural network based on the multi-stage attention mechanism is as follows: first, initialize the various parameters, learning rate, optimization algorithm and maximum number of iterations in the temporal convolutional neural network; then batch the image samples into a four-dimensional tensor and input it into the convolutional neural network for training, and calculate the cross entropy classification error function L between the output of the convolutional neural network and the actual scene classification number, and then continuously update the parameters in the entire neural network through the optimization algorithm until the error function L converges or reaches the maximum number of iterations, thereby completing the training and obtaining the prediction model.

[0064] (4) Practical prediction and experimental evaluation

[0065] The data set used in the test sample of this implementation is to collect scene images in the direction of vehicle travel, and to form an RGB image with a length and width of 299 pixels through decoding transformation.

[0066] The image is input into the trained risky road scene recognition model based on a multi-stage attention mechanism deep convolutional neural network, and the recognition likelihood values ​​of six types of risky road scenes are obtained at the same time: highway curves, highway ramps, highways near water, highways near cliffs, highway tunnels, and highway bridges. The category with the largest likelihood value is the final recognition result.

[0067] We use the average accuracy (AP) and AUC (Area Under ROC Curve) values ​​as quantitative indicators for evaluating models, and use the ROC curve to intuitively evaluate the model. Figure 6 shown.

[0068] Evaluation indicators Accuracy Average precision AP 0.9189873417721519 Identifying the AUC value of highway adjacent to water scenes 0.9431818181818181 Identify the AUC value of highway intersection scene 0.9995715509854327 AUC value for identifying highway ramp scenes 0.9589182493806772 Identify the AUC value of highway bridge scene 0.9910211267605634 Identify the AUC value of the highway tunnel scene 0.9902409134385461 AUC value for identifying highway curve scenes 0.9930758620689656 AUC value for identifying cliff-adjacent road sections 0.9342373312961548 Average AUC 0.972892

[0069] Compared with the prior art, the present invention has the following advantages:

[0070] (1) Through the analysis of traffic accident statistics, we obtained six scenarios where traffic accidents are prone to occur, and designed a deep learning recognition model that can uniformly identify six road scenes where accidents often occur: road curves, road ramps, roads near water, roads near cliffs, road tunnels, and road bridges.

[0071] (2) The model is composed of a convolutional network and a residual network, and an attention network is added in the convolution stage and the residual network stage respectively, which improves the significance of image features that are consistent with the scene semantics.

[0072] (3) Adding attention networks in the convolutional and residual network stages improves recognition efficiency and real-time performance without significantly increasing the network size.

[0073] (4) The model is trained on samples collected from driving recorders, which does not require high video image quality and has wide adaptability.

Claims

1. A risky road scene recognition method based on multi-stage attention deep learning, comprising the following steps: (1) When the car is driving, the camera collects the image in front of the vehicle in real time and converts the image into digital image information using a video decoder. (2) using a linear interpolation method to transform the digital image information collected in step (1) into a three-dimensional tensor of size 299×299×3, (3) Establish a deep convolutional neural network with a multi-stage attention mechanism and use samples to train the network to obtain a risky road scene recognition model. (4) inputting the three-dimensional tensor transformed in step (2) into the risk road scene recognition model trained in step (3) to obtain recognition likelihood values ​​of six scene types: highway curves, highway bridges, highway tunnels, highway ramps, highways adjacent to water, and highways adjacent to cliffs. (5) The category with the largest recognition likelihood value is taken as the recognized scene category; in, The deep convolutional neural network of the multi-stage attention mechanism is composed of an input layer, a convolutional network, a residual network and a fully connected layer, which are connected in sequence, and an attention network is added to the convolutional network and the residual network respectively; The convolutional network is composed of three convolutional blocks connected in sequence, where the convolution kernel sizes of the three convolutional blocks are 3×3, 3×3 and 2×2 respectively, and the activation functions of the three convolutional blocks are all PReLU functions. Where c is the channel, x c is the feature of the image, a c is a parameter, batch normalization is used in all three convolution blocks; The attention network in the convolutional network is connected to the channel attention network and the spatial attention network in sequence after the three convolutional blocks, and then connected to the two-dimensional spatial attenuation layer SpatialDropout2D and the convolution layer with a convolution kernel size of 1×1 and an activation function of PReLU, wherein the output of the attention network is added to the output of the three convolutional blocks, and the output results are sequentially input into the two-dimensional spatial attenuation layer SpatialDropout2D and the convolution layer with a convolution kernel size of 1×1 and an activation function of PReLU as filtering of the output features of the attention mechanism; The residual network includes a 7-layer sequentially connected residual network structure, wherein the residual network structure of the first 6 layers is: the input and output of each residual network layer are superimposed and input into the maximum pooling layer MaxPooling2D, and this is repeated 6 times to form the first 6 layers of residual network structure, and the residual network structure of the 7th layer is to superimpose the output of the attention network and the output of the 7th residual network and then input into the global maximum pooling layer, and the output value is then input into the fully connected layer, and after generating a 6-dimensional vector, batch normalization is used to obtain features, and then a random decay algorithm is used to select 60% of the outputs of neurons to form a feature vector, and finally a softmax function is used. As a classifier, x is the output feature of the fully connected layer, θ is the classifier parameter obtained through training, and the classification likelihood values ​​of the six classes are obtained by taking the values ​​of the six softmax functions. Then, the class with the maximum likelihood value is taken as the final classification result of the sample.

2. The risky road scene identification method according to claim 1, wherein step (1) include: When the car is driving, the four cameras of the vehicle-mounted driving recorder collect the image in front of the car once every 1 second.

3. According to the risky road scene recognition method of claim 1, the deep convolutional neural network based on the multi-stage attention mechanism is trained by using a transfer learning method, and the model parameters of the deep convolutional neural network InceptionResnet V2 that has been trained on the ImageNet dataset are used for parameter initialization. The model training optimization is performed by constructing a sample library with a variety of road risk scenes, and the cross entropy is used as the loss function Loss(θ). Where m = 16 is the number of batch samples, which can be reset according to the needs of training and the training hardware conditions. k = 6 is the number of scene categories. is the true classification result of sample i for scene category j, x (i) is the characteristic of sample i, is based on the classification parameter θ j The classifier of scene category j; The risky road scene recognition model parameters are obtained through multiple iterations and convergence.

4. The risky road scene identification method according to claim 1, Features The attention mechanism in the deep convolutional neural network of the multi-stage attention mechanism in step (3) consists of a channel attention mechanism and a spatial attention mechanism, and the output of the channel attention mechanism is the input of the spatial attention mechanism.

5. The risky road scene identification method according to claim 1, Features: The specific process of training the multi-stage attention mechanism convolutional neural network in step (3) is as follows: first, initializing various parameters, learning rate, optimization algorithm and maximum number of iterations in the multi-stage attention mechanism convolutional neural network; then batching image samples into a four-dimensional tensor and inputting it into the multi-stage attention mechanism convolutional neural network for training, and calculating the cross entropy classification error function L between the output result of the multi-stage attention mechanism convolutional neural network and the actual scene classification number, and then continuously updating the parameters in the entire multi-stage attention mechanism convolutional neural network through the optimization algorithm until the cross entropy classification error function L converges or reaches the maximum number of iterations, thereby completing the training and obtaining the prediction model.

6. The risky road scene identification method according to claim 1, in, The samples in step (3) are obtained in the following manner: collecting road scene images through a driving recorder, and then manually labeling the collected road scene images with labels of six risky road scenes, namely, road curves, road ramps, roads adjacent to water, roads adjacent to cliffs, road tunnels and road bridges, and these labeled image data sets are used as the samples.

Citation Information

Patent Citations

  • Driving scene classification method based on convolution neural network

    CN107609602A

  • Construction method of behavior recognition deep network model and behavior recognition method

    CN111985343A