A forward vehicle recognition and distance detection method based on deep learning
Through a deep learning-based method, combined with a multi-task attention network and a large-core attention mechanism, the problem of insufficient accuracy and positioning performance of forward vehicle identification and distance detection in the prior art is solved, and high-precision recognition and distance measurement in complex environments are achieved.
Patent Information
- Application Number
- CN202210979374.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-08-16
AI Technical Summary
The prior art cannot accurately identify forward vehicles and detect their distances, and the positioning performance is insufficient, especially in complex traffic environments.
Forward vehicle identification and distance detection methods based on deep learning are adopted, including data preprocessing, backbone network construction, object detection subnet construction, depth estimation subnet construction, multi-task attention network and large-nuclear attention mechanism, and K-Means-optimized distance measurement method.
Improves the accuracy and positioning performance of forward vehicle identification and distance detection, and can accurately identify vehicles and distance measurement in complex traffic environments.
Smart Images

Figure CN115424237B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle detection, and in particular relates to a forward vehicle recognition and distance detection method based on deep learning. Background Art
[0002] Forward vehicle target detection and distance detection are an indispensable part of the environmental perception technology of intelligent driving systems. By timely and accurately perceiving environmental information, combining scene data for analysis and decision-making, predicting potential traffic accidents and issuing warnings, the active safety performance of the car can be significantly improved. However, in real-world scenarios, the types, models, and sizes of vehicles vary in complexity, and the varying degrees of occlusion between vehicles, mutual occlusion between vehicles and the environment, and changes in the pitch angle of the road all pose great challenges to forward vehicle recognition and distance detection. Therefore, how to quickly and accurately perform vehicle target recognition and distance detection has become a major difficulty in the current research on intelligent driving systems.
[0003] The premise of environmental perception is to understand the environmental information around the vehicle, and obtaining the distance information of the vehicle ahead is also the key to making decision-making and control. According to the different detection equipment and detection methods, the forward vehicle distance detection methods can be divided into the following categories: millimeter wave radar ranging, laser ranging, ultrasonic ranging, visual ranging, etc. Among these methods, although ultrasonic, millimeter wave radar, lidar, etc. are far superior to visual ranging in ranging accuracy, active sensors based on lidar and millimeter wave radar are expensive, have limited ranging scanning range, and are sensitive to signal interference. In contrast, the method based on visual ranging is low in cost, easy to install, and has a high penetration rate, so the industry prefers the visual ranging method.
[0004] In visual ranging, the ranging method based on monocular vision has the advantages of simple model and low consumption of computing resources. It is the standard configuration of ordinary vehicles and has good application prospects. However, the current traffic environment is changeable and there are many influencing factors. It is difficult to obtain good vehicle detection accuracy by directly applying the target detection algorithm, which brings great challenges to vehicle target detection.
[0005] At present, the mainstream distance detection methods based on monocular vision are generally based on the principle of similar geometry, combining the internal and external parameters of the camera to estimate the distance to the front vehicle. However, this type of method requires the actual height or width of the object to be measured, and the camera parameter acquisition process is complicated and the matching process is complicated. On unstructured roads, the effective distance of the distance measurement is short, and the distance measurement error is large on curves.
[0006] Therefore, there is an urgent need for a method that can accurately identify the vehicle ahead, detect the distance to the vehicle ahead, and perform precise positioning. Summary of the invention
[0007] In view of this, the purpose of the present invention is to provide a forward vehicle recognition and distance detection method based on deep learning. The present invention aims to solve the problem that the existing method cannot accurately detect the forward vehicle and has poor positioning performance.
[0008] To achieve the above object, the present invention provides a forward vehicle recognition and distance detection method based on deep learning, comprising the following steps:
[0009] S1. Obtain the data set required for forward vehicle identification and distance detection, and preprocess the data set;
[0010] S2. Build a backbone network;
[0011] S3. Build the target detection sub-network;
[0012] S4. Build a depth estimation subnetwork;
[0013] S5. Training of forward vehicle recognition and distance detection network based on deep learning;
[0014] S6. Optimize forward vehicle distance detection based on K-Means.
[0015] Furthermore, the data set of step S1 is a KITTI data set, and the KITTI data set includes vehicle training pictures, annotation files, and point cloud files.
[0016] Furthermore, in step S1, the preprocessing steps of the data set are as follows:
[0017] S1.1 Convert point cloud files into depth map labels;
[0018] S1.2 cleans the KITTI dataset processed in step S1.1, filters out the images with incorrect annotations and removes them;
[0019] S1.3 uses K-means k-means clustering algorithm to determine the number of anchor boxes and aspect ratio;
[0020] S1.4 uses 90% of the data set as the training set and the remaining 10% as the test set
[0021] Furthermore, in step S2, the steps of building the backbone network are as follows:
[0022] S2.1 introduces the multi-task attention network MTAN with VGG-16 as the backbone, and constructs the target detection task and depth estimation task;
[0023] S2.2 introduces the large kernel attention mechanism LKA to replace the 1×1 convolution layer, BN layer and ReLu activation function in the MTAN attention module introduced in step S2.1;
[0024] S2.3 extracts the outputs Conv4-3-1, Conv7-1, Conv4-3-2 and Conv7-2 of the shared network Conv4-3 and Conv7 corresponding to the replaced attention module in step S2.2, and the outputs Conv4-3-1, Conv7-1, Conv4-3-2 and Conv7-2 are used as inputs for subsequent tasks;
[0025] S2.4 Upsample Conv4-3-1, Conv4-3-2, Conv7-1 and Conv7-2 by 2 times and then concatenate them in the channel dimension to obtain the feature map ψ 1 , 2 .
[0026] Furthermore, in step S3, the steps for building the target detection subnetwork are as follows:
[0027] S3.1 will ψ 1 As the input of the parallel multi-scale receptive field fusion module, the multi-scale receptive field fusion module is connected in parallel with the ASPP module, and the hole rate of the ASPP module is set to 1, 6, and 12 respectively. Then, the feature map φ passing through the ASPP module is extracted 1 ,φ 2 and φ 3 ;
[0028] S3.2 Feature map φ 1 ,φ 2 ,φ 3 As a benchmark, 4 additional groups of convolutions are added to each feature map. The first group of convolutions is a 3×3 convolution with a step size of 1, and the following 3 groups of convolutions are composed of 3×3 convolutions with a step size of 2. The feature maps after adding convolutions are extracted to construct feature pyramids.
[0029] S3.3 selects pyramids of the same resolution size from the three pyramids of different receptive field scales for channel dimension splicing, then introduces the SE module for learning, and uses the final feature pyramid as the initial inspection network of the target detection network;
[0030] Based on the initial inspection network, S3.4 uses weighted deformable convolution to process feature maps of various scales, thereby improving the regression accuracy of the detection box.
[0031] Furthermore, in step S4, the steps for building the depth estimation subnetwork are as follows:
[0032] S4.1 will ψ 2 As input for the DORN depth estimation task;
[0033] S4.2 adds a scene understanding module, wherein the scene understanding module includes a full image encoding module, a cross-channel information compression module, and a dilated spatial convolution pooling pyramid module;
[0034] S4.3 uses an ordinal regression module to classify discrete depth values into multiple categories.
[0035] Further, in step S5, the training steps are as follows:
[0036] S5.1 Design of overall loss function L total , the overall loss function L total Including the target detection loss function L detect And the depth estimation loss function L depth ;
[0037] S5.2 sets the network input image size, initial learning rate and number of iterations;
[0038] S5.3 uses the loss function adaptive strategy to train the network model.
[0039] Further, in step S6, the optimization detection steps are as follows:
[0040] S6.1 Input the image to be predicted and obtain the coordinates of the vehicle detection frame and the depth value of each pixel in the image;
[0041] S6.2 calculates the coordinates of the center point of the detection frame according to the coordinates of the vehicle detection frame, and then uses the coordinates of the center point as the center point of the depth extraction area, and constructs a depth value extraction area with half the height and width of the detection frame;
[0042] S6.3 introduces the K-Means clustering algorithm to detect the distance of the forward vehicle target.
[0043] The beneficial effects of the present invention are:
[0044] 1. The present invention provides a forward vehicle recognition and distance detection method based on deep learning. Through deep learning technology, target detection and target distance measurement are driven. Depth information can be used to characterize the true distance value in the image, and monocular depth estimation can improve the accuracy of distance detection. The present invention improves the multi-scale target detection effect and positioning performance through a multi-scale receptive field fusion module and an improved cascade SSD vehicle detection algorithm.
[0045] 2. The present invention provides a forward vehicle recognition and distance detection method based on deep learning, introduces a multi-task attention network MTAN, connects the target detection task and the depth estimation task in the forward vehicle distance detection network in parallel, and proposes an end-to-end target detection and monocular depth estimation multi-task learning model, which solves the problem of the difficulty in balancing the correlation and difference between the target detection task and the depth estimation task. The present invention also introduces a large-core attention mechanism and a multi-task loss function adaptive strategy to further improve the accuracy of target detection and depth estimation. At the same time, the present invention also proposes a ranging method based on ranging feature point fitting and K-Means optimization, which solves the problem that the depth value of the non-vehicle area in the 2D vehicle boundary box interferes with the distance detection.
[0046] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Schematic diagram of the network structure of the forward vehicle recognition and distance detection method based on deep learning of the present invention. DETAILED DESCRIPTION
[0048] In order to make the technical solutions, advantages and purposes of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of this application.
[0049] like Figure 1 As shown, the present invention provides a forward vehicle recognition and distance detection method based on deep learning, comprising the following steps:
[0050] S1. Obtain the dataset of forward vehicle recognition and distance detection object detection and depth estimation alignment, and preprocess the dataset, which is mainly divided into the following five parts:
[0051] S1.1 Download the KITTI dataset, which contains vehicle training images, annotation files, and point cloud files.
[0052] The KITTI dataset was jointly created by the Karlsruhe Institute of Technology in Germany and Toyota Research Institute of America. It is currently the world's largest computer vision algorithm evaluation dataset for autonomous driving scenarios.
[0053] The data collection platform of the KITTI dataset is equipped with 2 grayscale cameras, 2 color cameras, a Velodyne 64-line 3D lidar, 4 optical lenses, and a GPS navigation system.
[0054] Point cloud is a collection of massive points that express the spatial distribution and surface characteristics of the target in the same spatial reference system. After obtaining the spatial coordinates of each sampling point on the surface of the object, a collection of points is obtained, which is called a "point cloud".
[0055] S1.2 uses the KITTI target detection original RGB image and the corresponding camera parameter matrix and radar point cloud data to generate a depth map label in uint8 data format, in meters.
[0056] S1.3 uses the 2D target detection annotation file of the KITTI dataset as the basis, cleans the data, filters out images containing vehicle information annotations, and performs vehicle target detection branch training.
[0057] By counting the distance of each vehicle target, the farthest target in the KITTI data set of this embodiment is 86.18m. In order to keep consistent with the depth value estimation range of 0-80m set by the DORN model, the image needs to be screened twice.
[0058] S1.4 Refer to the K-Means (k-means clustering algorithm) method in the YOLOv2 algorithm to cluster the target boxes into 2-11 categories, and obtain 10 average IoUs. Then use the number of cluster categories as the horizontal axis and the average IoU as the vertical axis to obtain the category number-average IoU graph. Then find the inflection point that best balances speed and accuracy and determine the number of categories.
[0059] IoU is the intersection over union ratio, which calculates the ratio of the intersection and union of the "predicted bounding box" and the "real bounding box".
[0060] S1.5 randomly divides the original RGB images into a training set and a test set in a ratio of 9:1. In this embodiment, there are about 6,800 images after data cleaning.
[0061] S2. Building a backbone network mainly includes the following four parts:
[0062] S2.1 introduces the multi-task attention network MTAN with VGG-16 as the backbone, and constructs two branches: target detection task and depth estimation task.
[0063] VGG is the abbreviation of Visual Geometry Group Network, a deep convolutional neural network. The number "16" in VGG16 means that there are 13 convolutional layers and 3 fully connected layers in the VGG structure.
[0064] MTAN (Multi-Task Attention Netwrok) consists of a shared network with global feature pooling and task-specific soft-attention modules. These modules learn task-specific features from global shared features while allowing features to be shared between different tasks. The architecture can be trained end-to-end and can be built on any forward neural network. It is simple to implement and has high parameter efficiency.
[0065] S2.2 introduces the large kernel attention mechanism LKA to replace the 1×1 convolution layer, BN layer and ReLu activation function in the MTAN attention module introduced in step S2.1.
[0066] LKA consists of three parts: spatial local convolution (Depth-wise Convolution), spatial long-range convolution (Depth-wise Dilation Convolution) and channel convolution (1×1 Convolution). The spatial local convolution is a 5×5 depth convolution, and the spatial long-range convolution is a 7×7 depth convolution with a dilation rate of 3.
[0067] S2.3 extracts the outputs Conv4-3-1, Conv7-1, Conv4-3-2 and Conv7-2 of the shared network Conv4-3 and Conv7 corresponding to the replaced attention module in step S2.2, and uses the outputs Conv4-3-1, Conv7-1, Conv4-3-2 and Conv7-2 as the input for subsequent tasks.
[0068] Among them, the output Conv4-3-1 (represents the output of the soft attention mask with the shared network Conv4-3 in the target detection subnetwork) and Conv4-3-2 (represents the output of the soft attention mask with the shared network Conv4-3 in the depth estimation subnetwork).
[0069] S2.4 Upsample Conv4-3-1, Conv4-3-2, Conv7-1 and Conv7-2 by 2 times and then concatenate them in the channel dimension to obtain the feature map ψ 1 , 2 .
[0070] Upsampling, alias: enlarging the image, also called image interpolation.
[0071] S3. Build the target detection sub-network, which is mainly divided into the following four parts:
[0072] S3.1 The obtained feature map ψ 1 As the input of the Multi-scale Receptive Field Fusion Module (MRFFM), the Multi-scale Receptive Field Fusion Module is connected in parallel with the Res-ASPP module. The hole rates of the Res-ASPP module are set to 1, 6, and 12 respectively. Then, the feature maps of the three scale receptive fields are obtained through the Res-ASPP module: φ 1 ,φ 2 and φ 3 .
[0073] S3.2 Feature map φ 1 ,φ 2 ,φ 3 As a benchmark, 4 additional groups of convolutions are added to each feature map. The first group of convolutions is a 3×3 convolution with a stride of 1, and the following 3 groups are composed of 3×3 convolutions with a stride of 2. The feature maps of these 4 scales are extracted to construct a feature pyramid. The resolution of each layer of the pyramid is shown in Table 1.
[0074] Table 1 Feature pyramid resolution
[0075]
[0076] S3.3 selects pyramids of the same resolution size from three pyramids of different receptive field scales for channel dimension splicing, and then introduces the SE module to learn the importance of each channel in different feature layers. The final feature pyramid is used as the preliminary inspection network of the target detection network, and the preliminary inspection network is used to classify the background and foreground.
[0077] S3.4, based on the initial inspection network, uses weighted deformable convolution for the feature maps of each scale, and uses a learning-based method to obtain the offset of the convolution sampling point to complete the weighted deformable convolution calculation.
[0078] Step S3.4 can be used to alleviate the feature misalignment problem caused by the anchor frame mechanism and improve the regression accuracy of the detection frame. The initial inspection network outputs four variables: (dx, dy, dh, dw), where (dx, dy) corresponds to the offset of the spatial position, and (dh, dw) represents the offset on the scale. The present invention designs a weighted feature alignment module WFAM. The offset required by WFAM is obtained by multiplying dx, dy by an initialization weight weight and then by a convolution. The whole process of WFAM is as follows:
[0079] Δp=f(weight·(dx,dy))
[0080] a′=WeightDeformableConv(a,Δp)
[0081] In the formula, f represents convolution, a and a′ represent the original feature map and the aligned feature map respectively, and Δp is the deviation of dx, dy, and weight learned by convolution. For the initialization weight weight, its initial value is set to 0.5. The initial value of the deviation of dx and dy is set to 0. WeightDeformable represents the deformable convolution with weight.
[0082] S4. Build the depth estimation subnetwork, which is mainly divided into the following four parts:
[0083] S4.1 The obtained feature map ψ 2 As the input of the DORN model depth estimation task;
[0084] S4.2 adds a scene understanding module consisting of a full-image encoding module, a cross-channel information compression module, and an atrous spatial convolutional pyramid pooling module (ASPP), thereby enabling the network to fully understand the input image.
[0085] The full-image encoding module can capture global context information and thus reduce the local aliasing problem.
[0086] The cross-channel information compression module aims to enhance the nonlinearity of features through 1×1 convolution, compress the feature dimensions, and interactively share information between channels.
[0087] The ASPP module expands the receptive field of the feature map by setting the expansion rates to 6, 12, and 18 while ensuring that the resolution of the feature map remains unchanged, thereby capturing contexts of different proportions for the same input, which helps to extract scene features.
[0088] S4.3 introduces a spatial increasing discretization strategy (SID) into the DORN model, continuously samples depth values, and then uses the ordinal regression module to divide the discrete depth values into multiple categories. The depth estimation problem is regarded as an ordered regression problem, and an ordered loss is used to learn its network parameters.
[0089] S5. The training of the forward vehicle recognition and distance detection network based on deep learning mainly includes the following three parts:
[0090] S5.1 Design of overall loss function L total , the overall loss function L total Including target detection loss function L detect And the monocular estimation network loss function L depth,λ 1 and λ 2 is a hyperparameter, and the overall loss function L total The calculation expression is as follows:
[0091] L total =L detect +L depth
[0092] Among them, the loss function of cascade SSD target detection is divided into two parts: Anchor Refinement Module (ARM) loss function and Object Detection Module (ODM) loss function. The classification loss function adopts Softmax cross entropy loss function, and the regression loss function adopts smooth-L1 loss function.
[0093] The calculation expression of the softmax cross entropy loss function is as follows:
[0094]
[0095] Where, L cls (x,c) is the softmax cross entropy loss function, i represents the candidate box number, j represents the true label box number, p is the category number, p = 0 represents the background, where Taking 1 means that the i-th candidate box matches the j-th annotation box, and the category of this annotation box is p. It represents the probability value of the predicted category p for the i-th candidate box. The first half of the formula is the loss of the positive sample (Pos), that is, the loss of being classified as a certain category (excluding the background), and the second half is the loss of the negative sample (Neg), that is, the loss of the background category.
[0096] Object detection loss function L detect The calculation expression is as follows:
[0097]
[0098]
[0099] Loss detect =Loss ARM +Loss ODM
[0100] In the formula, i represents the number of Anchor, p i With x i They represent the probability that the model predicts that the i-th Anchor is the foreground target and the regression value of the Anchor in the ARM stage, respectively. i With t iRespectively represent the predicted target category and the coordinates of the prediction box in the ODM stage. Loss ARM With Loss ODM The tables are the loss functions in the ARM and ODM stages, L b-cls and L cls Represents the binary classification and full category Softmax cross entropy, L reg represents the CIoU loss function based on Centerness weighting, is the category label of the i-th Anchor, is a sign function, If it is greater than or equal to 1, it is 1, otherwise it is 0. is the position and size of the target matched by the i-th Anchor, and λ is set to 1 with reference to RefineDet and Cas-SSD.
[0101] Depth estimation uses the DORN depth estimation loss function and the monocular estimation network loss function L depth The calculation expression is as follows:
[0102]
[0103]
[0104]
[0105] l (w,h) ∈{0, 1, ..., K-1}
[0106] Where χ is the feature map obtained, W is the width of the feature map, H is the height of the feature map, Θ is the weight vector, N is the number of pixels in the feature map, and l is (w,h) Indicates the discrete value label of the pixel (w, h) corresponding to the SID strategy, l (w,h) is the estimated discrete value decoded in the output of the ordered regression, and P is the predicted probability value.
[0107] S5.2 sets the network input image size, initial learning rate and number of iterations. In this embodiment, the network input image size is set to 320*320, the initial learning rate is set to 0.0004, and the number of iterations is set to 120 epochs (the forward vehicle distance detection based on the improved MTAN is set to 250 epochs).
[0108] The first three epochs use the warm-up strategy, with the learning rate increasing from 10 -6At the beginning, it linearly increases to 0.0004 after 3 epochs, and the learning rate of subsequent epochs is decayed using the cosine annealing algorithm. The optimizer uses SGD (stochastic gradient descent) with momentum, where the momentum is set to 0.9 and the weight decay is set to 0.0005 to prevent the model from overfitting. The warmup strategy and cosine annealing algorithm are defined as follows:
[0109]
[0110]
[0111] Among them, lr min According to experience, it is set to 10 -6 ; lr base Indicates the initial learning rate is 0.0004; epoch_size is set to 3, indicating that the first three epochs use the warm-up WarmUp strategy; iter and Iter respectively indicate the number of iterations required for an epoch and the current number of iterations; T cur and T sum The sub-table shows the current number of iterations and the total number of iterations.
[0112] S5.3 uses the loss function adaptive strategy to train the network model. The calculation expression is as follows:
[0113] L total =λ 1 L detect +λ 2 L depth
[0114] In the formula, λ 1 With λ 2 They are used as weight parameters for target detection and monocular depth estimation tasks respectively.
[0115] According to the multi-task loss function adaptive weight strategy, λ 1 With λ 2 The calculation expression is as follows:
[0116]
[0117]
[0118] In the formula, λ k represents the learning weight corresponding to task k, k∈{1,2}; w k (·) represents the relative rate of decrease calculated in the interval (0, +∞); t represents the current number of iterations; T represents the variable that controls the softness of task weights. A larger T will make the weight distribution between different tasks more uniform. If T is large enough, then λ k≈1, the learning weights of all tasks tend to be equal, so this embodiment sets T to 2 so that the network can find a balance between the two tasks; finally, the softmax operation is multiplied by K (the total number of tasks) to ensure that ∑ i λ i (t) = K. In this embodiment, L k (t) is the average of all the iteration loss in each epoch. This can reduce the uncertainty caused by stochastic gradient descent and random training data selection. For t = 1, 2, w k (t) is initialized to 1.
[0119] S6. Forward vehicle distance detection based on K-Means optimization mainly includes the following three parts:
[0120] S6.1 Input the image to be predicted and obtain the coordinates of the vehicle detection frame and the depth value of each pixel in the image;
[0121] S6.2 calculates the center point coordinates of the vehicle detection frame according to the detection frame coordinates (x1, y1, x2, y2). The calculation expression of the center point coordinates is:
[0122]
[0123] This center point is used as the center point of the depth value extraction area, and the depth value extraction area is constructed with half the height and width of the detection box.
[0124] S6.3 introduces the K-Means clustering algorithm to detect the distance of the forward vehicle target. The detection process is as follows:
[0125] ① Consider all bounding boxes output by the target detection model of the forward vehicle distance detection model as a set The output of monocular depth estimation is a depth map M;
[0126] ②Calculation set The IoU between any two boxes in the image is used to determine whether there is occlusion between cars. If the IoU between the two boxes is greater than the set threshold (such as 0.3), it means that the occlusion between the cars is high, and K-Means clustering analysis is required for the pixels in the depth map area corresponding to the two boxes; otherwise, jump to step ③;
[0127] 1) Filter out the top two clusters in each bounding box, and record their cluster "center point" and cluster number as c 1 、c 2 n 1 、n 2 ,
[0128] 2) When n 1 ≥1.5n2 When c 1 ≤80, then select c 1 As the forward distance detection value of the vehicle corresponding to the bounding box; otherwise, c 2 As the forward distance detection value of the vehicle corresponding to the bounding box;
[0129] 3) When n 1 <1.5n 2 When c 1 and c 2 The minimum value of is used as the forward distance detection value of the vehicle corresponding to the bounding box to alleviate the interference of the background depth value;
[0130] ③ In order to improve the speed of forward distance detection, the method of calculating the mean depth value in the depth value extraction area is used to collect The distance detection is performed on all vehicle bounding boxes in the calculation expression:
[0131]
[0132] Where d(w,h) represents the depth value corresponding to the pixel point (w,h) in the depth map; N represents the total number of pixels in the depth value information extraction area; W and H represent the height and width of the depth value information extraction area respectively; Distance is the forward distance of the center point of the fitted 3D vehicle.
[0133] In summary, the present invention realizes forward vehicle recognition and distance detection from five aspects: data set, network structure design, model building, loss function design and target ranging feature point fitting.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the protection scope of the present invention.
Claims
1. A forward vehicle recognition and distance detection method based on deep learning, It is characterized in that The following steps are involved: S1. Obtain the data set required for forward vehicle identification and distance detection, and preprocess the data set; S2. Build a backbone network; S2.1 introduces the multi-task attention network MTAN with VGG-16 as the backbone, and constructs the target detection task and depth estimation task; S2.2 introduces the large kernel attention mechanism LKA to replace the 1×1 convolution layer, BN layer and ReLu activation function in the MTAN attention module introduced in step S2.1; S2.3 extracts the outputs Conv4-3-1, Conv7-1, Conv4-3-2 and Conv7-2 of the shared network Conv4-3 and Conv7 corresponding to the replaced attention module in step S2.2, and the outputs Conv4-3-1, Conv7-1, Conv4-3-2 and Conv7-2 are used as inputs for subsequent tasks; S2.4 Upsample Conv4-3-1, Conv4-3-2, Conv7-1 and Conv7-2 by 2 times and then concatenate them in the channel dimension to obtain the feature map ψ 1 , 2 ; S3. Build the target detection sub-network; S3.1 will ψ 1 As the input of the parallel multi-scale receptive field fusion module, the multi-scale receptive field fusion module is connected in parallel with the ASPP module, and the hole rate of the ASPP module is set to 1, 6, and 12 respectively. Then, the feature map φ passing through the ASPP module is extracted 1 ,φ 2 and φ 3 ; S3.2 Feature map φ 1 ,φ 2 ,φ 3 As a benchmark, 4 additional groups of convolutions are added to each feature map. The first group of convolutions is a 3×3 convolution with a step size of 1, and the following 3 groups of convolutions are composed of 3×3 convolutions with a step size of 2. The feature maps after adding convolutions are extracted to construct feature pyramids. S3.3 selects pyramids of the same resolution size from the three pyramids of different receptive field scales for channel dimension splicing, then introduces the SE module for learning, and uses the final feature pyramid as the initial inspection network of the target detection network; S3.4 uses weighted deformable convolution to process feature maps of various scales based on the initial inspection network, thereby improving the regression accuracy of the detection frame; S4. Build a depth estimation subnetwork; S4.1 will ψ 2 As input for the DORN depth estimation task; S4.2 adds a scene understanding module, wherein the scene understanding module includes a full image encoding module, a cross-channel information compression module, and a dilated spatial convolution pooling pyramid module; S4.3 uses an ordinal regression module to classify discrete depth values into multiple categories; S5. Training of forward vehicle recognition and distance detection network based on deep learning; S6. Optimize forward vehicle distance detection based on K-Means.
2. According to the method for forward vehicle recognition and distance detection based on deep learning in claim 1, It is characterized in that The data set of step S1 is a KITTI data set, and the KITTI data set includes vehicle training pictures, annotation files, and point cloud files.
3. According to the method for forward vehicle recognition and distance detection based on deep learning in claim 2, It is characterized in that In step S1, the preprocessing steps of the data set are as follows: S1.1 Convert point cloud files into depth map labels; S1.2 cleans the KITTI dataset processed in step S1.1, filters out the images with incorrect annotations and removes them; S1.3 uses K-means k-means clustering algorithm to determine the number of anchor boxes and aspect ratio; S1.4 uses 90% of the data set as the training set and the remaining 10% as the test set.
4. According to the method for forward vehicle recognition and distance detection based on deep learning in claim 1, It is characterized in that In step S5, the training steps are as follows: S5.1 Design of overall loss function L total , the overall loss function L total Including the target detection loss function L detect And the depth estimation loss function L depth ; S5.2 sets the network input image size, initial learning rate and number of iterations; S5.3 uses the loss function adaptive strategy to train the network model.
5. According to the method for forward vehicle recognition and distance detection based on deep learning in claim 1, It is characterized in that In step S6, the optimization detection steps are as follows: S6.1 Input the image to be predicted and obtain the coordinates of the vehicle detection frame and the depth value of each pixel in the image; S6.2 calculates the coordinates of the center point of the detection frame according to the coordinates of the vehicle detection frame, and then uses the coordinates of the center point as the center point of the depth extraction area, and constructs a depth value extraction area with half the height and width of the detection frame; S6.3 introduces the K-Means clustering algorithm to detect the distance of the forward vehicle target.
Citation Information
Cited By
Target distance and speed measuring method based on monocular vision
CN120926942A