A target detection method based on the combination of Gm-APD lidar and deep learning
By combining Gm-APD lidar and deep learning, using adaptive histogram equalization and data set fusion technology, an improved YOLOv5s network is established, which solves the accuracy and speed problems in long-distance object detection and achieves efficient long-distance vehicle detection.
Patent Information
- Application Number
- CN202210627946.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-06
AI Technical Summary
The prior art has problems with accuracy and speed in long-distance target detection, especially in vehicle detection in military operations.
The object detection method based on Gm-APD lidar and deep learning is adopted to preprocess the original image data through the adaptive histogram equalization method, enhance the intensity image, and fuse the distance image with the enhanced intensity image in the data set. Then, an improved YOLOv5s network is established, an attention module SE and Bottleneck structure are introduced, and the network structure is optimized to improve detection accuracy.
It improves the accuracy and speed of long-distance target detection, can realize real-time detection, and has practical value in military combat.
Smart Images

Figure CN115019094B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Gm-APD target detection, and particularly to a target detection method based on the combination of Gm-APD lidar and deep learning. Background Art
[0002] In recent years, lidar technology has developed rapidly. As a three-dimensional imaging detection means, it is currently widely used in military and civilian fields such as precision guidance, military strikes, resource exploration, urban planning, water conservancy projects, and environmental monitoring. In the early stage, laser active three-dimensional imaging technology mainly adopted the scanning method. To obtain the three-dimensional information of the target by scanning the target area point by point, a scanning device needs to be added to the laser emission and reception system. Therefore, the system has a large volume, increased power consumption and weight, which is not conducive to the miniaturization and integration of the overall system. At the same time, it limits the imaging speed of the system. In recent years, these unfavorable factors have prompted countries to vigorously carry out research on non-scanning laser active three-dimensional imaging technology. Among them, the Geiger mode Avalanche Photo Diode (Gm-APD) is a single-photon detection device that adopts non-scanning active three-dimensional imaging technology and has advantages such as good concealment, strong anti-interference ability, small volume, light weight, high sensitivity, high precision, and high integration. It can meet the requirements of system miniaturization, integration, and high-speed imaging, and is one of the important development directions of target detection technology in military operations.
[0003] Currently, with the large-scale application of deep learning in the field of target detection, the accuracy and speed of target detection technology have been rapidly improved and have been widely used in fields such as pedestrian detection, face detection, text detection, traffic sign and signal detection, and remote sensing image detection. Therefore, aiming at the target detection problem in military operations, taking vehicle detection as an example, a target detection method based on the combination of Gm-APD lidar and deep learning is designed to improve the accuracy of long-distance target detection. Summary of the Invention
[0004] The present invention provides a target detection method based on the combination of Gm-APD lidar and deep learning to solve the problem of difficult long-distance target detection in the prior art and improve the detection accuracy.
[0005] The present invention is achieved through the following technical solutions:
[0006] A target detection method based on the combination of Gm-APD lidar and deep learning, the target detection method comprising the following steps:
[0007] Step 1: Preprocess the original image data collected by the Gm-APD lidar using the adaptive histogram equalization method;
[0008] Step 2: Enhance the intensity image in the original image data preprocessed in Step 1, and perform dataset fusion on the distance image and the enhanced intensity image;
[0009] Step 3: Establish an improved YOLOv5s network;
[0010] Step 4: Use the data fused in Step 2 to train and validate the improved YOLOv5s network established in Step 3;
[0011] Step 5: Perform target detection using the improved YOLOv5s network.
[0012] A target detection method based on the combination of Gm-APD lidar and deep learning. In Step 1, the adaptive histogram equalization method is used to preprocess the original image data collected by the Gm-APD lidar. Specifically,
[0013] Step 1.1: Divide the original image data, i.e., the picture, collected by the Gm-APD lidar into several sub-block regions;
[0014] Step 1.2: Add contrast limits to each small block region, and clip the histogram obtained by statistics in the sub-block so that its amplitude is lower than a certain upper limit.
[0015] A target detection method based on the combination of Gm-APD lidar and deep learning. The specific steps for enhancing the intensity image in the original image data preprocessed in Step 2 are as follows,
[0016] Step RH2.1: Perform grayscale transformation on the single-channel two-dimensional matrices of the 64x64 intensity image and the distance image;
[0017] Step RH2.2: Perform normalization processing on the grayscale transformation of the single-channel two-dimensional matrix in Step RH2.1;
[0018] Step RH2.3: For the data after normalization processing in Step RH2.2, combine the two-channel 64x64 intensity image with the one-channel 64x64 distance to obtain an IIR three-channel image;
[0019] Step RH2.4: Use 8-fold nearest neighbor interpolation to obtain a 512x512 IIR image of a long-distance target.
[0020] A target detection method based on the combination of Gm-APD lidar and deep learning. The specific normalization processing in Step RH2.2 is as follows: The normalization calculation of the two-dimensional matrix is:
[0021]
[0022] Among them, I(i,j) is the pixel value at the position (i,j), y max = 255, y min = 0, x max is the maximum value of I(i,j), x min is the minimum value of I(i,j).
[0023] A target detection method based on the combination of Gm-APD lidar and deep learning. Specifically, the improved YOLOv5s network established in step 3 is to modify the BottleneckCSP structure into a Bottleneck structure, and add an attention module SE one step before the data output. An attention mechanism is introduced into the network structure to establish dynamic weight parameters, so as to strengthen key information and weaken useless information, thereby improving the efficiency of the deep learning algorithm.
[0024] A target detection method based on the combination of Gm-APD lidar and deep learning. The improved YOLOv5s network includes a Conv structure I, a Conv structure II, a Bottleneck structure, a Conv structure III, a Bottleneck structure, a Conv structure IV, a Bottleneck structure, a Conv structure V, and an attention module SE. The data passes through the Conv structure I, the Conv structure II, the Bottleneck structure, the Conv structure III, the Bottleneck structure, the Conv structure IV, the Bottleneck structure, the Conv structure V, and the attention module SE in sequence from left to right and then outputs.
[0025] A target detection method based on the combination of Gm-APD lidar and deep learning. For the attention module SE, the input is x, and the number of its feature channels is c1. Through a series of convolutions, a feature with the number of feature channels c2 is obtained. The importance of each feature channel is automatically obtained through learning, and then the useful features are enhanced and the features that are not very useful for the current task are suppressed according to this importance.
[0026] A target detection method based on the combination of Gm-APD lidar and deep learning. The loss function of the improved YOLOv5s network specifically includes classification loss, localization loss, and object confidence loss. The binary cross-entropy loss function is used to calculate the loss of the class probability and the object confidence score, and GIOU is used as the loss function for bounding box regression. The calculation formula is as follows:
[0027]
[0028] Among them, A is the predicted box, B is the ground truth box, C is the smallest enclosing box of A and B, GIOU lossThe value range of is [0, 2], taking the minimum value of 0 when the two boxes completely overlap, the loss function value is 1 when the sides of the two boxes are externally tangent, and the loss function value is 2 when the two boxes are separated and far apart.
[0029] A target detection method based on the combination of Gm-APD lidar and deep learning. The specific detection result of the improved YOLOv5s network established in step 3 using the dataset fused in step 2 in step 4 is as follows:
[0030] In target detection, the accuracy P, recall rate R, and F1 score are mainly used to evaluate the detection performance of the model. The formula is as follows:
[0031]
[0032]
[0033]
[0034] Among them, P is the accuracy, R is the recall rate, TP is the number of true positive samples, FP is the number of false positive samples, and FN is the number of false negative samples.
[0035] A computer-readable storage medium stores a computer program therein. When the computer program is executed by a processor, the above-mentioned method steps are implemented.
[0036] The beneficial effects of the present invention are:
[0037] In the present invention, aiming at the similarity between the target gray level and the background gray level at a long distance in the intensity image, the method of limited contrast adaptive histogram equalization is adopted to locally enhance the weak target, and the intensity image is combined with the distance image.
[0038] The present invention improves the YOLOv5 network, redesigns a lightweight backbone network, can be deployed in mobile devices, and introduces an SE structure to improve the performance of the backbone network. Compared with the original network, the detection accuracy of vehicles at a long distance is improved, and the target can be detected in real time, which has practical value in military operations. Description of the Drawings
[0039] Figure 1 is a schematic structural diagram of the present invention.
[0040] Figure 2 is a schematic structural diagram of the improved YOLOv5s network of the present invention.
[0041] Figure 3 is a schematic structural diagram of the SE structure of the present invention.
[0042] Figure 4 is a schematic diagram of the GIOU of the present invention.
[0043] Figure 5 is the result curve graph during the training process of the present invention, where Figure 5 -(a) Precision curve graph, Figure 5 -(b) Recall curve graph, Figure 5 -(c) Localization loss curve graph, Figure 5 -(d) Class loss curve graph.
[0044] Figure 6 are the detection results of the present invention, where Figure 6 -(a) Target scene graph, Figure 6 -(b) Intensity image, Figure 6 -(c) Processed IIR image, Figure 6 -(d) Target detection result. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] The present invention utilizes a Gm-APD lidar system device with 64×64 pixels, and combines the collected distance image and intensity image with deep learning methods for target detection.
[0047] A target detection method based on the combination of Gm-APD lidar and deep learning, the target detection method comprising the following steps:
[0048] Step 1: Preprocess the original image data collected by the Gm-APD lidar using the adaptive histogram equalization method;
[0049] Step 2: Enhance the intensity image in the original image data preprocessed in Step 1, and perform dataset fusion on the distance image and the enhanced intensity image;
[0050] Step 3: Establish an improved YOLOv5s network;
[0051] Step 4: Train and validate the improved YOLOv5s network established in Step 3 with the data fused in Step 2;
[0052] Step 5: Detect the target using the improved YOLOv5s network.
[0053] A target detection method based on the combination of Gm-APD lidar and deep learning. In step 1, the adaptive histogram equalization method is used to preprocess the original image data collected by the Gm-APD lidar. Specifically, in the lidar intensity image, due to the similar contrast between weak targets at long distances and the background, local contrast enhancement can improve the contrast of weak targets, improve the accuracy of target detection, increase the number of misdetected target samples, increase the number of sample features learned by the network, solve the problem of sample imbalance, and thus improve the accuracy of target detection. Local contrast enhancement algorithm;
[0054] Step 1.1: Divide the original image data, i.e., the picture, collected by the Gm-APD lidar into several sub-block regions;
[0055] Step 1.2: Add a contrast limit to each small block region, and clip the histogram obtained by statistics in the sub-block so that its amplitude is lower than a certain upper limit. The clipped part is evenly distributed over the entire gray level interval to ensure that the total area of the histogram remains unchanged.
[0056] A target detection method based on the combination of Gm-APD lidar and deep learning. The enhancement of the intensity image in the original image data preprocessed in step 1 in step 2 specifically includes the following steps,
[0057] Since the Gm-APD lidar can generate a distance image and an intensity image. The intensity image is the relative intensity information of the target obtained through the amplitude of the echo data, and the distance image is the three-dimensional distance information obtained by using the time-of-flight method based on the time difference of the received laser. According to the analysis of the characteristics of the lidar intensity image and the distance image, the intensity image is affected by complex environments such as sunlight and the backscattering of particles in the air. Most of the geometric and structural characteristics represented by the distance image do not change with vision. As the target distance increases, the distance resolution of the distance image will decay, and the distance boundary only describes the structural attributes of the target, lacking the appearance discrimination of the target and being insufficient to detect the target. Therefore, the fusion of the distance image and the intensity image is very necessary.
[0058] Since the main feature of the intensity image is the gradient change of gray level, and the main feature of the distance image is the contour information of the target.
[0059] Step RH2.1: Perform gray level transformation on the single-channel two-dimensional matrix of the 64x64 intensity image and the distance image;
[0060] Step RH2.2: Perform normalization processing on the gray level transformation of the single-channel two-dimensional matrix in step RH2.1;
[0061] Step RH2.3: Normalize the data processed in Step RH2.2 to the range of 0 - 255, and obtain an IIR three-channel image by combining the 64x64 intensity images of two channels with the 64x64 distance of one channel; the IIR three-channel image contains not only intensity information but also distance information;
[0062] Step RH2.4: Use 8-fold nearest neighbor interpolation to obtain a 512x512 IIR image of a distant target.
[0063] A target detection method based on the combination of Gm-APD lidar and deep learning. The normalization process in Step RH2.2 is specifically as follows: the normalization calculation of a two-dimensional matrix is:
[0064]
[0065] where I(i, j) is the pixel value at the position (i, j), y max = 255, y min = 0, x max is the maximum value of I(i, j), and x min is the minimum value of I(i, j).
[0066] A target detection method based on the combination of Gm-APD lidar and deep learning. The specific process of establishing the improved YOLOv5s network in Step 3 is as follows: modify the BottleneckCSP structure to the Bottleneck structure, and add an attention module SE one step before the data output. Introduce an attention mechanism into the network structure to establish dynamic weight parameters, so as to strengthen key information and weaken useless information, thereby improving the efficiency of the deep learning algorithm. Modifying the BottleneckCSP structure to the Bottleneck structure can greatly simplify the complexity of the network, reduce the parameters of the network, and reduce the number of layers of the network structure. While simplifying the network, it is also necessary to ensure the detection accuracy.
[0067] A target detection method based on the combination of Gm-APD lidar and deep learning. The improved YOLOv5s network includes a Conv structure Ⅰ, a Conv structure Ⅱ, a Bottleneck structure, a Conv structure Ⅲ, a Bottleneck structure, a Conv structure Ⅳ, a Bottleneck structure, a Conv structure Ⅴ, and an attention module SE. The data passes through the Conv structure Ⅰ, the Conv structure Ⅱ, the Bottleneck structure, the Conv structure Ⅲ, the Bottleneck structure, the Conv structure Ⅳ, the Bottleneck structure, the Conv structure Ⅴ, and the attention module SE in sequence from left to right and then is output.
[0068] A target detection method based on the combination of Gm-APD lidar and deep learning. The attention module SE has a network structure as follows: Figure 3 As shown, the input is x with the number of feature channels c1. Through a series of convolutions, a feature with the number of feature channels c2 is obtained. The importance degree of each feature channel is automatically obtained through learning, and then the useful features are enhanced and the features that are not very useful for the current task are suppressed according to this importance degree.
[0069] A target detection method based on the combination of Gm-APD lidar and deep learning. The loss function of the improved YOLOv5s network specifically includes classification loss, localization loss, and object confidence loss. The binary cross-entropy loss function is used to calculate the loss of class probability and object confidence score, and GIOU is used as the loss function for bounding box regression. The calculation formula is:
[0070]
[0071] where A is the predicted box, B is the ground truth box, C is the smallest enclosing box of A and B. The relationship between A, B, and C is specifically as follows: Figure 4 As shown; GIOU loss ranges from [0, 2]. When the two boxes completely overlap, the minimum value 0 is taken. When the sides of the two boxes are externally tangent, the loss function value is 1. When the two boxes are separated and far apart, the loss function value is 2. Using the method of the circumscribed rectangle can not only reflect the area of the overlapping region but also calculate the proportion of the non-overlapping region. Therefore, the GIOU loss function can effectively reflect the overlapping degree and the distance between the ground truth box and the predicted box.
[0072] A target detection method based on the combination of Gm-APD lidar and deep learning. The detection result of using the dataset fused in step 2 to detect the improved YOLOv5s network established in step 3 in step 4 is as follows: The data used in the present invention is the dynamic vehicle data on the road about 1 km away at a fixed position of the Gm-APD lidar at night. The field of view angle of the lidar is 1.8°, and the repetition frequency is 20 kHz. The detailed information of the data is shown in Table 1. The specific acquisition time is from 20:37:19 to 21:23:50 at night, 12 groups of data, a total of 6000 images.
[0073] Table 1 Detailed information of the data
[0074]
[0075]
[0076] 1260 IIR images of 512x512 are used as the training set, 540 images as the validation set, and 4200 images for testing the network. The parameters for network training are set as epochs = 100, batchsize = 32, and the initial learning rate = 0.01. The curve graph of the object detection results during the training process is as shown in Figure 5 shown. In Figure (a), the detection accuracy increases with the increase of the training times and gradually stabilizes. The highest detection accuracy is 98.35%. In Figure (b), the highest recall rate also increases with the increase of the training times and finally stabilizes. The highest recall rate can reach 98.59%. Figures (c) and (d) are the object bounding box loss and object class loss respectively. As the network training times increase, the losses decrease. The localization loss decreases to about 0.03 and stabilizes, and the class loss decreases to about 0.04 and stabilizes. Therefore, when the model is trained for 100 epochs, both the loss function and the detection accuracy can tend to be stable, achieving a relatively high detection accuracy.
[0077] In object detection, the accuracy P (precision), recall R (recall), and F1 score are mainly used to evaluate the detection performance of the model. The formulas are as follows:
[0078]
[0079]
[0080]
[0081] Among them, P is the accuracy, R is the recall, TP is the number of true positive samples, FP is the number of false positive samples, and FN is the number of false negative samples. By validating the trained network model on the test data set, the detection accuracy is 97.36%, the recall rate is 98.35%, and the F1 score is 0.9785. The comparison of the detection results for the improved part with the results of the original network is shown in Table 2. Through image interpolation and comparison for enhancement processing, the learnable features of the network can be better enhanced and the features of the object and the background can be better distinguished. As can be seen from Table 2, compared with the original 64×64 images, the detection accuracy has increased by 3.14%, the recall rate has increased by 1.03%, and the F1 has increased by 0.0208. By improving the backbone network, not only can the network complexity be reduced, but also the training parameters are reduced, the network convergence speed is accelerated, and the detection accuracy of the network is improved. By adding the SE attention mechanism module, the network can pay more attention to the target area, effectively learn the target features, reduce the interference of background features, and the detection accuracy has increased by 1.3% compared with the original network, the recall rate has increased by 3.02%, and the F1 has increased by 0.0217.
[0082] Table 2 Comparison of different improved modules with the original network
[0083] Method P R F1 YOLOv5s_64x64 92.65% 93.62% 0.9313 YOLOv5s_BC_IIR_512x512 95.79% 94.65% 0.9521 YOLOv5s_backbone_BC_IIR_512x512 95.95% 95.40% 0.9568 YOLOv5s_SE_BC_IIR_512x512 97.09% 97.67% 0.9738
[0084] A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the above-described method steps are implemented.
Claims
1. A target detection method based on the combination of Gm-APD lidar and deep learning, characterized in that, the target detection method includes the following steps: Step 1: Preprocess the original image data collected by the Gm-APD lidar using the adaptive histogram equalization method; Step 2: Enhance the intensity image in the original image data preprocessed in Step 1, and fuse the distance image and the enhanced intensity image into a dataset; The specific steps for enhancing the intensity image in the original image data preprocessed in Step 1 in Step 2 are as follows, Step RH2.1: Perform gray-scale transformation on the single-channel two-dimensional matrix of the 64x64 intensity image and the distance image; Step RH2.2: Normalize the gray-scale transformation of the single-channel two-dimensional matrix in Step RH2.1; Step RH2.3: For the data after normalization in Step RH2.2, combine the two-channel 64x64 intensity image with the one-channel 64x64 distance to obtain an IIR three-channel image; Step RH2.4: Use 8-fold nearest neighbor interpolation to obtain a 512x512 IIR image of a long-distance target; Step 3: Establish an improved YOLOv5s network; The improved YOLOv5s network includes Conv Structure Ⅰ, Conv Structure Ⅱ, Bottleneck Structure, Conv Structure Ⅲ, Bottleneck Structure, Conv Structure Ⅳ, Bottleneck Structure, Conv Structure Ⅴ and the attention module SE. The data passes through Conv Structure Ⅰ, Conv Structure Ⅱ, Bottleneck Structure, Conv Structure Ⅲ, Bottleneck Structure, Conv Structure Ⅳ, Bottleneck Structure, Conv Structure Ⅴ and the attention module SE from left to right and then outputs; Step 4: Use the data fused in Step 2 to train and validate the improved YOLOv5s network established in Step 3; Step 5: Use the improved YOLOv5s network to detect targets.
2. The target detection method according to claim 1, characterized in that, in Step 1, the original image data collected by the Gm-APD lidar is preprocessed using the adaptive histogram equalization method, specifically, Step 1.1: Divide the original image data, i.e., the picture, collected by the Gm-APD lidar into several sub-block regions; Step 1.2: Add contrast limits to each small block region, and clip the histogram statistically obtained in the sub-block so that its amplitude is lower than the threshold upper limit.
3. The target detection method according to claim 1, characterized in that, the specific normalization process in Step RH2.2 is as follows: The normalization calculation of the two-dimensional matrix is: Where I(i,j) is the pixel value at position (i,j), y max = 255, y min = 0, x max is the maximum value of I(i,j), x min is the minimum value of I(i,j).
4. The target detection method according to claim 1, characterized in that, for the attention module SE, the input is x, the number of its characteristic channels is c1, and a feature with a number of characteristic channels of c2 is obtained through a series of convolutions. The importance of each characteristic channel is automatically obtained through learning, and then the useful features are enhanced according to the importance and the features that are not very useful for the current task are suppressed.
5. The object detection method according to claim 2, wherein, the loss function of the improved YOLOv5s network specifically includes classification loss, localization loss, and object confidence loss. The binary cross-entropy loss function is used to calculate the loss of class probability and object confidence score, and GIOU is used as the loss function for bounding box regression. The calculation formula is as follows: Among them, A is the predicted bounding box, B is the ground truth bounding box, C is the smallest enclosing box of A and B, and GIOU loss ranges from [0, 2]. When the two bounding boxes completely overlap, the minimum value of 0 is taken. When the sides of the two bounding boxes are externally tangent, the loss function value is 1. When the two bounding boxes are separated and far apart, the loss function value is 2.
6. The object detection method according to claim 5, wherein, the detection result of using the dataset fused in step 2 in step 4 to establish the improved YOLOv5s network in step 3 is that the detection performance of the model is mainly evaluated by accuracy P, recall R, and F1 score in object detection. The formula is expressed as follows: where P is the accuracy, R is the recall, TP is the number of true positive samples, FP is the number of false positive samples, and FN is the number of false negative samples.
7. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method steps described in any one of claims 1-5.
Citation Information
Patent Citations
Gm-APD array laser radar imaging method and Gm-APD array laser radar imaging system under strong background noise
CN110554404A
Lightweight target detection method
CN114120019A