Lane Line Detection and Lane Auxiliary Keeping Method Based on UAV Aerial Images

Through the improved deep learning model and the dual-branch network structure of multi-scale U-Net, the problem of lane line detection accuracy and image quality in aerial images of drones is solved, and high-precision and robust lane line detection is achieved, supporting lane keeping of autonomous vehicles.

CN116434088BActive Publication Date: 2025-06-17CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310406160.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-17
Publication Date
2025-06-17
Estimated Expiration
2043-04-17

AI Technical Summary

Technical Problem

The existing lane line detection algorithm is mainly based on road image data from the perspective of ground vehicles. When applied to aerial images of drones, the detection accuracy is reduced, and the images acquired by drones are susceptible to noise interference, affecting detection tasks.

Method used

The improved deep learning model is used for lane line detection, including the construction and training of lane line detection deep learning model from aerial viewing angle, the image denoising is used with a multi-scale U-Net dual-branch network structure, and combined with a lane-assisted retention method of air-ground collaboration.

Benefits of technology

It improves the problem of field of view during lane line detection, improves detection accuracy and robustness, solves the problem of image quality degradation, enhances the accuracy of lane line detection, and provides strong guarantees for lane maintenance of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116434088B_ABST
    Figure CN116434088B_ABST
Patent Text Reader

Abstract

The present invention relates to a lane line detection and lane auxiliary keeping method based on UAV aerial images, belonging to the field of computer vision. Aiming at the problem of limited vision in traditional lane line detection based on the perspective of in-vehicle cameras, this method uses the UAV perspective to replace the in-vehicle camera perspective for road image acquisition, designs a new lane line detection network and loss function, and at the same time designs an aerial image denoising model for the problem of UAV image transmission interference; then uses a ground workstation loaded with a lane line detection module and an aerial image denoising module to receive the road images transmitted by the UAV for lane line detection; next, calculates the offset between the lane center point and the image center point based on the output result of the lane line detection module, and finally generates an offset signal according to the offset and sends it to the driverless vehicle to assist it in lane keeping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and relates to a lane line detection and lane assist keeping method based on drone aerial images. Background Art

[0002] With the development of the Intelligent Transportation System (ITS), urban intelligent transportation has been gradually popularized, and intelligent transportation technologies represented by vehicle autonomous driving have also received extensive attention. At the same time, computer vision technology has been widely applied in the field of vehicle autonomous driving. By using a vision sensor to obtain road image data of an autonomous driving vehicle during driving for lane line detection and using the detection result to assist the autonomous driving vehicle in lane keeping, the safety and reliability of the autonomous driving vehicle can be greatly improved.

[0003] Currently, the lane assist keeping method for autonomous driving vehicles mainly provides road area and boundary information for autonomous driving vehicles by detecting lane lines on the road. The mainstream lane line detection methods include radar-based lane line detection and image-based lane line detection. Among them, the image-based lane line detection method is mainly realized by processing road images captured by an in-vehicle camera. Since the shooting field of view of the in-vehicle camera is small, the line of sight is easily blocked when the following distance is too close or the height of the vehicle in front is high. Therefore, it will be hindered in practical applications. With the development of drone technology, considering the long-distance vision advantage of drones, using drones to replace in-vehicle cameras to obtain road images for lane line detection has become a feasible solution.

[0004] However, the existing lane line detection algorithms are mainly based on road image data from the perspective of ground vehicles. When applying them to drone aerial images, the lane line detection accuracy often decreases. At the same time, due to the particularity and complexity of the drone working environment, the road images obtained from drones are also vulnerable to noise interference during imaging and transmission, which further leads to a decrease in image quality and affects the subsequent lane line detection task. In addition, how to use the detection result of the lane line detection method based on drone aerial images to assist the autonomous driving vehicle in lane keeping is also an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a lane line detection and lane assist keeping method based on drone aerial images, which has the characteristics of wide sensing range and strong robustness, and has strong practical significance and good application prospects in the field of lane line detection and vehicle autonomous driving.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] Lane line detection and lane auxiliary keeping method based on UAV aerial images, the method comprising the following steps:

[0008] S1: Lane line detection model construction part:

[0009] S11 Construct and train a lane line detection deep learning model from the aerial perspective:

[0010] S12: Train the model with at least 2000 labeled lane line pictures from the UAV aerial perspective, and set them as the training set and the test set according to a ratio of 10:1;

[0011] S13: Design a classification loss function L cls and an instance segmentation loss function L seg for the main network and the auxiliary network respectively, where the main network uses the cross entropy between the predicted value of the lane line probability distribution and the true lane line position in the training data set as the loss function L cls, as shown in the following formula:

[0012]

[0013] In the formula, P i,j is the predicted value of the lane line position in the main network, T i,j is the label of the lane line position in the training set, the subscript i represents the serial number of a certain lane line in the image, the subscript j represents the serial number of a certain preset row anchor, C is the number of lane lines in the image, h is the number of preset row anchors, and N is the number of samples predicted by the model;

[0014] Secondly, considering the continuity and smoothness of the lane lines, use L cont as the continuity loss function of the main network, as shown in the following formula:

[0015]

[0016] Use this loss function to calculate the predicted values of the probability distribution of the same lane line position between adjacent rows of the image, and aim to narrow the gap between the predicted values of the probability distribution of the same lane line position between adjacent rows when training the network;

[0017] In addition, to make the prediction of the lane lines smoother, the increasing or decreasing relationship of the probability distribution values of the same lane line position between adjacent rows of the image should be made consistent. For this purpose, use L str as the structural loss function of the main network, as shown in the following formula:

[0018]

[0019] Select the cross entropy as the loss function of the auxiliary network, as shown in the following formula:

[0020]

[0021] where p i,j is the segmentation prediction output of the auxiliary network, and t i,j is the instance segmentation label in the dataset. Combining the loss functions of the main network and the auxiliary network, the overall loss function L of the lane detection model is:

[0022] L = W cls L cls + W cont L cont + W str L str + W seg L seg (5)

[0023] where W cls , W cont , W str , W seg are the weight coefficients of each loss function component;

[0024] S14: When training the lane detection model, select Adam as the network training optimizer and set the number of iterations epoch to 1000 times;

[0025] S2: Drone aerial image denoising model construction part:

[0026] S21: Construct and train the drone aerial image denoising model:

[0027] In order to retain more details while denoising the image, the denoising network adopts a dual-branch network structure based on multi-scale U-Net. The specific architecture is as follows:

[0028] The first branch network first performs 3 consecutive downsamplings on the input noisy image through 3 consecutive 3*3 convolutional kernels and a max pooling layer with a scale of 2 to extract features, and uses the Gaussian Error Linear Unit GELU as the activation function. Subsequently, it performs 3 consecutive upsamplings through transposed convolutions of the same scale, and the activation function also uses GELU, thus obtaining the output of the first network branch;

[0029] In order to retain more gradient information, residual connections are established between the second layer and the penultimate layer of the first branch network;

[0030] The second branch network first operates on the input image using 3 consecutive 3*3 dilated convolutional kernels to expand the receptive field of the denoising model. Subsequently, the output result of the dilated convolution is input into the same number of conventional 3*3 convolutional kernels, thereby obtaining the output of the second branch network;

[0031] To improve the performance of the denoising model, a channel attention module CA is introduced between three consecutive 3×3 dilated convolution modules in the second branch network;

[0032] For the outputs of the first branch network and the second branch network, the maximum and average values in the channel dimension are calculated respectively, the obtained maximum value matrix and average value matrix are concatenated, and a 1×1 convolutional kernel is used to perform a convolution operation on the result, thereby obtaining the fused aerial photo denoised image;

[0033] S22: The training of the denoising model requires at least 3000 pairs of noisy / clean image pairs, where the size of each image is 200×200;

[0034] S3: The lane assistance maintenance part based on air-ground cooperation:

[0035] The trained lane line detection model and the unmanned aerial vehicle (UAV) aerial photo denoising model are loaded into the on-vehicle ground workstation. At the same time, the UAV is made to fly behind the autonomous driving vehicle traveling on the ground and capture road scene images in real time;

[0036] During hovering and flight, the UAV receives the real-time GPS information of the autonomous driving vehicle and adjusts its own position accordingly to achieve following the vehicle;

[0037] The on-board control panel of the UAV transmits the road scene image to the on-vehicle ground workstation through socket, and the socket transmission protocol of the image is TCP;

[0038] After receiving the road scene image captured by the UAV, the ground workstation uses the natural image quality evaluation index (Natural Image Quality Evaluator, NIQE) to evaluate the quality of the captured image and compares it with the preset upper threshold value of NIQE. The calculation formula of NIQE is as follows:

[0039]

[0040] In the formula, v1, v2 and ∑1, ∑2 represent the average vectors and covariance matrices of the natural image multi-view geometry MVG model and the MVG model of the low-quality image to be evaluated respectively;

[0041] If the NIQE value of the road scene image is greater than the threshold, it is input into the trained UAV aerial photo denoising model for image denoising, and the denoised road image is input into the lane line detection model;

[0042] If the NIQE value of the road scene image is less than the threshold, it is directly input into the trained lane line detection model;

[0043] The ground workstation filters the lane line identification points output by the lane line detection model, filters out the identification points that do not belong to the lane on which the autonomous vehicle is currently driving, and distinguishes between the left and right lane lines. The steps are as follows:

[0044] Establish a Cartesian coordinate system for the image output by the lane line detection model, and set the four boundary anchor points (X1, Y1), (X2, Y2), (X3, Y3), (X4, Y4) of the ROI area, where the ROI area is the region of interest;

[0045] Compare the size of the coordinates (X, Y) of the lane line identification points output by the lane line detection model with the four boundary anchor points of the ROI area to determine whether it falls within the ROI area;

[0046] If the lane line identification point falls within the ROI area, retain the X value and Y value of its coordinates;

[0047] Arrange the discrete lane line identification points that fall within the ROI area in ascending order according to the size of their X-axis coordinate values;

[0048] Calculate the Euclidean distance D between adjacent lane line identification points in turn i , as shown in the following formula:

[0049]

[0050] In the formula, the subscript i represents the index of the lane line identification point, X i , X i+1 represents the abscissa of the lane line identification point, Y i , Y i+1 represents the ordinate of the lane line identification point;

[0051] Take the lane line identification point represented by the subscript where the Euclidean distance D i is the maximum as the demarcation point between the left and right lane lines;

[0052] Perform polynomial fitting on the coordinates (X, Y) of the left and right lane lines respectively, and calculate the lane line curvature based on this. At the same time, calculate the position coordinates of the center point of the lane according to the fitted coordinate values;

[0053] According to the height during the flight of the drone, set the ratio of the pixel position coordinate difference to the lane width to convert the pixel space to meters;

[0054] Calculate the distance between the center point of the image and the center point of the lane according to the position coordinates. At the same time, set an offset threshold. When the distance between the center point of the image and the center point of the lane is greater than this threshold, the ground workstation will send an offset signal to the vehicle;

[0055] The vehicle offset direction in the offset signal is determined by the difference between the X-axis coordinate value of the center point of the screen and the X-axis coordinate value of the center point of the lane, where the X-axis coordinate value of the center of the screen is X C , and the X-axis coordinate value of the center point of the lane is X L , when X C -X L <0, the vehicle offset signal is a right offset signal, when X C -X L >0, the vehicle offset signal is a left offset signal;

[0056] The ground workstation encapsulates the offset information into a data frame and sends it to the autonomous driving vehicle that has established a TCP connection with it. The data frame encapsulation rule should be based on the control information interface of the autonomous driving vehicle.

[0057] Optionally, the S11 is specifically: using an improved Residual Network-34 residual network as the main network to extract features from the input image. The size of the input image is uniformly 640*360. The feature extraction network is composed of four residual blocks connected in sequence. The composition of each residual block is as follows:

[0058] The first residual block is composed of 3 groups of dual 3*3 convolutional kernels connected, and the output channel is 64;

[0059] The second residual block is composed of 4 groups of dual 3*3 convolutional kernels connected, and the output channel is 128;

[0060] The third residual block is a dilated residual block, which is composed of 6 groups of dual 3*3 convolutional kernels with dilation coefficients connected. The dilation coefficients are alternately set to 2, 3, 5, and the output channel is 256;

[0061] The fourth residual block is a dilated residual block, which is composed of 3 groups of dual 3*3 convolutional kernels with dilation coefficients connected. Its dilation coefficient setting is the same as that of the third residual block, and the output channel is 512;

[0062] Introduce a CBAM connection between each residual block, where CBAM is composed of a channel attention module and a spatial attention module;

[0063] Further use a fully connected layer after the improved residual network to expand the receptive field of the network;

[0064] To enhance the lane line visual features during the network training process, an auxiliary network is added to the main network. The last layer of each residual block is upsampled to the same size as the input feature map and concatenated on the channel, and then it is input into the CBAM module. Finally, a 1*1 convolutional kernel is used to merge 4 channels, and then the segmentation output of the lane line is obtained

[0065] The beneficial effects of the present invention are as follows: The present invention proposes a lane line detection and lane auxiliary keeping method based on UAV aerial images. This method uses the UAV perspective to detect lane lines through an improved deep learning model, and integrates the lane offset calculation and lane auxiliary keeping functions on the basis of lane line detection, which not only improves the problem of limited vision in the lane line detection process, but also ensures a relatively fast detection speed and robustness. At the same time, an image denoising and restoration method based on the U-Net framework is adopted, which can well solve the problem of image quality degradation during the transmission of UAV-captured images, thereby improving the accuracy of lane line detection and providing a strong guarantee for subsequent lane keeping of assisted autonomous vehicles.

[0066] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent description, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. Brief Description of the Drawings

[0067] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0068] Figure 1 is a schematic flowchart of the lane line detection and lane auxiliary keeping method based on UAV aerial images provided by the present invention;

[0069] Figure 2 is a neural network structure diagram designed for the lane line detection part of the present invention;

[0070] Figure 3 is a model structure diagram of the UAV aerial image denoising and restoration module of the present invention;

[0071] Figure 4 is the denoising and restoration result of the noisy image from the aerial perspective of the present invention; (a) is the noisy image, and (b) is the denoised and restored image;

[0072] Figure 5 is a schematic diagram of the detection result of the lane line detection part of the present invention; (a), (b), and (d) are the lane line detection result diagrams in the curved road scenario, and (c) is the lane line detection result diagram in the straight road scenario;

[0073] Figure 6 is a schematic diagram of the UAV assisting an autonomous vehicle to keep in the lane of the present invention. Detailed Embodiment

[0074] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0075] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.

[0076] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the attached drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0077] Please refer to Figure 1 , which is the overall process block diagram of the present invention, mainly composed of an unmanned aerial vehicle (UAV) aerial photography image denoising module and a lane line detection module. In actual processing, two network modules can be designed and trained separately, but it is necessary to ensure that the execution of the lane line detection module is after calculating the NIQE value of the image or the execution of the denoising module.

[0078] S1. Construct and train a deep learning model for lane line detection from an aerial perspective as shown in Figure 2 :

[0079] S11. Use an improved ResidualNetwork-34 (ResNet-34) residual network as the main network to extract features from the input image. The size of the input image is uniformly 640*360. The feature extraction network is composed of four residual blocks connected in sequence. The composition of each residual block is as follows:

[0080] The first residual block is composed of 3 groups of double 3*3 convolutional kernels connected, and the output channel is 64;

[0081] The second residual block is composed of 4 groups of dual 3×3 convolutional kernels connected, and the output channels are 128;

[0082] The third residual block is a dilated residual block, which is composed of 6 groups of dual 3×3 convolutional kernels with dilation coefficients connected. The dilation coefficients are alternately set to 2, 3, and 5, and the output channels are 256;

[0083] The fourth residual block is a dilated residual block, which is composed of 3 groups of dual 3×3 convolutional kernels with dilation coefficients connected. The dilation coefficients are set the same as those of the third residual block, and the output channels are 512;

[0084] A CBAM (Convolutional Block Attention Module) connection is introduced between each residual block, where CBAM is composed of a channel attention module and a spatial attention module combined;

[0085] A fully connected layer is further used after the improved residual network to expand the receptive field of the network;

[0086] To enhance the lane line visual features during the network training process, an auxiliary network is added to the main network. The last layer of each residual block is upsampled to the same size as the input feature map and concatenated on the channels, and then it is input into the CBAM module. Finally, a 1×1 convolutional kernel is used to merge 4 channels, and then the segmentation output of the lane line is obtained.

[0087] S12. The training of the model requires at least 2000 labeled lane line pictures from the perspective of UAV aerial photography, and they are set as the training set and the test set according to a ratio of 10:1.

[0088] S13. For the main network and the auxiliary network, a classification loss function L cls and an instance segmentation loss function L seg are designed respectively, where the main network uses the cross-entropy between the predicted value of the lane line probability distribution and the true lane line position in the training data set as the loss function L cls , as shown in the following formula:

[0089]

[0090] In the formula, P i,j is the predicted value of the lane line position in the main network, T i,j is the label of the lane line position in the training set. The subscript i represents the serial number of a certain lane line in the image, the subscript j represents the serial number of a certain preset row anchor, C is the number of lane lines in the image, h is the number of preset row anchors, and N is the number of samples predicted by the model;

[0091] Secondly, considering the continuity and smoothness of the lane line, Lcont As the continuity loss function of the main network, it is shown as follows:

[0092]

[0093] Using this loss function to calculate the predicted values of the probability distribution of the position of the same lane line between adjacent rows of the image, the goal is to narrow the gap between the predicted values of the probability distribution of the position of the same lane line between adjacent rows during network training;

[0094] In addition, to make the prediction of the lane line smoother, the increasing or decreasing relationship of the probability distribution values of the position of the same lane line between adjacent rows of the image should be made consistent. For this purpose, L str is the structural loss function of the main network, shown as follows:

[0095]

[0096] Select cross-entropy as the loss function of the auxiliary network, shown as follows:

[0097]

[0098] In the formula, p i,j is the segmentation prediction output of the auxiliary network, t i,j is the instance segmentation label in the dataset. Combining the loss functions of the main network and the auxiliary network, the overall loss function L of the lane line detection model is:

[0099] L = W cls L cls + W cont L cont + W str L str + W seg L seg (5)

[0100] where W cls , W cont , W str , W seg are the weight coefficients of each loss function component. In the embodiments of this specification, the initial values are set as W cls = 1, W cont = 1, W str = 0.02, W seg = 1.

[0101] S14. When training the lane line detection model, select Adam as the network training optimizer and set the number of iterations epoch to 1000 times.

[0102] S2 Drone aerial image denoising model construction part:

[0103] S21, construct and train a drone aerial image denoising model as shown in Figure 3 :

[0104] To retain more details while denoising the image, the denoising network adopts a dual-branch network structure based on multi-scale U-Net. The specific architecture is as follows:

[0105] The first branch network first performs three consecutive downsamplings on the input noisy image to extract features through three consecutive 3*3 convolutional kernels and a max pooling layer with a scale of 2, and uses the Gaussian Error Linear Unit (GELU) as the activation function. Subsequently, it performs three consecutive upsamplings through transposed convolutions of the same scale, and the activation function also uses GELU, thereby obtaining the output of the first network branch;

[0106] To retain more gradient information, residual connections are established between the second layer and the penultimate layer of the first branch network;

[0107] The second branch network first operates on the input image using three consecutive 3*3 dilated convolutions to expand the receptive field of the denoising model. Subsequently, the output result of the dilated convolution is input into the same number of conventional 3*3 convolutional kernels, thereby obtaining the output of the second branch network;

[0108] To improve the performance of the denoising model, a Channel Attention (CA) module is introduced between three consecutive 3*3 dilated convolution modules in the second branch network;

[0109] For the outputs of the first branch network and the second branch network, calculate their maximum and average values in the channel dimension respectively, concatenate the obtained maximum value matrix and average value matrix, and perform a convolution operation on the result using a 1*1 convolutional kernel, thereby obtaining the fused aerial denoised image.

[0110] In the embodiments of this specification, the loss function of the denoising network adopts a smooth loss function:

[0111]

[0112] Use this loss function to calculate the error between the training-generated image and the actual clean image, and narrow the gap between the generated image and the clean image, where and are the horizontal and vertical gradient operators, is the output of the denoising network, is the clean image, and ξ is the smoothness, whose default value is 1.

[0113] S22. The training of the denoising model requires at least 3,000 pairs of noisy / clean image pairs, where the size of each image is 200*200. For the specific denoising effect, please refer to Figure 4 . Figure 4 This is the denoising and restoration result of the noisy image from the aerial view of the present invention; (a) is the noisy image, and (b) is the denoised and restored image.

[0114] S3 Lane line assisted maintenance part based on ground-air cooperation:

[0115] Load the trained lane line detection model and the UAV aerial image denoising model into the on-vehicle ground workstation. At the same time, make the UAV fly behind the autonomous driving vehicle driving on the ground and capture road scene images in real time. The cooperation method between the UAV and the autonomous driving vehicle is as Figure 6 shown;

[0116] During hovering and flight, the UAV receives the real-time GPS information of the autonomous driving vehicle and adjusts its own position accordingly to achieve following the vehicle;

[0117] The on-board control panel of the UAV transmits the road scene image to the on-vehicle ground workstation through socket. The socket transmission protocol of the image is TCP;

[0118] After receiving the road scene image captured by the UAV, the ground workstation calculates its NIQE value and compares it with the preset upper threshold of NIQE. The calculation formula of NIQE is as follows:

[0119]

[0120] In the formula, v1, v2 and ∑1, ∑2 represent the average vector and covariance matrix of the natural image Multiple View Geometry (MVG) model and the MVG model of the low-quality image to be evaluated respectively;

[0121] If the NIQE value of the road scene image is greater than the threshold, input it into the trained UAV aerial image denoising model for image denoising, and input the denoised road image into the lane line detection model;

[0122] If the NIQE value of the road scene image is less than the threshold, directly input it into the trained lane line detection model. The output result of the model is as Figure 5 shown. Figure 5 This is the schematic diagram of the detection result of the lane line detection part of the present invention; (a), (b), (d) are the lane line detection result diagrams in the curved road scene, and (c) is the lane line detection result diagram in the straight road scene.

[0123] The ground workstation screens the lane line identification points output by the lane line detection model, filters out the identification points that do not belong to the lane on which the autonomous vehicle is currently driving, and distinguishes the left and right lane lines. The steps are as follows:

[0124] Establish a Cartesian coordinate system for the image output by the lane line detection model, and set the 4 boundary anchor points (X1, Y1), (X2, Y2), (X3, Y3), (X4, Y4) of the ROI region, where ROI is the Region of Interest;

[0125] Compare the size of the coordinates (X, Y) of the lane line identification points output by the lane line detection model with the 4 boundary anchor points of the ROI region to determine whether it falls within the ROI region;

[0126] If the lane line identification point falls within the ROI region, retain the X value and Y value of its coordinates.

[0127] Arrange the discrete lane line identification points that fall within the ROI region in ascending order according to the magnitude of their X-axis coordinate values;

[0128] Successively calculate the Euclidean distance D between adjacent lane line identification points i , as shown in the following formula:

[0129]

[0130] In the formula, the subscript i represents the index of the lane line identification point, X i , X i+1 represents the abscissa of the lane line identification point, and Y i , Y i+1 represents the ordinate of the lane line identification point;

[0131] Take the lane line identification point represented by the subscript where the maximum value of the Euclidean distance D i is located as the demarcation point between the left and right lane lines.

[0132] Perform polynomial fitting on the coordinates (X, Y) of the left and right lane lines respectively, and calculate the lane line curvature based on this. At the same time, calculate the position coordinates of the center point of the lane according to the fitted coordinate values.

[0133] According to the height during the flight of the drone, set the ratio of the pixel position coordinate difference to the lane width to convert the pixel space to meters;

[0134] Calculate the distance between the center point of the image and the center point of the lane according to the position coordinates, and at the same time set an offset threshold. When the distance between the center point of the image and the center point of the lane is greater than this threshold, the ground workstation will send an offset signal to the vehicle;

[0135] The vehicle offset direction in the offset signal is determined by the difference between the X-axis coordinate value of the center point of the screen and the center point of the lane, where the X-axis coordinate value of the center of the screen is X C , and the X-axis coordinate value of the center point of the lane is X L , when X C -X L < 0, the vehicle offset signal is a right offset signal. When X C -X L > 0, the vehicle offset signal is a left offset signal;

[0136] The ground workstation encapsulates the offset information into a data frame and sends it to the autonomous driving vehicle that has established a TCP connection with it. The data frame encapsulation rule should be based on the control information interface of the autonomous driving vehicle.

[0137] It should be noted that the image stream collected by the drone is cut into single-frame images and transmitted to the ground workstation in the form of pictures instead of video streams. In order to improve the real-time performance of the model, it is possible to choose whether to load the drone aerial photography image denoising module according to the specific usage scenario.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A lane line detection and lane assist keeping method based on UAV aerial images, characterized in that: The method includes the following steps: S1: Lane line detection model construction part: S11: Construct and train a lane line detection deep learning model from an aerial perspective: S12: The model is trained with at least 2000 labeled drone aerial perspective lane line images, which are set as the training set and the test set according to a ratio of 10:1; S13: Design classification loss functions \(L\) for the main network and the auxiliary network respectively cls and the instance segmentation loss function \(L\) seg , where the main network uses the cross-entropy between the predicted value of the lane line probability distribution and the true lane line position in the training data set as the loss function \(L\) cls, as shown in the following formula: where P i,j is the predicted value of the lane line position in the main network, T i,j is the label of the lane line position in the training set. The subscript i represents the serial number of a certain lane line in the image, the subscript j represents the serial number of a certain preset row anchor, C is the number of lane lines in the image, h is the number of preset row anchors, and N is the number of samples predicted by the model; Secondly, considering the continuity and smoothness of the lane lines, L cont is used as the continuity loss function of the main network, as shown in the following formula: Use this loss function to calculate the predicted values of the probability distribution of the same lane line position between adjacent rows of the image. When training the network, the goal is to narrow the gap between the predicted values of the probability distribution of the same lane line position for adjacent rows; In addition, to make the prediction of the lane line smoother, the increasing or decreasing relationship of the position probability distribution values of the same lane line between adjacent rows of the image is made consistent, and L str is used as the structural loss function of the main network, as shown in the following formula: Select cross-entropy as the loss function for the auxiliary network, as shown in the following formula: where p i,j is the segmentation prediction output of the auxiliary network, and t i,j is the instance segmentation label in the dataset. Combining the loss functions of the main network and the auxiliary network, the overall loss function L of the lane detection model is: L = W cls L cls + W cont L cont + W str L str + W seg L seg (5) Among which, W cls , W cont , W str , W seg is the weight coefficient of each loss function component; S14: When training the lane line detection model, select Adam as the network training optimizer, and set the number of iterations epoch to 1000 times; S2: Drone aerial image denoising model construction part: S21: Construct and train a drone aerial image denoising model: In order to retain more details while denoising the image, the denoising network adopts a dual-branch network structure based on multi-scale U-Net. The specific architecture is as follows: The first branch network first performs 3 consecutive downsamplings on the input noisy image through 3 consecutive 3*3 convolutional kernels and a max pooling layer with a scale of 2 to extract features, and uses the Gaussian error linear unit GELU as the activation function. Subsequently, it performs 3 consecutive upsamplings through transposed convolutions of the same scale, and the activation function also uses GELU, thereby obtaining the output of the first network branch; In order to retain more gradient information, establish residual connections between the second layer and the penultimate layer of the first branch network; The second branch network first operates on the input image using 3 consecutive 3*3 dilated convolutional kernels to expand the receptive field of the denoising model. Subsequently, the output result of the dilated convolution is input into the same number of conventional 3*3 convolutional kernels, and then the output of the second branch network is obtained; In order to improve the performance of the denoising model, a channel attention module CA is introduced between 3 consecutive 3*3 dilated convolution modules in the second branch network; For the outputs of the first branch network and the second branch network, calculate their maximum and average values in the channel dimension respectively, splice the obtained maximum value matrix and average value matrix, and perform a convolution operation on the result using a 1*1 convolutional kernel, thereby obtaining the fused aerial denoised image; S22: The training of the denoising model requires at least 3000 pairs of noise / clean image pairs, where the size of each image is 200*200; S3: Lane assistance maintenance part based on ground-air cooperation: Load the trained lane line detection model and the drone aerial image denoising model into the on-vehicle ground workstation. At the same time, make the drone fly behind the autonomous driving vehicle traveling on the ground and capture road scene images in real time; The drone receives the real-time GPS information of the autonomous driving vehicle during hovering and flight and adjusts its own position accordingly to achieve following the vehicle; The drone-borne control panel transmits the road scene image to the on-vehicle ground workstation through socket, and the socket transmission protocol of the image is TCP; After the ground workstation receives the road scene image captured by the drone, it uses the natural image quality evaluation index NIQE to evaluate the quality of the captured image and compares it with the pre-set upper threshold of NIQE. The calculation formula of NIQE is as follows: In the formula, v1, v2 and ∑1, Σ2 represent the average vector and covariance matrix of the natural image multi-view geometry MVG model and the MVG model of the low-quality image to be evaluated respectively; If the NIQE value of the road scene image is greater than the threshold, it will be input into the trained drone aerial photography image denoising model for image denoising, and the denoised road image will be input into the lane line detection model; If the NIQE value of the road scene image is less than the threshold, it will be directly input into the trained lane line detection model; The ground workstation screens the lane line identification points output by the lane line detection model, filters out the identification points that do not belong to the lane where the autonomous vehicle is currently driving, and distinguishes the left and right lane lines. The steps are as follows: Establish a Cartesian coordinate system for the image output by the lane line detection model, and set 4 boundary anchor points (X1, Y1), (X2, Y2), (X3, Y3), (X4, Y4) of the ROI area, where the ROI area is the region of interest; Compare the size of the coordinates (X, Y) of the lane line identification points output by the lane line detection model with the 4 boundary anchor points of the ROI area to determine whether it falls within the ROI area; If the lane line identification point falls within the ROI area, retain the X value and Y value of its coordinates; Arrange the discrete lane line identification points falling within the ROI area in ascending order according to the size of their X-axis coordinate values; Calculate the Euclidean distance D between adjacent lane line identification points in sequence i , as shown in the following formula: In the formula, the subscript i represents the index of the lane line identification point, X i , X i+1 represents the abscissa of the lane line identification point, Y i , Y i+1 represents the ordinate of the lane line identification point; Take the Euclidean distance D i The lane line identification point represented by the subscript where the maximum value is located is used as the demarcation point between the left and right lane lines; Perform polynomial fitting on the coordinates (X, Y) of the left and right lane lines respectively, and calculate the lane line curvature based on this. At the same time, calculate the position coordinates of the lane center point according to the fitted coordinate values; According to the height of the drone during flight, set the ratio of the pixel position coordinate difference to the lane width to convert the pixel space to meters; Calculate the distance between the center point of the image and the center point of the lane according to the position coordinates, and set an offset threshold at the same time. When the distance between the center point of the image and the center point of the lane is greater than this threshold, the ground workstation will send an offset signal to the vehicle; The vehicle offset direction in the offset signal is determined by the difference between the X-axis coordinate value of the center point of the screen and the center point of the lane, where the X-axis coordinate value of the center of the screen is X C , and the X-axis coordinate value of the center point of the lane is X L , when X C -X L < 0, the vehicle offset signal is a right offset signal, when X C -X L > 0, the vehicle offset signal is a left offset signal; The ground workstation encapsulates the offset information into a data frame and sends it to the autonomous vehicle that has established a TCP connection with it. The data frame encapsulation rule is based on the control information interface of the autonomous vehicle.

2. The lane line detection and lane auxiliary keeping method based on UAV aerial images according to claim 1, characterized in that: The specific content of S11 is: use the improved Residual Network-34 residual network as the main network to extract features from the input image. The size of the input image is uniformly 640*360. The feature extraction network is composed of four residual blocks connected in sequence. The composition of each residual block is as follows: The first residual block is composed of 3 groups of dual 3*3 convolutional kernels connected, and the output channel is 64; The second residual block is composed of 4 groups of dual 3*3 convolutional kernels connected, and the output channel is 128; The third residual block is a dilated residual block, which is composed of 6 groups of dual 3*3 convolutional kernels with dilation coefficients connected. The dilation coefficients are alternately set to 2, 3, 5, and the output channel is 256; The fourth residual block is a dilated residual block, which is composed of three groups of dual 3*3 convolutional kernels with dilation coefficients. The dilation coefficients are set the same as those of the third residual block, and the output channels are 512; A CBAM connection is introduced between each residual block, where CBAM is composed of a channel attention module and a spatial attention module; A fully connected layer is further used after the improved residual network to expand the receptive field of the network; To enhance the lane line visual features during the network training process, an auxiliary network is added to the main network. The last layer of each residual block is upsampled to the same size as the input feature map and concatenated on the channel, and then it is input into the CBAM module. Finally, a 1*1 convolutional kernel is used to merge four channels to obtain the segmentation output of the lane line.