Road recognition method for aerial images based on deep learning

By fusion of multi-source data and an improved U2Net network, combined with the SimAM attention mechanism, the accuracy and robustness of road recognition in high-altitude aerial images are improved, which solves the problem of insufficient multi-source data fusion in existing technologies and provides more accurate road recognition results.

CN119131577BActive Publication Date: 2025-09-05XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411134998.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-09-05
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

Existing high-altitude aerial image road recognition methods based on deep learning have insufficient multi-source data fusion and recognition accuracy when dealing with complex backgrounds and diverse ground features.

Method used

Through multi-source data collection, image enhancement, feature extraction and rasterization processing, an improved U2Net network was built and the SimAM attention mechanism module was introduced. The road prediction results of trajectory raster images and high-altitude aerial images were fused using a pixel-level weighted combination method.

Benefits of technology

The accuracy and robustness of road recognition in high-altitude aerial images are improved, noise interference is reduced, and more comprehensive and accurate road recognition information is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119131577B_ABST
    Figure CN119131577B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for road recognition in aerial images based on deep learning, comprising the following steps: Step 1, multi-source data acquisition; Step 2, aerial image enhancement; Step 3, feature extraction and rasterization of the collected vehicle trajectory data; Step 4, adding a SimAM attention mechanism module after the encoder of the U2Net network to build an improved U2Net network, and obtaining road prediction images corresponding to the aerial image and the trajectory raster image, respectively; Step 5, fusing the trajectory raster image obtained in Step 4 with the road prediction image corresponding to the aerial image using a pixel-level weighted combination method to obtain a merged road recognition prediction result. The present invention belongs to the field of visual and remote sensing image processing technology, and solves the problem of insufficient multi-source data fusion and insufficient recognition accuracy in existing technologies when dealing with complex backgrounds and diverse land features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vision and remote sensing image processing, and relates to a high-altitude aerial image road recognition method based on deep learning. Background Art

[0002] With the rapid development of drone and satellite remote sensing technologies, aerial imagery has become widely used in urban planning, traffic management, and disaster response. Aerial photographs and drone images are widely used to generate and update road information. However, due to the diversity of land use and land cover types and the intricate details of the roads to be extracted, such as complex structures and texture features, quickly and accurately detecting and extracting road features from aerial imagery remains a technical challenge.

[0003] In recent years, deep learning technology has made significant progress in image processing and computer vision. In particular, convolutional neural networks (CNNs) have demonstrated excellent performance in object detection and image segmentation tasks. While some existing deep learning-based road recognition methods have improved recognition accuracy to a certain extent, they still suffer from problems such as the high demand for large-scale datasets, long model training time, and insufficient robustness in various environments.

[0004] Therefore, there is an urgent need to develop a new method that can perform data enhancement on original high-altitude aerial images, can cope with complex image backgrounds through the fusion of multi-source data, and improve the performance of road recognition in high-altitude aerial images. Summary of the Invention

[0005] The purpose of the present invention is to provide a road recognition method for high-altitude aerial images based on deep learning, which solves the problems of insufficient multi-source data fusion and insufficient recognition accuracy in the existing technology when dealing with complex backgrounds and diverse land features.

[0006] The technical solution adopted by the present invention is a method for road recognition in high-altitude aerial images based on deep learning, which is implemented according to the following steps:

[0007] Step 1: multi-source data collection;

[0008] Step 2: Enhance the aerial image;

[0009] Step 3: extract features and perform rasterization processing on the collected vehicle trajectory data;

[0010] Step 4: Add the SimAM attention mechanism module after the encoder of the U2Net network to build an improved U2Net network, and obtain the road prediction images corresponding to the aerial image and the trajectory raster image respectively;

[0011] In step 5, the trajectory grid image obtained in step 4 and the road prediction image corresponding to the high-altitude aerial image are fused using a pixel-level weighted combination method to obtain a merged road recognition prediction result.

[0012] The beneficial effect of the present invention is that, by collecting vehicle trajectory data and high-altitude aerial photography data, the trajectory data and aerial images are combined to provide more comprehensive and accurate road recognition information; at the same time, in order to reduce the influence of noise and irrelevant features on the recognition results, the fractional-order calculus image enhancement method and the SimAM attention mechanism module are introduced in the image enhancement and feature extraction module to selectively enhance the useful features of the road and suppress irrelevant features; then, the fusion of the trajectory data and aerial image road recognition results is achieved through a pixel-level weighted fusion method. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a schematic diagram of the overall process of the method of the present invention;

[0014] Figure 2 is the original high-altitude aerial image mentioned in the method of the present invention;

[0015] Figure 3 The method of the present invention is directed to Figure 2 Image enhanced using fractional differentials;

[0016] Figure 4 is a flow chart of road prediction and recognition in the method of the present invention;

[0017] Figure 5 The aerial image of the urban road in Example 1 of the present invention;

[0018] Figure 6 This is the road recognition result of Example 1 of the present invention using the deeplabv3+ algorithm;

[0019] Figure 7 This is the road recognition result of Example 1 of the present invention using the SegFormer algorithm;

[0020] Figure 8 is the road recognition result of Example 1 of the present invention under the method of the present invention;

[0021] Figure 9 This is the aerial image of the mountain road in Example 2 of the present invention;

[0022] Figure 10 This is the road recognition result of Example 2 of the present invention using the deeplabv3+ algorithm;

[0023] Figure 11 This is the road recognition result of Example 2 of the present invention using the SegFormer algorithm;

[0024] Figure 12 is the road recognition result of Example 2 of the present invention under the method of the present invention;

[0025] Figure 13 This is the aerial image of the rural road in Example 3 of the present invention;

[0026] Figure 14 This is the road recognition result of Example 3 of the present invention using the deeplabv3+ algorithm;

[0027] Figure 15 This is the road recognition result of Example 3 of the present invention using the SegFormer algorithm;

[0028] Figure 16 This is the road recognition result of Example 3 of the present invention using the method of the present invention. DETAILED DESCRIPTION

[0029] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] The high-altitude aerial image road recognition method of the present invention based on deep learning is implemented according to the following steps:

[0031] Step 1: Multi-source data collection,

[0032] The collected data includes vehicle trajectory data (GPS data) and high-altitude aerial photography data (drone) including high-resolution visible light imagery. By combining trajectory data and aerial imagery, more comprehensive and accurate road identification information is provided.

[0033] Step 2: Enhance the aerial image.

[0034] The fractional-order calculus image enhancement method is used to enhance the high-altitude aerial image data. The specific process is as follows:

[0035] 2.1) Determine whether the image is a grayscale image. If not, convert the image into a grayscale image.

[0036] 2.2) Find the main direction of the image to be enhanced and determine the weighted multiple of the main direction;

[0037] 2.3) Modify the fractional order template accordingly according to the weighted multiple of the main direction of the image to be enhanced;

[0038] 2.4) Apply the new fractional-order template to the image to be enhanced to obtain the enhanced result image.

[0039] Step 3: trajectory feature selection and trajectory data rasterization.

[0040] The collected vehicle trajectory data is subjected to feature extraction and rasterization processing to be converted into image data. The specific process is as follows:

[0041] 3.1) Format the vehicle trajectory data into a standard form, filter and denoise the standard form trajectory data to improve data accuracy. The expression of vehicle trajectory data is:

[0042] Trajectory={t1, t2, t3…t n}

[0043] Among them, t k Represents a trajectory point in the trajectory data, k = 1, 2, 3, ..., n, whose attributes consist of vehicle number, collection time, longitude coordinate, latitude coordinate, driving speed and driving direction at the time of collection;

[0044] 3.2) Calculate and obtain the characteristics of vehicle trajectory data, including speed, direction, density, average speed, standard deviation of speed, etc.

[0045] 3.3) Using the rasterization method, the discrete trajectory point data is converted into grayscale image data as the input of the neural network according to the characteristics of the trajectory data. The function F(x,y) is used to represent the digital image. F(x,y) represents the grayscale of the image at the coordinate (x,y). The longitude and latitude coordinates (p lon ,p lat ) is converted to digital image coordinates (x, y) as shown below:

[0046] x=floor((p lon -Lon min )*Interval lon )

[0047] y=floor((Lat max -p lat )*Interval lat )

[0048] Among them, floor is the rounding function, Lon min Represents the minimum longitude value, Lat max Represents the maximum latitude value; Interval lon Indicates the number of pixels corresponding to 1 longitude, Interval lat Represents the number of pixels corresponding to 1 degree of latitude.

[0049] Step 4: Build the U2Net network to extract image features.

[0050] Adding the SimAM attention mechanism module after the encoder of the U2Net network enhances the existing U2Net network's focus on key features and builds an improved U2Net network. The specific process is as follows:

[0051] 4.1) Build a U2Net network (full name: two-layer nested U-shaped network) and use it to obtain feature information of high-altitude aerial photography data. The residual U-block captures multi-scale information without reducing the resolution of the feature map, to some extent compensating for the lack of attention to local details compared to global information. The dilation mode of the residual U-block is used in the 5th layer (En5) and 6th layer (En6) of the U2Net network encoder and the 5th layer (De5) of the U2Net network decoder. Dilated convolution replaces pooling and upsampling operations, resulting in a residual U-block with four feature levels (RSU-4F). All intermediate feature maps of the residual U-block have the same resolution as its input feature map.

[0052] 4.2) To filter out feature combinations that are helpful for feature recognition, enhance the U2Net network's ability to express road features, and reduce the interference of background noise, a SimAM attention mechanism module is added after the U2Net network encoder. The SimAM attention mechanism module is a 3D attention module that directly infers 3-D weights from neurons. That is, the SimAM attention mechanism module considers both spatial and channel dimensions, refines neurons, and infers effective 3-D weights by defining an energy function to derive a closed-form solution. The energy function of each neuron is defined as follows:

[0053]

[0054] Where, and is a linear transformation of t and xi, where t is the target neuron in a single channel of the input feature, and xi is the other neurons in a single channel of the input feature; i is the index in the spatial dimension, M = H × W is the number of neurons in the channel, and ω t and b t are the weights and biases of the transformation, where all values ​​are scalars, and the improved U2Net network is obtained;

[0055] 4.3) The aerial image and trajectory grid image enhanced in step 2 are input into the improved U2Net network respectively to obtain the road prediction images corresponding to the aerial image and trajectory grid image respectively.

[0056] Step 5: fusion of results.

[0057] The trajectory grid image obtained in step 4 and the road prediction image corresponding to the high-altitude aerial image are fused using a pixel-level weighted combination method to obtain the merged road recognition prediction result. The specific process is:

[0058] The road recognition results of the trajectory raster image and the aerial image are fused using a pixel-level weighted fusion method. The intensity value of each pixel in the fused image is the weighted average of the intensity values ​​of the corresponding pixels in the trajectory data road extraction image and the remote sensing image road extraction image.

[0059] Assume that the pixel intensity value matrix of the trajectory data road extraction image is T, the pixel intensity value matrix of the remote sensing image road extraction image is R, and the pixel intensity value matrix of the fused image is F. The function of the fusion process is as follows:

[0060] F(i,j)=α·T(i,j)+β·R(i,j)

[0061] Where F(i, j) represents the pixel intensity value of the fused image at position (i, j), T(i, j) represents the pixel intensity value of the trajectory data road extraction image at position (i, j), R(i, j) represents the pixel intensity value of the remote sensing image road extraction image at position (i, j), α and β represent the weighting coefficients of the trajectory data and remote sensing image, respectively, and satisfy α + β = 1;

[0062] From then on, the final road recognition prediction result is obtained.

[0063] Example 1

[0064] The implementation object of this embodiment 1 is a high-altitude aerial image of urban roads. The high-altitude aerial image road recognition method based on deep learning is implemented according to the following steps:

[0065] Step 1: Multi-source data collection: Collect the road vehicle trajectory data and (UAV) high-altitude aerial photography data of Example 1. The trajectory data comes from the vehicle's GPS record, and the high-altitude aerial photography data includes high-resolution visible light images.

[0066] The collected data includes vehicle trajectory data and high-altitude aerial photography data. The trajectory data is derived from the vehicle's GPS records and includes the vehicle number, acquisition time, longitude and latitude coordinates, speed, and direction at the time of acquisition. The high-altitude aerial photography data includes high-resolution visible light imagery. By combining trajectory data and aerial imagery, more comprehensive and accurate road identification information is provided.

[0067] Step 2: High-altitude aerial image enhancement: The high-altitude aerial image data of Example 1 is enhanced using a fractional-order calculus image enhancement method.

[0068] 2.1) After confirming that it is not a grayscale image, convert the image into a grayscale image first.

[0069] 2.2) Find the main direction of the aerial data and determine the weighted multiple of the main direction. The main direction is determined by the direction histogram of the gradient of the aerial data. Let the image gradient be Calculate the horizontal gradient of an image and vertical gradient Its direction θ is:

[0070]

[0071] 2.3) Modify the fractional template accordingly based on the weighted multiple of the main direction of the high-altitude aerial photography data. Assuming the weighted multiple of the main direction is k, the formula of the fractional template H(x,y) is as follows:

[0072] H(x,y)=H0(x,y)·k(θ)

[0073] Among them, H0(x,y) is the original fractional-order template, k(θ) is the weighted multiple adjusted according to the main direction, and (x,y) represents the pixel coordinates of the image.

[0074] 2.4) The new fractional order template is applied to the high-altitude aerial photography data of Example 1 to obtain an enhanced result. The specific operation is to perform a convolution operation on the image as shown in the following formula:

[0075] I enhanced (x,y)=I(x,y)*H(x,y)

[0076] Among them, * is the convolution operation, I(x,y) is the original image, I enhanced (x,y) is the enhanced image.

[0077] Step 3: Track feature selection and track data rasterization. The collected vehicle track data is subjected to feature extraction and rasterization processing to convert it into image data.

[0078] 3.1) Format the vehicle trajectory data into a standard format, and then perform filtering and denoising on the standard trajectory data. The vehicle's driving trajectory data is represented as:

[0079] Trajectory={t1, t2, t3…t n}

[0080] Among them, t k Represents a trajectory point in the trajectory data. Its attributes consist of vehicle number, collection time, longitude coordinate, latitude coordinate, driving speed and driving direction at the time of collection.

[0081] 3.2) Calculate and obtain trajectory data features, including speed, direction, density, average speed, standard deviation of speed, etc.

[0082] 3.3) Use rasterization method to convert the discrete trajectory point data into grayscale image data as the input of neural network according to the characteristics of trajectory point data. Use function F(x,y) to represent digital image, F(x,y) represents the grayscale of the image at coordinate (x,y), and the longitude and latitude coordinates (p lon ,p lat ) is converted to digital image coordinates (x, y) as shown below:

[0083] x=floor((p lon -Lon min )*Interval lon )

[0084] y=floor((Lat max -p lat )*Interval lat )

[0085] Among them, floor is the rounding function, Lon min and Lat max Represents the minimum longitude value and the maximum latitude value respectively. lon Indicates the number of pixels corresponding to 1 longitude, Interval lat Represents the number of pixels corresponding to 1 degree of latitude.

[0086] Step 4: Build the U2Net network. Add the SimAM attention mechanism module after the U2Net encoder to enhance the U2Net network's focus on key features and build an improved U2Net network.

[0087] 4.1) U2Net (full name: double-layer nested U-shaped network) uses a double-layer nested U-shaped network to obtain feature information of high-altitude aerial photography data. The residual U block captures multi-scale information without reducing the resolution of the feature map, to a certain extent making up for the lack of attention to local details compared to global information. The expansion mode of the residual U block is used in the 5th layer (En5) and 6th layer (En6) of the U2Net network encoder and the 5th layer (De5) of the U2Net network decoder network. The expansion convolution is used instead of the pooling and upsampling operations, so that all intermediate feature maps of the residual U block with four feature levels (RSU-4F) have the same resolution as its input feature map.

[0088] 4.2) To filter out feature combinations that are helpful for feature recognition, enhance the U2Net network's ability to express road features, and reduce the interference of background noise, the SimAM attention mechanism module is integrated after the U2Net network encoder. The SimAM attention mechanism module is a 3D attention module that directly infers 3-D weights from neurons. That is, the SimAM attention mechanism module considers both spatial and channel dimensions, refines neurons, and infers effective 3-D weights by defining an energy function to derive a closed-form solution. The energy function of each neuron is defined as follows:

[0089]

[0090] Where, and is a linear transformation of t and xi, where t and xi are the target neuron and other neurons in a single channel of the input feature. i is the index in the spatial dimension, M = H × W is the number of neurons in the channel, ω t and b t are the weights and biases of the transformation, where all values ​​are scalars.

[0091] 4.3) The trajectory raster image after rasterization of the trajectory data and the high-altitude aerial image enhanced by step 2 are respectively input into the improved U2Net network to obtain the road prediction images corresponding to the high-altitude aerial image and the trajectory raster image.

[0092] Step 5: Result fusion. The trajectory raster image and the aerial photography data enhanced in step 2 are fused with the corresponding road prediction image obtained in step 4 using a pixel-level weighted combination method to obtain the combined road recognition prediction result.

[0093] Using pixel-level weighted fusion, the trajectory raster image after rasterization of the trajectory data and the road recognition results of the high-altitude aerial photography data after image enhancement in step 2 are fused. Each pixel intensity value in the fused image is the weighted average of the intensity values ​​of the corresponding pixels in the road extraction image of the trajectory raster image after rasterization of the trajectory data and the high-altitude aerial photography image after image enhancement in step 2. Let the pixel intensity value matrix of the road extraction image of the trajectory raster image after rasterization of the trajectory data be T, the pixel intensity value matrix of the high-altitude aerial photography data after image enhancement in step 2 be R, and the pixel intensity value matrix of the fused image be F. The fusion formula is as follows:

[0094] F(i,j)=α·T(i,j)+β·R(i,j)

[0095] Where F(i, j) represents the pixel intensity value of the fused image at position (i, j), T(i, j) represents the pixel intensity value of the trajectory grid image at position (i, j) after the trajectory data is rasterized, R(i, j) represents the pixel intensity value of the high-altitude aerial photography data at position (i, j) after the image enhancement in step 2, α and β represent the weighting coefficients of the trajectory grid image and the high-altitude aerial photography data after the image enhancement in step 2, respectively, and satisfy α+β=1.

[0096] Below is Figure 5-Figure 8 The experimental description is made by taking the image shown as an example. According to the above steps, the road recognition results of the method of the present invention for urban high-altitude aerial images are obtained, and compared with the road recognition results of the existing high-altitude aerial road recognition methods deeplabv3+ algorithm and SegFormer algorithm for urban high-altitude aerial images. Figure 5 High-altitude aerial images of town roads taken by drones; Figure 6 This is the road recognition result of the existing deeplabv3+ algorithm; Figure 7 This is the road recognition result of the existing SegFormer algorithm; Figure 8 2 is the road recognition result of the method of the present invention. It can be seen that the method of the present invention has better road recognition effect for urban high-altitude aerial images.

[0097] At the same time, this Example 1 compares the method of the present invention with the existing deeplabv3+ algorithm and SegFormer algorithm through precision (P), recall (R), F1 score (F1), and mean absolute error (MAE). Precision (P) measures the proportion of pixels predicted by the algorithm as roads that are actually roads. A higher precision means that the algorithm has fewer errors in road prediction. The calculation formula is:

[0098]

[0099] Among them, TP (True Positives) is the number of pixels correctly detected as roads, FP (False Positives) is the number of false positives, the number of pixels incorrectly detected as roads. The recall rate measures the proportion of pixels that are correctly detected as roads among the actual pixels. The calculation formula is:

[0100]

[0101] TP (True Positives) represents true positives, which are the number of pixels correctly detected as roads. FN (False Negatives) represents false negatives, which are the number of pixels not correctly detected as roads. The F1 score (F1) is used to strike a balance between precision and recall. When one indicator is high and the other is low, the F1 score can provide a comprehensive evaluation. The calculation formula is as follows:

[0102]

[0103] Among them, P is the precision rate and R is the recall rate.

[0104] The mean absolute error (MAE) measures the average of the absolute values ​​of the differences between the predicted value and the true value, and is used to evaluate the accuracy of road detection. The calculation formula is as follows:

[0105]

[0106] Where n is the total number of pixels, y i is the true value of the i-th pixel (road or non-road), is the predicted value (road or non-road) of the i-th pixel.

[0107] Table 1 shows a comparison of experimental results between the proposed method and the existing deeplabv3+ and SegFormer algorithms. As can be seen from Table 1, the proposed method achieves higher precision, recall, and F1 scores than the two existing methods, while its MAE is lower. This indicates that the proposed method is more effective for road recognition in aerial imagery of urban areas.

[0108] Table 1. Comparison of experimental results of different methods in Example 1

[0109]

[0110] Example 2

[0111] According to the steps of Example 1, road recognition is performed on the aerial image of the mountain road.

[0112] Below Figures 9-12 The experimental description is made using the image shown as an example. The road recognition results of the method of the present invention for high-altitude aerial images of mountain roads are obtained according to the steps of Example 1, and compared with the road recognition results of the existing deeplabv3+ algorithm and SegFormer algorithm for high-altitude aerial images of mountain roads. Figure 9 This is an aerial image of a mountain road; Figure 10 This is the road recognition result of the existing deeplabv3+ algorithm; Figure 11This is the road recognition result of the existing SegFormer algorithm; Figure 12 This is the road recognition result of the present invention. It can be seen that the method of the present invention has better road recognition results for high-altitude aerial images in mountainous areas.

[0113] Table 2 shows a comparison of experimental results between the proposed method and the existing deeplabv3+ and SegFormer algorithms. As can be seen from Table 2, the proposed method achieves higher precision, recall, and F1 scores than the two existing methods, while its MAE is lower. This indicates that the proposed method is more effective for road recognition in aerial imagery of mountainous roads.

[0114] Table 2. Comparison of experimental results of different methods in Example 2

[0115]

[0116] Example 3

[0117] According to the steps of Example 1, road recognition is performed on the aerial image of rural roads.

[0118] Below is Figure 13-16 The experimental description is made using the image shown as an example. The road recognition results of the method of the present invention for aerial images of rural roads are obtained according to the steps of Example 1, and compared with the road recognition results of the prior art deeplabv3+ algorithm and SegFormer algorithm for high-altitude aerial images of rural roads. Figure 13 High-altitude aerial images of rural roads; Figure 14 This is the road recognition result of the existing deeplabv3+ algorithm; Figure 15 This is the road recognition result of the existing SegFormer algorithm; Figure 16 This is the road recognition result of the present invention. It can be seen that the method of the present invention has better road recognition results for high-altitude aerial images of rural roads.

[0119] Table 3 shows a comparison of experimental results between the proposed method and the existing deeplabv3+ and SegFormer algorithms. As can be seen from Table 3, the proposed method achieves higher precision, recall, and F1 scores than the two existing methods, while its MAE is lower. This indicates that the proposed method is more effective for road recognition in aerial imagery of rural roads.

[0120] Table 3. Comparison of experimental results of different methods in Example 3

[0121]

[0122] In summary, the method of the present invention uses the fractional calculus image enhancement method to enhance the original image data, effectively enhances the road features and reduces noise interference; at the same time, the SimAM attention mechanism module is introduced into the feature extraction module to further improve the model's ability to express key features; through the pixel-level weighted fusion method, the trajectory data and the recognition results of the aerial image are fused to optimize the final road recognition effect, which significantly improves the problem of inaccurate road feature extraction in complex environments in the existing high-altitude aerial image road recognition method, and improves the accuracy and robustness of the model.

Claims

1. A road recognition method for aerial images based on deep learning, characterized by: Follow these steps to implement: Step 1: Multi-source data collection, The collected data includes vehicle trajectory data and high-altitude aerial photography data. The vehicle trajectory data comes from the vehicle's GPS records, and the high-altitude aerial photography data includes high-resolution visible light images. Step 2: High-altitude aerial image enhancement. The high-altitude aerial image data is enhanced using the fractional order calculus image enhancement method. 2.1) Determine whether the image is a grayscale image. If not, convert the image into a grayscale image. 2.2) Find the main direction of the image to be enhanced and determine the weighted multiple of the main direction; 2.3) Modify the fractional order template accordingly based on the weighted multiple of the main direction of the image to be enhanced; 2.4) Apply the new fractional-order template to the image to be enhanced to obtain the enhanced result image; Step 3: Extract features and rasterize the collected vehicle trajectory data. 3.1) Format the vehicle trajectory data into a standard form, filter and denoise the standard form trajectory data. The expression of the vehicle trajectory data is: in, represents a trajectory point in the trajectory data, k =1,2,3,…, n , whose attributes consist of vehicle number, collection time, longitude coordinate, latitude coordinate, driving speed and driving direction at the time of collection; 3.2) Calculate the characteristics of vehicle trajectory data, including speed, direction, density, average speed, and standard deviation of speed; 3.3) Using the rasterization method, the discrete trajectory point data is converted into grayscale image data as the input of the neural network according to the characteristics of the trajectory data, using the function Represents a digital image, Representing coordinates The grayscale of the image at Convert to digital image coordinates , as shown below: in, floor is the floor function, Represents the minimum longitude value, Represents the maximum latitude value; Indicates the number of pixels corresponding to 1 longitude, Represents the number of pixels corresponding to 1 latitude; Step 4: Add the SimAM attention mechanism module after the encoder of the U2Net network to build an improved U2Net network, and obtain the road prediction images corresponding to the high-altitude aerial image and the trajectory raster image respectively. 4.1) Build a U2Net network and use it to obtain feature information from aerial photography data. The residual U-block captures multi-scale information without reducing feature map resolution. The dilation mode of the residual U-block is used in the 5th and 6th layers of the U2Net encoder and the 5th layer of the U2Net decoder. Dilated convolutions replace pooling and upsampling operations to obtain a residual U-block with four feature levels. All intermediate feature maps of the residual U-block have the same resolution as its input feature map. 4.2) A SimAM attention mechanism module is added after the encoder of the U2Net network. The SimAM attention mechanism module infers 3-D weights directly from neurons. That is, it considers both spatial and channel dimensions, refines neurons, and infers effective 3-D weights by defining an energy function to derive a closed-form solution. The energy function of each neuron is defined as follows: Where, and yes and The linear transformation of is the target neuron in a single channel of the input feature, are the other neurons in a single channel of the input feature; is the index on the spatial dimension, is the number of neurons on this channel, and are the weights and biases of the transformation, where all values ​​are scalars, and the improved U2Net network is obtained; 4.3) The enhanced aerial image and trajectory raster image from step 2 are input into the improved U2Net network to obtain the road prediction images corresponding to the aerial image and trajectory raster image respectively; In step 5, the trajectory grid image obtained in step 4 and the road prediction image corresponding to the high-altitude aerial image are fused using a pixel-level weighted combination method to obtain a merged road recognition prediction result.

2. The method for road recognition in aerial images based on deep learning according to claim 1, characterized in that: In step 5, the specific process is: The road recognition results of the trajectory raster image and the aerial image are fused using a pixel-level weighted fusion method. The intensity value of each pixel in the fused image is the weighted average of the intensity values ​​of the corresponding pixels in the trajectory data road extraction image and the remote sensing image road extraction image. Assume that the pixel intensity value matrix of the trajectory data road extraction image is , the pixel intensity value matrix of the remote sensing image road extraction image is , the pixel intensity value matrix of the fused image is , the functional formula of the fusion process is as follows: in, Indicates that the fused image is at position The pixel intensity value at Represents the trajectory data road extraction image at position The pixel intensity value at Indicates the remote sensing image road extraction image at location The pixel intensity value at and Represent the weighting coefficients of trajectory data and remote sensing images respectively, and satisfy ; From then on, the final road recognition prediction result is obtained.

Citation Information

Patent Citations

  • Adaptive gain image enhancement method based on fractional order multi-scale entropy fusion

    CN110889806A

  • Urban road extraction method based on multi-source data fusion, storage medium and equipment

    CN118094471A