A method for improving the quality of road network recognition based on superimposed cutting
The method of double segmentation and result combination addresses the issue of incomplete edge pixel information in road extraction from remote sensing images, improving detection accuracy by ensuring comprehensive contextual information is obtained.
Patent Information
- Application Number
- CN202210243293.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-11
AI Technical Summary
In the prior art, after the road remote sensing image is cut, edge pixels cannot obtain sufficient information, resulting in low road extraction quality, discontinuity and misalignment problems, and the cost of improving hardware performance is high.
Two different cutting methods are used to cut the remote sensing image, and the superimposed cutting results are used in the prediction stage. The information acquisition of edge pixels is improved through linear superimposition methods, including adding all-zero pixel bars around the original image during the prediction stage, renumbering and inputting it into the neural network for cutting, and finally superimposing the two cutting results to make up for the missing parts.
The quality of road extraction is improved, especially the extraction accuracy near the cutting line, reduces road discontinuity and dislocation, and reduces the cost requirement for hardware improvement.
Smart Images

Figure CN114743091B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image road recognition, and particularly relates to a method for improving the recognition quality of road networks based on superimposed cutting. Different cutting methods are used to generate different road extraction maps when extracting roads, and then they are superimposed to improve the road extraction quality at the cutting points. Background Art
[0002] Road remote sensing information technology is an essential part of the modern urban and social economic development in our country. Studying the application of remote sensing road images and the methods of road information extraction has important social science and technical significance. The design scope of automatic remote sensing road information extraction technology covers multiple fields and tasks such as urban planning, intelligent rail transit, update of geographic information systems, and detection of comprehensive utilization of land resources.
[0003] In recent years, the use of remote sensing images for large-scale automatic road extraction has received extensive attention. However, due to the large size and high resolution of road remote sensing images, and at the same time because of hardware limitations, it is difficult to perform road extraction on the complete remote sensing image. The currently common method is to first cut the image, then perform road extraction on the cut objects separately, and finally splice the extracted road network information. However, the current general focus is on how to improve the extraction accuracy of the algorithm, while ignoring the impact of cutting on road extraction. Performing road extraction on the cut image will cause the model to be unable to obtain the information around the edge pixels, thereby reducing the prediction quality of the edge pixels and resulting in situations such as discontinuous roads and misalignment. One way to solve this problem is to improve the hardware performance, but the cost is relatively high. A better way should be to implement an algorithm that can enable the model to fully obtain the information around it when predicting edge pixels. Summary of the Invention
[0004] The purpose of the present invention is to propose a method for improving the recognition quality of road networks based on superimposed cutting, aiming to solve the problem that in the road recognition algorithm, it is necessary to perform cutting prediction on the image, resulting in insufficient information being obtained by the pixels near the cutting during prediction, resulting in a low quality of the road network predicted at the cutting point and causing situations such as discontinuous roads and misalignment.
[0005] Based on the above purpose, the technical solution of the present invention is as follows: A method for improving the recognition quality of road networks based on superimposed cutting, characterized in that different cutting methods are used to cut the complete remote sensing image in the prediction stage, so that the pixels located at the edge of the cutting line in the first cutting method are located inside the small image in the second cutting method, thereby obtaining relatively complete information, and finally linearly superimposing the results of the two cutting predictions.
[0006] Specifically, in the prediction stage, first, the high-resolution remote sensing image is cut. Here, it is assumed that the size of the original image is M*N, and the image is cut into small images of k*k. To make the size of each small image equal, a strip of all-zero pixels with a width of is added at the bottom of the original image, and a strip of all-zero pixels with a width of is added on the right side of the original image; after cutting, a total of small images can be obtained; to finally assemble the small images into the original image, the small images are numbered. The pixel coordinates of the upper left corner of the original image are set to (0, 0). Therefore, the pixel with coordinates (m, n) is located in the m-th row and the n-th column of the original image. Use img i,j to represent the small image with coordinates (i, j). The coordinates of the upper left corner pixel in the original image are (i*k, j*k). Then, the cut image is used as the input of the neural network to obtain the semantic segmentation result of the first cut. Then, according to the coordinates of the small image and the cutting size, calculate the position of the small image in the original image to obtain the complete road extraction map. To restore it to the original size, the pixels with a width of at the bottom and the pixels with a width of on the right side need to be removed;
[0007] Before the second cut, similar to the first cut, a strip of all-zero pixels with a width of is added at the bottom of the original image, and a strip of all-zero pixels with a width of is added on the right side of the original image. At this time, the size of the image becomes However, in addition, strips of all-zero pixels need to be added around the original image. Strips of all-zero pixels with a width of are added on the left and upper sides of the image, and strips of all-zero pixels with a width of are added on the right and lower sides of the image. Therefore, before the second cut, the size of the image becomes After cutting, small images will be generated. At this time, the small images need to be re-numbered, and then they are respectively input into the semantic segmentation network to obtain the semantic segmentation result of the second cut. Then, the segmentation results of the small images are assembled to obtain the complete road extraction map. Similarly, to restore it to the original size, the pixels with a width of on the left and upper sides need to be removed, the pixels with a width of at the lower side need to be removed, and the pixels with a width of on the right side need to be removed;
[0008] Finally, the semantic segmentation results obtained by different cutting methods are superimposed to complement each other's missing parts. Suppose the complete semantic segmentation map generated by the first cut is I1, and the complete semantic segmentation map generated by the second cut is I2. The superimposed result of the two is I new , I x,y represents the pixel value at coordinates (x, y) in the semantic segmentation map.
[0009] Furthermore, the superimposing method of the two semantic segmentation maps adopts taking the average, taking the maximum value, and weighted average, and finally selects the one with the highest prediction quality as the final prediction result.
[0010] Among them, the superimposing method of the two semantic segmentation maps takes the average method, and takes the average value of the pixel values of the corresponding pixels in the semantic segmentation maps generated by the two cuts as the new pixel value, that is
[0011]
[0012] where is the new pixel value, is the pixel value of the first semantic segmentation map, is the pixel value of the second semantic segmentation map.
[0013] Among them, when the superimposing method of the two semantic segmentation maps does not consider the pixel surrounding environment, it takes the maximum value. This method takes the maximum value of the pixel values of the corresponding pixels in the semantic segmentation maps generated by the two cuts as the new pixel value, that is
[0014]
[0015] where is the new pixel value, is the pixel value of the first semantic segmentation map, is the pixel value of the second semantic segmentation map.
[0016] Among them, when the superimposing method of the two semantic segmentation maps considers the pixel surrounding environment, it takes the average. This method does not directly calculate according to equal weights, but uses the predicted value of the pixel as the weight. The weight of the pixel value in the p-th semantic segmentation map is calculated as:
[0017]
[0018]
[0019] where Q represents the number of semantic segmentation maps generated by cutting, which is 2 here. The constant 1 is to prevent the result of a certain pixel in multiple prediction maps from being 0, resulting in a denominator of 0.
[0020] By adopting the method of two - time cutting, for the first cutting method, the pixels located at the edge of the cutting line are inside the small image in the second cutting method. Similarly, for the second cutting method, the pixels located at the edge of the cutting line are inside the small image in the first cutting method. In this way, relatively complete information can be ensured, thereby improving the quality of road extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is the remote sensing image to be predicted.
[0022] Figure 2 is the road semantic segmentation map generated after only the first cutting.
[0023] Figure 3 is the road semantic segmentation map generated after only the second cutting.
[0024] Figure 4 is Figure 2 and Figure 3 the result of taking the average value.
[0025] Figure 5 is Figure 2 and Figure 3 the result of taking the maximum value.
[0026] Figure 6 is Figure 2 and Figure 3 the result of taking the weighted average. DETAILED IMPLEMENTATION MANNER
[0027] The following further illustrates the technical solution of the present invention according to specific embodiments.
[0028] Due to the limitation of hardware conditions, it is difficult to use a whole high - resolution remote sensing image as input to predict road information. Therefore, first, the high - resolution remote sensing image is cut. Here, it is assumed that the size of the original image is M * N, and the image is cut into small images of k * k. To make the size of each small image equal, we add a strip of all - zero pixels with a width of at the bottom of the original image, and add a strip of all - zero pixels with a width of at the right side of the original image. After segmentation, we can obtain a total of small images. To finally assemble the small images into the original image, we need to number the small images. We set the pixel coordinates of the upper - left corner of the original image as (0, 0). Therefore, the pixel with coordinates (m, n) is located in the m - th row and the n - th column of the original image. We use img i,j to represent the small image with coordinates (i, j). The coordinates of the top - left pixel in the original image are (i * k, j * k). Then, the cut - out image is used as the input to the neural network to obtain the semantic segmentation result of the first cut. Then, based on the coordinates of the small image and the cutting size, the position of the small image in the original image is calculated to obtain the complete road extraction map. To restore it to the original size, it is also necessary to remove the pixels with a width of at the bottom, and the pixels with a width of on the right.
[0029] Before the second cut, similar to the first cut, a strip of all - zero pixels with a width of is added to the bottom of the original image, and a strip of all - zero pixels with a width of is added to the right of the original image. At this time, the size of the image becomes However, in addition, strips of all - zero pixels need to be added around the original image. Strips of all - zero pixels with a width of are added to the left and top of the image, and strips of all - zero pixels with a width of are added to the right and bottom of the image. Therefore, before the second cut, the size of the image becomes After cutting, small images are generated. At this time, the small images need to be re - numbered, and then they are respectively input into the semantic segmentation network to obtain the semantic segmentation result of the second cut. Then, the segmentation results of the small images are assembled to obtain the complete road extraction map. Similarly, to restore it to the original size, it is necessary to remove the pixels with a width of on the left and above, remove the pixels with a width of at the bottom, and remove the pixels with a width of on the right.
[0030] Finally, the semantic segmentation results obtained by different cutting methods are superimposed to complement each other's missing parts. Suppose the complete semantic segmentation map generated by the first cut is I1, the complete semantic segmentation map generated by the second cut is I2, and the superimposed result of the two is I new ,I x,y represents the pixel value at the coordinate (x, y) in the semantic segmentation map. Specific superimposing methods can adopt different methods, such as taking the average, taking the maximum value, and weighted average, which will be introduced separately below.
[0031] Taking the average: This method takes the average of the pixel values of the corresponding pixels in the semantic segmentation maps generated by the two cuts as the new pixel value, that is
[0032]
[0033] Taking the maximum value without considering the pixel's surrounding environment: This method takes the maximum value of the pixel values of the corresponding pixels in the semantic segmentation maps generated by two cuts as the new pixel value, that is
[0034]
[0035] Taking the average considering the pixel's surrounding environment: This method does not calculate directly with equal weights, but uses the predicted value of the pixel as the weight. The calculation method of the weight w is:
[0036]
[0037]
[0038] Among them, Q represents the number of semantic segmentation maps generated by cutting, which is 2 here. The constant 1 is to prevent the result of a certain pixel in multiple prediction maps from being 0, resulting in a denominator of 0.
[0039] Select a binary threshold for binarization. Predict pixels with predicted values greater than the threshold as roads and set the pixel value to 1. Conversely, predict as background and set the pixel value to 0.
[0040] Now explain why using the method of two cuts can improve the quality of road extraction. Assume that the coordinates of a certain pixel P in the original image are (m, nk), where m, n ∈ N, 0 ≤ m, nk ≤ M. Therefore, in the first cut method, pixel P is located at the left boundary of When is input into the semantic segmentation network for prediction, the network cannot obtain the information on the left side of pixel P, so the prediction quality of this pixel will be reduced. In the second cut method, due to adding all-zero pixel strips on the left and above of the original image, the position of pixel P has changed. At this time, the coordinates of pixel P are Therefore, pixel P is located at img new represents the small image generated by the second cut. At this time, pixel P is inside the small image. Therefore, when making a prediction, the semantic segmentation network can obtain the information around pixel P and improve the prediction quality of this pixel.
[0041] Similarly, in the second cut method, pixels located near the cut line are inside the small image in the first cut method. Finally, linearly superimpose the two road extraction maps to improve the extraction quality near the cut line.
[0042] The following is an embodiment of the present invention. Figure 1 is the original remote sensing image, with a size of 7772 * 13719, which is a remote sensing image of a certain area in Shanghai.
[0043] First, fill the right and bottom sides of the complete remote sensing image with zero pixel bars. In the experiment, the original image was cut into small images of 256*256. Therefore, the filling width on the right side is The calculation result in the experiment is 105. The filling width at the bottom is The calculation result in the experiment is 164. The smaller the k, the lower the prediction accuracy may be when predicting the current small image because insufficient environmental information cannot be obtained. However, if k is too large, although more information will be obtained, noise will also be generated. In the experiment, k was taken as 128, 256, 512, and 1024 for comparison, and finally it was found that the best effect was obtained when k = 256. Then, the image was cut, and the cut images were input into the semantic segmentation network for prediction to generate the prediction results of the first cut. Then, the images were stitched together, and the effect is as Figure 2 shown. Then, based on the first filling, add zero pixel bars with a width of 128 around the image, and cut the image into small images of 256*256 again. This makes the cutting line in the first cut located at the center of the small images in the second cut. Similarly, the cutting line in the second cut is also located at the center of the small images in the first cut. Then, prediction is performed again to generate the prediction results of the second cut, as Figure 3 shown. Then, by using different superposition methods, the semantic segmentation maps generated by different cutting methods are superposed to make up for the poor road extraction quality near the cutting line. Three superposition methods are used respectively: taking the average, taking the maximum value, and taking the weighted average.
[0044] In this embodiment, the symbol refers to rounding up A.
[0045] Taking the average value is to take the average of the pixel values of the corresponding pixels in multiple segmentation maps as the pixel values in the new segmentation map. This method will first calculate the sum of the road pixels and the background pixels, and then calculate the average value. Suppose the predicted values at a certain position in the two cutting methods are 200 and 40 respectively, then the superposition result is 120. However, there will be such a situation: after the first cut, the predicted pixel value of a certain pixel is relatively high, that is, it is predicted as a road. After the second cut, the predicted pixel value of this pixel is relatively low, that is, it is predicted as the background. After averaging the two, the predicted pixel value of the road will become lower. Just like the above example, in the first cut result, the predicted result is 200, but after superposition, it is 120. After the final binarization step, it may be considered as the background, resulting in a reduction in accuracy. The effect is as Figure 4 shown. The road represented by the blue line in it was not predicted in the first cut (i.e., Figure 2 ), so using overlapping prediction can significantly improve the road prediction quality.
[0046] When taking the maximum value without considering the surrounding environment of the pixel, the pixel value of the corresponding pixel in multiple segmentation maps is taken as the maximum value as the pixel value in the new segmentation map. Suppose the predicted values at a certain position in two cutting methods are 200 and 40 respectively, then the superimposed result is 200. This method will not misjudge the pixels originally predicted as roads, but there may be more fragmented roads. As Figure 5 shown, the blue color indicates the roads that were not predicted in the first cutting result. Compared with Figure 4 , although more roads are predicted, the fragmented roads in the prediction are also significantly increased.
[0047] Taking the weighted average uses the current pixel value as the weight, and the result is between taking the average value and taking the maximum value. Suppose the predicted values at a certain position in two cutting methods are 200 and 40 respectively, then the superimposed result is 173. When the prediction results of the two cutting methods are close at a certain pixel position, the effects obtained by using the above three superimposing methods are similar. However, when the difference between the two prediction results is large, the result obtained by taking the weighted average will be greater than the result obtained by taking the average value with equal weights. Compared with Figure 4 , more roads are predicted. Compared with Figure 5 , there are fewer fragmented roads.
[0048] Comparing the three different superimposing methods, select the one with the highest prediction quality as the final prediction result.
Claims
1. A method for improving the quality of road network recognition based on superposition cutting, characterized in that During the prediction stage, the complete remote sensing image is cut using different cutting methods. In the first cutting method, the pixels on the edge of the cutting line are inside the small image in the second cutting method, so as to obtain relatively complete information. Finally, the results of the two cutting predictions are linearly superimposed; In the prediction stage, first, the high-resolution remote sensing image is cut. Here, it is assumed that the size of the original image is M*N, and the image is cut into small images of k*k. To make the size of each small image equal, a strip of all-zero pixels with a width of is added to the bottom of the original image, and a strip of all-zero pixels with a width of is added to the right side of the original image; after segmentation, a total of small images can be obtained; to finally assemble the small images into the original image, the small images are numbered. The pixel coordinates of the upper left corner of the original image are set to (0,0). Therefore, the pixel with coordinates (m,n) is located in the m-th row and the n-th column of the original image. Use img i,j to represent the small image with coordinates (i,j). The pixel coordinates of the upper left corner in the original image are (i*k,j*k). Then, the cut image is used as the input of the neural network to obtain the semantic segmentation result of the first cut. Then, according to the coordinates of the small image and the cutting size, the position of the small image in the original image is calculated to obtain the complete road extraction map. To restore it to its original size, the pixels with a width of at the bottom and the pixels with a width of on the right side need to be removed. Before the second cut, similar to the first cut, a strip of all-zero pixels with a width of was added to the bottom of the original image, and a strip of all-zero pixels with a width of was added to the right side of the original image. At this time, the size of the image became However, in addition to this, strips of all-zero pixels need to be added around the original image. Strips of all-zero pixels with a width of are added to the left and top of the image, and strips of all-zero pixels with a width of are added to the right and bottom of the image. Therefore, before the second cut, the size of the image becomes After cutting, small images will be generated. At this time, the small images need to be re-numbered, and then input into the semantic segmentation network separately to obtain the semantic segmentation results of the second cut. Then, the segmentation results of the small images are assembled to obtain the complete road extraction map. Similarly, in order to restore it to its original size, pixels with a width of need to be removed from the left and top, pixels with a width of need to be removed from the bottom, and pixels with a width of need to be removed from the right; Finally, the semantic segmentation results obtained by different cutting methods are superimposed to complement each other's missing parts. Suppose the complete semantic segmentation map generated by the first cut is I1, and the complete semantic segmentation map generated by the second cut is I2. The superimposed result of the two is I new , I x,y represents the pixel value at the coordinates (x, y) in the semantic segmentation map.
2. The method for improving the recognition quality of road networks based on superimposed cutting according to claim 1, wherein The superimposing methods of the two semantic segmentation maps include taking the average, taking the maximum value, and weighted average. Finally, the one with the highest prediction quality is selected as the final prediction result.
3. The method for improving the recognition quality of road networks based on superimposed cutting according to claim 2, wherein, The superimposing method of the two semantic segmentation maps is the average method. The average value of the pixel values of the corresponding pixels in the two semantic segmentation maps generated by the two cuts is taken as the new pixel value, that is wherein is the new pixel value, is the pixel value of the first semantic segmentation map, is the pixel value of the second semantic segmentation map.
4. The method for improving the recognition quality of road networks based on superimposed cutting according to claim 2, characterized in that, The superimposing method of the two semantic segmentation maps takes the maximum value without considering the surrounding environment of the pixels. In this method, the maximum value of the pixel values of the corresponding pixels in the two semantic segmentation maps generated by the two cuts is taken as the new pixel value, that is wherein is the new pixel value, is the pixel value of the first semantic segmentation map, is the pixel value of the second semantic segmentation map.
5. The method for improving the recognition quality of road networks based on superimposed cutting according to claim 2, wherein The superimposing method of the two semantic segmentation maps takes the average considering the surrounding environment of the pixels. This method does not calculate directly according to equal weights, but uses the predicted value of the pixels as the weights. The calculation method of the weight w is: Among them, Q represents the number of semantic segmentation maps generated by cutting, which is 2 here.
Citation Information
Patent Citations
Semantic segmentation method, device and system
CN107977624A