Road extraction method combining OSM data and optical image and related equipment
By combining OSM data and optical images, using image segmentation and superpixel segmentation models, the conversion from sparse labels to fully supervised labels is realized, solving the problem of artificial label data dependence in deep learning methods, and realizing the automation and intelligence of road extraction.
Patent Information
- Application Number
- CN202510090580.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
AI Technical Summary
Existing deep learning methods require a large amount of manually annotated label data in road interpretation, resulting in insufficient automation and intelligence, especially in the training data.
Combining OSM data and optical images, the semantic and structural-level features of optical images are obtained using image segmentation models and superpixel segmentation models, and the conversion from sparse labels to fully supervised labels is achieved through the label propagation of multi-level features and the prediction capabilities of network models.
The full process automation and intelligence of road extraction is realized, the coverage and accuracy of tag data are improved, the problems of inaccurate spread of sparse tags and incomplete coverage are solved, and the accuracy and stability of road extraction are improved.
Smart Images

Figure CN119942122A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road detection, and in particular to a road extraction method combining OSM data and optical images and related equipment. Background Art
[0002] As a data-driven method, the application effect of deep learning is often closely related to the quality of label data. As an intensive prediction task, the performance of the network model of road interpretation is even more strongly dependent on the fine annotation of training labels at the pixel-level granularity, so as to provide accurate trend preferences for parameter optimization in the back-propagation process. In traditional network model training, the basic rules of its loss constraints are formulated around the quantification of prediction results and label data, so the category attribution of label data needs to traverse all pixels. However, the generalization ability of traditional network models in different regions is usually limited, resulting in the need for corresponding label data for intelligent interpretation of roads in different regions, while the current fully supervised label data is mainly manually annotated. Therefore, in the actual application scenarios of road interpretation, deep learning methods have not achieved automation and intelligence of the entire process, especially in terms of training data, which involves a lot of manual work. Summary of the invention
[0003] In view of the problem that existing deep learning methods require a large amount of manually annotated label data as training data in the process of road interpretation, the present invention proposes a road extraction method and related equipment that combines OSM data and optical images. OSM data is used as the initial label of the road, which is used as a seed to propagate pixel-level labels, and then road extraction is achieved with the help of a network model to make up for the defects of automation and intelligence in the process.
[0004] In a first aspect, the present invention provides a road extraction method combining OSM data and optical images, comprising:
[0005] Acquire OSM data corresponding to the target optical image, and preprocess the OSM data;
[0006] Generate a first road buffer and a second road buffer according to the preprocessed OSM data; wherein the pixel attributes in the first road buffer are all roads, and the second road buffer should cover all pixels with road attributes;
[0007] The target optical image is processed by an image segmentation model to obtain a first segmentation result; the target optical image is processed by a superpixel segmentation model to obtain a second segmentation result;
[0008] Performing rough optimization of road labels according to the first road buffer, the first segmentation result, and the second segmentation result to generate a rough optimization result;
[0009] Inputting the target optical image into a trained network model to perform road prediction and obtain a road prediction result;
[0010] The road prediction result, the rough optimization result, the first road buffer zone and the second road buffer zone are combined to perform road label fine optimization to generate a fine optimization result.
[0011] Furthermore, the OSM data is preprocessed, specifically including:
[0012] A grid matrix is constructed according to the spatial resolution of the target optical image, and the attributes of all pixels of the grid matrix are initialized as background; the geographic coordinates of the road in the OSM data are converted into pixel coordinates and the attributes of the pixels at the corresponding positions in the grid matrix are modified to roads to generate a road mask image; the road mask image is aligned with the target optical image.
[0013] Further, performing rough optimization of road labels according to the first road buffer, the first segmentation result and the second segmentation result to generate a rough optimization result specifically includes:
[0014] Traverse each object in the first segmentation result, and calculate the ratio T of the number of pixels in the intersection area between the object and the first road buffer to the number of pixels of the object area1 , if the ratio T area1 If the value is greater than a preset threshold, the object is considered to be a road, and all road objects are merged to generate a first rough optimization intermediate result;
[0015] Traverse each superpixel in the second segmentation result, and calculate the ratio T of the number of pixels in the intersection area between the superpixel and the first road buffer to the number of pixels in the superpixel area2 , if the ratio T area2 If the value is greater than a preset threshold, the superpixel is considered to be a road, and all road superpixels are merged to generate a second rough optimization intermediate result;
[0016] The first rough optimization intermediate result and the second rough optimization intermediate result are combined to obtain the rough optimization result.
[0017] Furthermore, the training process of the network model includes: using the rough optimization result of the optical image as a label, training a preset network in combination with the optical image, and obtaining a trained network model.
[0018] Furthermore, combining the road prediction result, the rough optimization result, the first road buffer zone, and the second road buffer zone to perform road label fine optimization to generate a fine optimization result specifically includes:
[0019] The rough optimization result is subjected to an intersection operation with the first road buffer zone, the road prediction result is subjected to an intersection operation with the second road buffer zone, and the two intersection operation results are subjected to a union operation to obtain the refined optimization result.
[0020] Further, generating a first road buffer and a second road buffer according to the preprocessed OSM data specifically includes:
[0021] The preprocessed OSM data is expanded using two convolution templates with different convolution kernel sizes to obtain a first road buffer and a second road buffer; the convolution kernel size of the convolution template corresponding to the first road buffer is smaller than the convolution kernel size of the convolution template corresponding to the second road buffer.
[0022] Furthermore, the image segmentation model adopts a SAM model.
[0023] In a second aspect, the present invention provides a road extraction device combining OSM data and optical images, comprising:
[0024] A preprocessing module, used for acquiring OSM data corresponding to the target optical image and preprocessing the OSM data;
[0025] A buffer generation module is used to generate a first road buffer and a second road buffer according to the preprocessed OSM data; wherein the pixel attributes in the first road buffer are all roads, and the second road buffer should cover all pixels with road attributes;
[0026] A segmentation module, used to process the target optical image using an image segmentation model to obtain a first segmentation result; and to process the target optical image using a superpixel segmentation model to obtain a second segmentation result;
[0027] A rough optimization module, used for performing rough optimization of road labels according to the first road buffer, the first segmentation result and the second segmentation result, and generating a rough optimization result;
[0028] A prediction module, used for inputting the target optical image into a trained network model to perform road prediction and obtain a road prediction result;
[0029] A fine optimization module is used to combine the road prediction result, the rough optimization result, the first road buffer zone and the second road buffer zone to perform fine optimization of road labels and generate a fine optimization result.
[0030] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0031] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.
[0032] The beneficial effects of the present invention are:
[0033] (1) The present invention uses OSM data as the initial label of the road, and uses it as a seed to propagate pixel-level labels, and then uses the network model to realize road extraction, forming a sparse label growth and optimization method that is widely applicable to a variety of network models, making up for the automation and intelligent defects in the process, and realizing the automatic extraction of roads in the whole process.
[0034] (2) To address the problem of inaccurate seed information propagation of sparse labels, the present invention combines an image segmentation model with a superpixel segmentation model to capture the semantic and structural features of optical images, realizes label propagation of multi-level features, and constructs an object-level information propagation method based on this, which makes up for the shortcoming that pixel-level processing units are not suitable for the rich spatial details of optical images.
[0035] (3) To address the problem of incomplete coverage of sparse label training data, a coarse-to-fine label optimization strategy is constructed. “Coarse optimization” is performed around the label propagation of multi-level features, and “fine optimization” is achieved with the help of the network model’s own prediction ability to accurately expand the coverage of sparse labels.
[0036] (4) Experimental analysis of the method of the present invention was carried out using both simulated data and real data, including accuracy comparison and effect analysis with typical and advanced methods, as well as ablation experiments on important modules, to verify the accuracy and stability of the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 One of the flowcharts of a road extraction method combining OSM data and optical images provided by an embodiment of the present invention;
[0038] Figure 2 A second flowchart of a road extraction method combining OSM data and optical images provided by an embodiment of the present invention;
[0039] Figure 3 The process of preprocessing OSM data provided by the embodiment of the present invention;
[0040] Figure 4 A rough label optimization process provided by an embodiment of the present invention;
[0041] Figure 5Example images and their OSM overlay and segmentation effects provided for the embodiments of the present invention: (a) OSM overlay—example image 1; (b) example image 2; (c) SAM segmentation—example image 2; (d) SLIC segmentation—example image 2;
[0042] Figure 6 An example of label rough optimization provided by an embodiment of the present invention: (a) optical image; (b) rough optimization process result 1; (c) rough optimization process result 2; (d) rough optimization result;
[0043] Figure 7 The label optimization process provided by the embodiment of the present invention;
[0044] Figure 8 Examples of road extraction results on different images using different methods provided in the embodiments of the present invention: (a) Image 1; (b) Image 2; (c) Image 3; (d) Image 4; (e) Image 5;
[0045] Fig. 9 A schematic diagram of the structure of a road extraction device combining OSM data and optical images provided by an embodiment of the present invention;
[0046] Fig.10 The present invention is a structural block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] Thanks to the development of the concept of "ubiquitous mapping", the popularity and richness of OSM (OpenStreetMap) data have been significantly improved, which has promoted the comprehensive development of the ability to perceive surface information to a certain extent. In terms of spatial relationships, OSM data and remote sensing images can realize the association mapping of vector data and raster data through geographic coordinates; in terms of attribute information, OSM data can provide initial object category labels for remote sensing images. However, OSM data is expressed in the form of strips on the road. Although it can describe the topological structure of the road, it does not completely cover the road area, that is, OSM data is equivalent to the sparse label (center line) of the road. Therefore, OSM data can theoretically be used as the initial label of the road. If it is used as a seed and the pixel-level label is propagated, the road extraction can be realized with the help of the network model, making up for the defects of automation and intelligence in the process.
[0049] In the field of road extraction based on sparse samples, the road width is mainly preset with empirical values, and fully supervised road samples are obtained in this way, but this idea cannot match the scale diversity brought by differences in road grades. Based on this, researchers effectively diffused sparse samples by calculating the difference in preset road widths and information propagation based on clustering algorithms, and finally divided the training data into roads, backgrounds, and pending areas, and supervised the pending areas with constraints such as regularization losses. However, problems such as the position offset between OSM data and optical images and the coverage of objects in optical images will bring challenges to information dissemination, resulting in inaccurate label diffusion, thus affecting the accuracy of road extraction.
[0050] In order to solve the above problems, Figure 1 As shown, an embodiment of the present invention provides a road extraction method combining OSM data and optical images, comprising the following steps:
[0051] S101: Acquire OSM data corresponding to a target optical image, and pre-process the OSM data;
[0052] Specifically, the target optical image is the optical image that needs to be extracted. OSM (Open Street Map) data is open source data, and users can download it from a designated website. According to the coverage area of the target optical image, the scope of OSM data is clarified. The purpose of preprocessing OSM data in this step is to make OSM data meet the input requirements of subsequent neural networks.
[0053] S102: Generate a first road buffer and a second road buffer according to the preprocessed OSM data; wherein the pixel attributes in the first road buffer are all roads, and the second road buffer should cover all pixels with road attributes;
[0054] Specifically, for the convenience of description, the image where the first road buffer is located is recorded as the first image, and the image where the second road buffer is located is recorded as the second image. The pixel attributes of the first image and the second image only include the background and the road. For the first image, the first road buffer can be understood as a refinement of the road target itself, and its constituent pixels must belong to the road as much as possible; for the second image, the second road buffer can be understood as an expansion of the road target itself, and its constituent pixels must cover all roads as much as possible.
[0055] S103: using an image segmentation model to process the target optical image to obtain a first segmentation result; using a superpixel segmentation model to process the target optical image to obtain a second segmentation result;
[0056] Specifically, the image segmentation model can be an image segmentation algorithm based on traditional computer vision, such as region growing method, edge detection method, threshold segmentation method and watershed algorithm, etc.; it can also be an image segmentation algorithm based on deep learning, such as U-Net, D-LinkNet, DeepLab series, YOLOv8-seg and SAM model, etc. Similarly, the superpixel segmentation model can be a superpixel segmentation model based on traditional methods, such as SLIC (Simple Linear Iterative Clustering) algorithm, Felzenszwalb-Huttenlocher algorithm, Quick Shift algorithm, etc.; it can also be a superpixel segmentation model based on deep learning, such as SAM-Road, DeepLab series, etc.
[0057] S104: performing rough optimization of road labels according to the first road buffer, the first segmentation result, and the second segmentation result to generate a rough optimization result;
[0058] Specifically, the image segmentation model mainly performs segmentation based on high-level features of pixels (such as semantic information) to identify different objects and regions in the image, that is, the segmentation accuracy of the first segmentation result is at the object level. The superpixel segmentation model mainly performs clustering segmentation based on low-level features of pixels (such as color, brightness, texture, etc.) to generate several superpixels. Each superpixel is generally a small area composed of a series of pixels with adjacent positions and similar features such as color, brightness, texture, etc., that is, the superpixel segmentation model pays more attention to the similarity and spatial proximity between pixels, and usually does not involve the semantic information of pixels, that is, the segmentation accuracy of the second segmentation result is at the superpixel level. For example, for the same target object in the image, the image segmentation model aims to segment the target object as a whole from other different targets, while the superpixel segmentation model will also segment the same target object into multiple superpixels. The present invention utilizes this complementary relationship between the two segmentation models, and then combines the position prior information given by the first road buffer to achieve rough optimization of optical image road labels.
[0059] S105: Inputting the target optical image into the trained network model to perform road prediction and obtain a road prediction result;
[0060] Specifically, this embodiment does not limit the type of network model for road prediction, which can be a traditional machine learning network, such as a support vector machine, random forest, etc.; it can also be a traditional neural network, such as a BP neural network; it can also be a deep learning network, such as a convolutional neural network, a recurrent neural network, and a graph neural network. Taking the training process of the network model of a deep learning network as an example, it specifically includes: constructing an optical image training data set, designing a loss function, and optimizing the parameters of the network model through the loss function. It can be understood that the training set data format, label data, and training learning process of different network models are different. When training a preset network, the optical image training data set and label data should be processed into a format that meets the requirements of the preset network as input, and training should be performed according to the existing training process.
[0061] S106: Combine the road prediction result, the rough optimization result, the first road buffer zone, and the second road buffer zone to perform road label fine optimization to generate a fine optimization result.
[0062] Specifically, the rough optimization result can be regarded as a sparse road label. This step aims to use the prediction ability of the neural network itself, that is, to achieve "fine optimization" with the help of the road prediction results, and accurately expand the coverage of sparse labels. In addition, considering that the objects obtained in the rough optimization process cannot accurately match the boundaries of the objects due to the limited segmentation accuracy, and the first road buffer cannot fully estimate the large-scale roads, this step also uses the second road buffer to achieve "fine optimization" to supplement the large-scale road labels.
[0063] The road extraction method provided by the embodiment of the present invention is based on OSM data and optical images, and uses an image segmentation model and a superpixel segmentation model to obtain segmentation objects of the optical image, thereby combining the spatial position relationship between the OSM data and the object to convert line annotations into fully supervised label data to complete coarse optimization; on the basis of the coarse optimization result, the road prediction result of the network model and the second road buffer are used to remove noise in the coarse optimization result, supplement the road labels with incomplete coverage, and realize fine optimization of road labels, thereby ultimately achieving the purpose of road extraction.
[0064] In one embodiment, Figure 2As shown in the figure, taking the SAM (Segment Anything Model) model as an image segmentation model and the SLIC (Simple Linear Iterative Clustering) model as a superpixel segmentation model as an example, a road extraction method combining OSM data and optical images is provided. The method mainly includes two stages: the coarse optimization stage of labels and the fine optimization stage of labels. Among them, the coarse optimization stage covers two steps: the first step is to segment the optical image with the help of SAM, so as to transform the pixel-level processing unit of the image into the object level, and then based on the segmentation result and the narrow buffer (i.e., the first road buffer), the result of this step is obtained by combining the intersection operation and the area threshold (i.e., the first step result); the second step is to segment the optical image with SLIC, form superpixel units with the help of constraint information such as gradient and distance, and then based on the superpixel unit segmentation result and the narrow buffer, the coarse optimization result (i.e., the second step result) is obtained by combining the intersection operation, the area threshold and the first step result. In the fine optimization stage, the rough optimization results will be used as label data, and the network model will be trained in conjunction with optical images to obtain the road prediction results of the optical image training set. Finally, the final label fine optimization results (i.e., the third step results) are obtained through the intersection calculation of the prediction results and the wide buffer zone (i.e., the second road buffer zone) and the intersection calculation of the rough optimization results and the narrow buffer zone.
[0065] In the embodiment of the present invention, SAM, as a classic example of a large model in the field of images, can obtain better segmentation results under zero-sample conditions and achieve effective aggregation of pixels that are spatially adjacent and have similar features. Since the model does not require the intervention of training data and belongs to the open source computing power in the field of image segmentation, the processing results of this technology can often be used as initial information or prior conditions in back-end tasks such as ground object interpretation and change detection. The processing results of SLIC are often composed of a series of pixels that are spatially adjacent and have similar shallow features such as spectrum and texture. Compared with SAM, this technology also belongs to the category of image segmentation, but the obtained units are usually more refined than the segmentation results of SAM, so in the information dissemination of label data, the central area of the road can be better focused, thereby taking into account the accuracy.
[0066] In one embodiment, considering that optical images and OSM data are raster and vector data respectively, they cannot be directly input into the network model for training and testing. Therefore, this embodiment will first convert the data type of OSM data and complete the location information matching based on this. The process of rasterizing OSM data is to map the location information of the road to the two-dimensional plane of the raster image. The specific process is as follows: Figure 3 As shown:
[0067] S301: Initialize the grid matrix
[0068] Specifically, a grid matrix is constructed according to the spatial resolution of the target optical image, and the attributes of all pixels of the grid matrix are initialized as background;
[0069] S302: Raster Conversion
[0070] Specifically, the geographic coordinates of the road in the OSM data are converted to pixel coordinates, and the pixel attribute value at the coordinate position is converted to the road, and so on, to achieve the type conversion of OSM data and generate a road mask image. The conversion formula between geographic coordinates and pixel coordinates is shown in equations (1) and (2):
[0071] x=round((lon p -lon min )×num lon )(1)
[0072] y=num y -round((lat p -lat min )×num lat )(2)
[0073] In the formula, round is the rounding function, lon is min and lat min Respectively represent the minimum longitude and minimum latitude of the coverage area, num lon and num lat The number of pixels in each degree of longitude and latitude is related to the set spatial resolution. y is the number of pixels in the y direction of the coverage, lon p and lat p is the longitude and latitude coordinates of the point to be converted, and x, y represents the pixel coordinates after conversion.
[0074] S303: Position Alignment
[0075] Specifically, considering that the spatial resolution of open-source remote sensing images is not constant in both horizontal and vertical directions, the rasterized OSM data usually cannot achieve a one-to-one correspondence between pixels and optical images. At the same time, considering that the rasterized OSM data is road label data, its attribute categories only include road and background, and the error caused by grayscale interpolation is small, so the OSM data is directly interpolated and binarized with reference to the length and width of the optical image, and finally the label data that is completely consistent with the length and width of the optical image is obtained.
[0076] In one embodiment, a convolution module is used to perform a dilation operation on a binary image based on OSM data to generate a road buffer. Specifically, two convolution templates with different convolution kernel sizes are used to perform a dilation operation on the preprocessed OSM data to obtain a first road buffer and a second road buffer; the convolution kernel size of the convolution template corresponding to the first road buffer is smaller than the convolution kernel size of the convolution template corresponding to the second road buffer. The above dilation operation process can be expressed as follows:
[0077] L buffer ={x|(K) x ∩A≠Θ}(3)
[0078] In the formula, K, A and L buffer Respectively represent the convolution template, the input OSM data and the expanded road buffer data, and x is the specific processing pixel. In order not to introduce too much non-road information, a smaller convolution module is selected to obtain a narrow buffer, that is, when the OSM data is superimposed without error, the pixels in the buffer will belong to the road as much as possible. A larger convolution module is selected to obtain a wide buffer, that is, the buffer should contain all roads as much as possible.
[0079] In one embodiment, based on the above embodiments, an embodiment of the present invention provides a method for coarse optimization of labels by combining the SAM model and the SLIC model. From the annotation form, OSM data is equivalent to the line annotation of the road, which provides the category and partial location information of the road. However, if the data is directly used as a road label, a large number of positive samples will be mistakenly marked as negative samples and input into the network model, thereby leading to incorrect optimization of model parameters; if a certain width is set to establish a buffer zone based directly on the line annotation, the scale difference of the road in the image cannot be solved, and a large number of positive and negative sample confusion phenomena will also be caused. In summary, the embodiment of the present invention uses SAM and SLIC to perform multi-level feature extraction on the optical image, and clusters the pixels of the optical image into objects based on feature similarity and spatial distance, thereby constructing the initial boundary between the road and the background by cutting between objects, and defining the effective range for the information propagation of the label. The detailed process is as follows. Figure 4 As shown:
[0080] S401: Information propagation based on SAM and narrow buffer
[0081] Specifically, from the perspective of ideal data, if the segmented object is covered by OSM data, then the object must be a road, otherwise it is a background. However, real data often has two problems. One is the matching accuracy problem, such as Figure 5As shown in (a), although most OSM data will be superimposed on the road area of the optical image, due to factors such as acquisition accuracy and timeliness, the superimposed area of some OSM data is not accurate. If the labels are optimized directly through spatial superposition, other objects will be mistakenly marked as road labels. The second is the problem of object occlusion and overlap, such as Figure 5 (b) and Figure 5 As shown in (c), the road surface is blocked by vegetation and other objects. In SAM image segmentation, the blocked area and the non-blocked area of the road usually belong to two objects, and the object in the blocked area will extend outside the road according to the actual distribution of the objects. Therefore, similar to problem 1, this phenomenon will also cause negative samples to be mistakenly marked as positive samples.
[0082] Therefore, in order to solve the above two problems, in an exemplary embodiment, the SAM segmentation results are used to calculate the relevant area ratio for each object to eliminate the influence of overlapping objects and mismatching. The specific process is as follows: traverse each object in the first segmentation result, calculate the ratio T of the number of pixels in the intersection area of the object and the first road buffer to the number of pixels of the object area1 , if the ratio T area1 If the value is greater than the preset threshold, the object is considered to be a road, and all road objects are merged to generate the first rough optimization intermediate result. area1 The calculation formula is as follows:
[0083]
[0084] In the formula, O SAM(i) represents the number of pixels of object i in the first segmentation result, m is the number of objects in the first segmentation result, ∩ is the intersection operation, and L buffer represents the number of pixels in the first road buffer, T area1 Represents the ratio of the number of pixels in the intersection area to the number of pixels in the SAM object.
[0085] If the object is a non-road target, its boundary outline and coverage are usually quite different from those of the narrow buffer; if the object category is a road, considering the strip outline of the road, the narrow buffer can usually be regarded as a "thinning" result of the road, so its coverage will not be very different. Figure 5 The object in the red box in (c) is classified as vegetation, but part of the area blocks the road, causing the narrow buffer of the OSM data to overlap with the object. However, the ratio of the overlapping area to the object area is obviously smaller than the area ratio of the regular road. Therefore, this parameter can effectively alleviate the impact of matching errors and ground object occlusion on the propagation of label information.
[0086] S402: Information propagation based on SLIC and narrow buffer
[0087] Considering that the segmentation accuracy of SAM itself is limited, and the above process completely removes the covered area of the road, the integrity of its optimization result has direct defects. In order to further improve the integrity of the label data, the SLIC segmentation result and the narrow buffer are combined for optimization, that is, the SLIC segmentation result is calculated for each object. The relevant area ratio is supplemented with the label data of the road covered area. The specific process is as follows: traverse each superpixel in the second segmentation result, and calculate the ratio of the number of pixels in the intersection area of the superpixel and the first road buffer to the number of pixels of the superpixel. area2 , if the ratio T area2 If the value is greater than the preset threshold, the superpixel is considered to be a road, and all road superpixels are merged to generate the second rough optimization intermediate result. area2 The calculation formula is as follows:
[0088]
[0089] In the formula, O SLIC(i) represents the number of pixels of superpixel i in the segmentation result of SLIC, n is the number of superpixels in the segmentation result, T area1 The ratio of the number of pixels in the intersection area to the number of pixels in the SLIC superpixel. Figure 5 As shown in (d), the granularity of the SLIC segmentation result is significantly finer than that of the SAM segmentation result, that is, SLIC not only divides the target by category, but also divides the same target according to the spatial relationship and feature similarity of the pixels. Therefore, the superposition of the narrow buffer of the OSM data and the segmentation result will not spread to the ground objects outside the road, but can effectively spread to other ground object coverage areas.
[0090] S403: Coarse optimization of label data by combining SAM and SLIC;
[0091] SAM segmentation can effectively mine the semantic features of images, and SLIC segmentation can make full use of the structural features of images. Therefore, the constructed object-level information propagation strategy makes full use of the multi-level features of images. At the same time, label propagation based on SAM image segmentation effectively solves the matching error and object overlap problems, and information propagation based on SLIC image segmentation improves the integrity of label data on the basis of the former. The two are complementary to each other. In summary, this embodiment obtains the rough optimization result of label data according to formula (6):
[0092] L Sum1 =L SAM ∪L SLIC (6)
[0093] Where, L SAM and L SLIC represent the processing results of step S401 and step 402 respectively, ∪ is the union operation, LSum1 is the rough optimization result. After the above processing, the coverage of label data will be effectively improved, such as Figure 6 As shown in the figure, "rough optimization process result 1" is obtained by the intersection operation of SAM segmentation result and narrow buffer zone, which will inevitably introduce other interfering objects; "rough optimization process result 2" is screened by area ratio on the basis of "rough optimization process result 1", which effectively eliminates most of the interfering objects; "rough optimization result" combines SLIC segmentation and narrow buffer zone on the basis of "rough optimization process result 2" to effectively supplement the label data.
[0094] In one embodiment, an embodiment of the present invention provides a method for fine optimization of labels. Still taking SAM and SLIC as an example, the key technology of coarse optimization is to obtain the segmented objects of optical images with the help of SAM and SLIC, thereby combining the spatial position relationship between OSM data and objects to convert line annotations into fully supervised label data. However, there are two deficiencies in the coarse optimization process. First, the training data of SAM is natural images, which are different from the research data of the present invention (mainly optical images) in type and scene, and the accuracy of cross-domain processing of the model is limited; second, the spatial details of optical images are increasingly rich, and the structural features that SLIC relies on are difficult to cut in this type of data, so the acquired objects cannot accurately match the boundaries of objects, and the narrow buffer zone cannot fully take into account large-scale roads. In summary, this embodiment combines optical images and coarse optimization results, and uses the training and testing of network models to self-correct label data to achieve fine optimization. In this embodiment, taking a neural network as a network model as an example, Figure 7 As shown in the figure, the specific process of label optimization is as follows:
[0095] S701: Using the rough optimization result of the optical image as a label, input the optical image and the rough optimization result into the preset neural network for training, calculate the loss value of the road prediction result and the rough optimization result, and then continuously update the parameters of the neural network through the loss constraint until the number of training times or the loss decreases to meet the set conditions. The loss value calculation process can be expressed by the following formula:
[0096] Loss=F(M(I),L Sum1 )(7)
[0097] Where F is the loss function, M is the neural network, I is the optical image, and L is Sum1 is the rough optimization label, Loss is the loss value. The loss function can be cross entropy loss, focal loss, Dice loss, etc.
[0098] S702: Inputting the optical image into the trained neural network for road prediction;
[0099] S703: combining the road prediction result, the rough optimization result, the narrow buffer zone and the wide buffer zone to obtain the label fine optimization result;
[0100] Specifically, the rough optimization result is intersected with the first road buffer, the road prediction result is intersected with the second road buffer, and the two intersection results are unioned to obtain the refined optimization result. The above process can be expressed as follows:
[0101] L Sum2 =(L pre ∩L buffer2 )∪(L Sum1 ∩L buffer1 )(8)
[0102] Where, L Sum1 and L Sum2 They represent the rough optimization and fine optimization results of the labels, L buffer1 and L buffer2 They correspond to narrow buffer and wide buffer respectively, which are obtained by dilating convolution modules of different sizes. pre Corresponding to the road prediction result of step S402.
[0103] In the embodiment of the present invention, by using the sparse label data of the rough optimization result as the label data of the training data set, the network model in the fine optimization process is trained around its own data, so the mining of road information will be more comprehensive; at the same time, in order to avoid the noise interference of the sparse label data itself as much as possible, this embodiment also uses a wide buffer for elimination. In addition, considering that the rough optimization process always transmits information based on a narrow buffer, and this part is usually close to the center area of the road and is more likely to belong to the road, it is also screened with the help of a narrow buffer.
[0104] In order to verify the effectiveness of the solution of the present invention, the present invention also provides the following comparative experiments.
[0105] (I) Experimental data
[0106] (1) RoadNet dataset: The remote sensing images of this dataset come from Google Images. The image range is selected from the Ottawa area of Canada. The image spatial resolution is 0.21 meters. There is a lot of occlusion on the road surface. This experiment uses manual annotation to obtain the label data of the road surface, centerline and edge. In the heterogeneous data road extraction experiment, it is divided into 256 pixels × 256 pixels, of which 2627 images are used as training data and 375 images are used as test data. The marked road centerline is simulated as OSM data.
[0107] (II) Experimental details
[0108] The core hardware configuration of the experimental environment is 2 NVIDIA Tesla V100 graphics cards with a total of 64G video memory. Adam is selected as the optimizer for network training, and the initial learning rate is set to 2e-4. Whenever the loss value is higher than the current optimal loss value for 3 consecutive times, the learning rate is reduced by 5 times. The training data block size is 16, the iteration epoch value is 100, and the loss function uses the sum of BCE loss and dicecoefficient loss as the basis for supervised training. At the same time, in order to enhance the samples, the training data is randomly (50%) flipped vertically, horizontally, diagonally, and transformed radially. In addition, regarding the convolution module size of the OSM data buffer, this experiment sets the narrow and wide buffers to 5*5 and 40*40 respectively. The principle of setting is: the pixels in the narrow buffer should belong to the road as much as possible, and the pixels in the wide buffer should cover all roads as much as possible.
[0109] (III) Comparative experiments of different methods
[0110] This part of the experiment mainly uses two data sets to compare the method of the present invention (in this experiment, the SAM model is used as the image segmentation model by default, and the SLIC model is used as the superpixel segmentation model) with the existing typical methods. The comparison methods include ScRoadExtract and WeaklyOSM, and the network models used for road prediction include UNet, D-LinkNet, MANet and UNetFormer. The road extraction of the selected data sets is compared and analyzed from the two perspectives of quantitative accuracy and qualitative results.
[0111] Table 1 shows the accuracy statistics of the proposed method and the comparative method on the RoadNet dataset, where the fully supervised label data is obtained by manual sampling. The following conclusions can be drawn through accuracy statistics:
[0112] (1) The comprehensive evaluation indicators (F1 and IoU) of the comparison method and the proposed method are both lower than those of the fully supervised method, which proves that the comprehensiveness and accuracy of samples play an important role in data-driven network models;
[0113] (2) Compared with the fully supervised method, the proposed method has certain advantages in the P index and Eor index in D-LinkNet, MANet and UNetFormer, proving that when the network performance is relatively stable, the road misextraction of the proposed method is even better than that of the fully supervised method, but the completeness of road extraction needs to be further improved;
[0114] (3) Compared with ScRoadExtract and WeaklyOSM, the R index and Com of the proposed method are at a disadvantage, but the P index and Eor index are obviously superior, and the comprehensive evaluation index is dominant, which proves that although the two comparison methods ScRoadExtract and WeaklyOSM can extract more complete road information, they also introduce more non-road information, and the overall effect is worse.
[0115] Table 1 Comparison of road extraction results accuracy of different methods (RoadNet dataset)
[0116]
[0117] In addition to the comparative analysis of the above quantitative accuracy indicators, in order to more intuitively compare the road extraction effects of each method, the road extraction results of some typical test images are selected for comparative analysis. The selected 5 images are from different scenes, covering some difficulties in road extraction. In addition, this section no longer covers all network models, but selects MANet and UNetFormer, which have relatively outstanding comprehensive evaluation indicators, for analysis. The specific results are shown in Figure 8 , the detailed analysis is as follows:
[0118] (1) Image 1: The road in this image is a low-grade road, and part of the middle section is blocked by vegetation. Both comparison methods have the problem of mis-extraction in the MANet and UNetFormer network models. Although the topological structure of the road is not destroyed, the road capacity shown in the extraction results is obviously inconsistent with the actual situation. The overall performance of the method of the present invention is better, and the mis-extraction problem is better solved.
[0119] (2) Image 2: This image is located between residential areas and includes two parallel roads with a large road width. Similar to Image 1, the two comparison methods have serious mis-extraction problems. In particular, in some cases, the isolation belt between the two roads is misidentified as a road, causing the number of roads to change. The method of the present invention maintains a good extraction effect.
[0120] (3) Image 3: This image contains roads of different scales, among which some sections of large-scale roads are covered by vegetation and their shadows. The two comparison methods miss the extraction of small-scale roads in MANet and there is a mis-extraction problem in UNetFormer. The extraction effect of the method of the present invention is better.
[0121] (4) Image 4: The road in this image is located in a parking lot, and there is a railway covered by vegetation and shadows above the road. The two comparison methods, MANet and WeaklyOSM in UNetFormer, mistakenly extracted part of the railway information, while the method of the present invention accurately excluded similar objects.
[0122] (5) Image 5: The road in this image is located in the middle of vegetation and has a tortuous shape. Part of the road is covered by power lines, electric poles and their shadows. The extraction results of all methods effectively restore the topological structure of the road, but the extraction results of the comparison method obviously exceed the actual range of the road in terms of coverage area.
[0123] Based on the same inventive concept, Fig. 9 As shown, an embodiment of the present invention further provides a road extraction device combining OSM data and optical images, including a preprocessing module, a buffer generation module, a segmentation module, a rough optimization module, a prediction module and a fine optimization module.
[0124] Among them, the preprocessing module is used to obtain OSM data corresponding to the target optical image and preprocess the OSM data; the buffer generation module is used to generate a first road buffer and a second road buffer according to the preprocessed OSM data; wherein the pixel attributes in the first road buffer are all roads, and the second road buffer should cover all pixels with road attributes; the segmentation module is used to process the target optical image using an image segmentation model to obtain a first segmentation result; the target optical image is processed using a superpixel segmentation model to obtain a second segmentation result; the coarse optimization module is used to perform coarse optimization of road labels according to the first road buffer, the first segmentation result and the second segmentation result to generate a coarse optimization result; the prediction module is used to input the target optical image into a trained neural network for road prediction to obtain a road prediction result; the fine optimization module is used to combine the road prediction result, the coarse optimization result, the first road buffer and the second road buffer to perform fine optimization of road labels to generate a fine optimization result.
[0125] It should be noted that the road extraction device combining OSM data and optical images provided in the embodiment of the present invention is for realizing the above method. Its specific functions can be referred to the above method embodiments, which will not be described in detail here.
[0126] Fig.10 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.10As shown, the electronic device may include: a processor (processor) 1001, a communication interface (Communications Interface) 1002, a memory (memory) 1003 and a communication bus 1004, wherein the processor 1001, the communication interface 1002, and the memory 1003 communicate with each other through the communication bus 1004. The processor 1001 can call the logic instructions in the memory 1003 to execute the road extraction method, which includes: obtaining OSM data corresponding to the target optical image and preprocessing the OSM data; generating a first road buffer and a second road buffer according to the preprocessed OSM data; wherein the pixel attributes in the first road buffer are all roads, and the second road buffer should cover all pixels with road attributes; using an image segmentation model to process the target optical image to obtain a first segmentation result; using a superpixel segmentation model to process the target optical image to obtain a second segmentation result; performing rough optimization of road labels according to the first road buffer, the first segmentation result and the second segmentation result to generate a rough optimization result; inputting the target optical image into a trained network model to perform road prediction to obtain a road prediction result; and combining the road prediction result, the rough optimization result, the first road buffer and the second road buffer to perform fine optimization of road labels to generate a fine optimization result.
[0127] In addition, when the logic instructions in the above-mentioned memory 1003 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0128] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the road extraction method provided by the above-mentioned method embodiments.
[0129] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the road extraction method provided by the above-mentioned method embodiments is implemented.
[0130] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A road extraction method combining OSM data and optical images, characterized in that: include: Acquire OSM data corresponding to the target optical image, and preprocess the OSM data; Generate a first road buffer and a second road buffer according to the preprocessed OSM data; wherein the pixel attributes in the first road buffer are all roads, and the second road buffer should cover all pixels with road attributes; The target optical image is processed by an image segmentation model to obtain a first segmentation result; the target optical image is processed by a superpixel segmentation model to obtain a second segmentation result; Performing rough optimization of road labels according to the first road buffer, the first segmentation result, and the second segmentation result to generate a rough optimization result; Inputting the target optical image into a trained network model to perform road prediction and obtain a road prediction result; The road prediction result, the rough optimization result, the first road buffer zone and the second road buffer zone are combined to perform road label fine optimization to generate a fine optimization result.
2. The method for extracting roads by combining OSM data and optical images according to claim 1, characterized in that: Preprocessing the OSM data specifically includes: A grid matrix is constructed according to the spatial resolution of the target optical image, and the attributes of all pixels of the grid matrix are initialized as background; the geographic coordinates of the road in the OSM data are converted into pixel coordinates and the attributes of the pixels at the corresponding positions in the grid matrix are modified to roads to generate a road mask image; the road mask image is aligned with the target optical image.
3. The road extraction method combining OSM data and optical images according to claim 1 is characterized in that: The method of performing rough optimization of road labels according to the first road buffer, the first segmentation result, and the second segmentation result to generate a rough optimization result specifically includes: Traverse each object in the first segmentation result, and calculate the ratio T of the number of pixels in the intersection area between the object and the first road buffer to the number of pixels of the object area1 , if the ratio T area1 If the value is greater than a preset threshold, the object is considered to be a road, and all road objects are merged to generate a first rough optimization intermediate result; Traverse each superpixel in the second segmentation result, and calculate the ratio T of the number of pixels in the intersection area between the superpixel and the first road buffer to the number of pixels in the superpixel area2 , if the ratio T area2 If the value is greater than a preset threshold, the superpixel is considered to be a road, and all road superpixels are merged to generate a second rough optimization intermediate result; The first rough optimization intermediate result and the second rough optimization intermediate result are combined to obtain the rough optimization result.
4. The method for extracting roads by combining OSM data and optical images according to claim 1, characterized in that: The training process of the network model includes: taking the rough optimization result of the optical image as a label, training a preset network in combination with the optical image, and obtaining a trained network model.
5. The method for extracting roads by combining OSM data and optical images according to claim 1, characterized in that: Combining the road prediction result, the rough optimization result, the first road buffer zone, and the second road buffer zone to perform road label fine optimization to generate a fine optimization result specifically includes: The rough optimization result is subjected to an intersection operation with the first road buffer zone, the road prediction result is subjected to an intersection operation with the second road buffer zone, and the two intersection operation results are subjected to a union operation to obtain the refined optimization result.
6. The method for extracting roads by combining OSM data and optical images according to claim 1, characterized in that: Generating the first road buffer and the second road buffer according to the preprocessed OSM data specifically includes: The preprocessed OSM data is expanded using two convolution templates with different convolution kernel sizes to obtain a first road buffer and a second road buffer; the convolution kernel size of the convolution template corresponding to the first road buffer is smaller than the convolution kernel size of the convolution template corresponding to the second road buffer.
7. A road extraction method combining OSM data and optical images according to any one of claims 1 to 6, characterized in that: The image segmentation model adopts the SAM model.
8. A road extraction device combining OSM data and optical images, characterized in that: include: A preprocessing module, used for acquiring OSM data corresponding to the target optical image and preprocessing the OSM data; A buffer generation module is used to generate a first road buffer and a second road buffer according to the preprocessed OSM data; wherein the pixel attributes in the first road buffer are all roads, and the second road buffer should cover all pixels with road attributes; A segmentation module, used to process the target optical image using an image segmentation model to obtain a first segmentation result; and to process the target optical image using a superpixel segmentation model to obtain a second segmentation result; A rough optimization module, used for performing rough optimization of road labels according to the first road buffer, the first segmentation result and the second segmentation result, and generating a rough optimization result; A prediction module, used for inputting the target optical image into a trained network model to perform road prediction and obtain a road prediction result; A fine optimization module is used to combine the road prediction result, the rough optimization result, the first road buffer zone and the second road buffer zone to perform fine optimization of road labels and generate a fine optimization result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.