Remote sensing image building extraction method combining U-Net and segmentation model
By combining the remote sensing image building extraction method with U-Net and segmentation model, the initial mask is generated and iteratively segmented, the problems of incomplete mask recognition and missing details in remote sensing images are solved, and more complete and accurate building segmentation is achieved.
Patent Information
- Application Number
- CN202510422349.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The problem of incomplete identification of target masks and missing details extracted from remote sensing images.
Combining the remote sensing image building extraction method of U-Net and segmentation model, a first mask is generated by processing the remote sensing image through U-Net, and the point prompt is determined according to the first mask, and the SAM segmentation model is input for further segmentation, a second mask is generated, and iterated through the differential mask until it meets the preset area.
Without increasing the running time, the integrity of building segmentation is effectively improved, the error detection rate is reduced, and the problems of incomplete mask recognition and lack of details are avoided.
Smart Images

Figure CN120047457A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a method for extracting buildings from remote sensing images by combining U-Net and a segmentation model. Background Art
[0002] With the development of Earth observation satellites and aerial remote sensing technologies, a large amount of multi-source remote sensing data provides rich information for various applications. As an important part of geographical information, building information is important basic data for applications such as population density estimation, land use management, and urban planning. Before deep learning became the mainstream method, building features were mainly constructed manually based on the visual features of images. Features in the images were obtained through image segmentation and feature extraction, and then traditional machine learning methods were used for classification. Such methods are often only applicable to specific data sets and have low accuracy, and cannot handle buildings with slightly complex shapes.
[0003] In recent years, the emergence of deep learning technologies has greatly improved the accuracy and efficiency of automatically extracting buildings from high-resolution images. By automatically learning image features, the subjectivity of manual feature selection is avoided. Since the emergence of convolutional neural networks, a series of semantic segmentation networks for automatic building extraction have been developed. When using U-Net-like deep learning network models to extract buildings, skip connections are used for feature fusion, which improves the details of the extraction results. However, due to the large semantic and scale differences between high-level and low-level features, the feature fusion is not sufficient, and the target masks extracted still have problems such as incomplete recognition and missing details. Summary of the Invention
[0004] The main objective of the present invention is to provide a method for extracting buildings from remote sensing images by combining U-Net and a segmentation model, aiming to solve the problems of incomplete recognition and missing details of the target masks extracted from remote sensing images.
[0005] To achieve the above objective, the method for extracting buildings from remote sensing images by combining U-Net and a segmentation model proposed by the present invention includes:
[0006] Obtain a remote sensing image, process the remote sensing image through U-Net, and determine the first mask of each building;
[0007] Determine the point prompts of each building according to the first mask of each building;
[0008] Input each of the point prompts into SAM segmentation to determine the second mask of each building;
[0009] Determine the differential mask of each building according to each of the first masks and each of the second masks;
[0010] Judge whether the area of the differential mask meets a preset area;
[0011] When it meets the preset area, the second mask is the final mask;
[0012] When it does not meet the preset area, use the differential mask as the first mask and perform the step of determining the point prompts of each building according to the first mask of each building.
[0013] Preferably, the step of obtaining the remote sensing image, processing the remote sensing image through U-Net, and determining the first mask of each building includes:
[0014] Obtain the remote sensing image, process the remote sensing image through U-Net, and determine the binary labels of each pixel in the remote sensing image;
[0015] Perform a denoising operation on each pixel with the binary label to eliminate the error items in each pixel with the binary label;
[0016] Process each pixel after eliminating the error items through the connected component analysis algorithm to determine the first mask of each building.
[0017] Preferably, the step of determining the point prompts of each building according to the first mask of each building includes:
[0018] According to each first mask, determine the centroid point of the first mask and two preset acquisition points on the long side of the minimum circumscribed rectangle of the first mask as the point prompts of each building.
[0019] Preferably, the formula for the point prompt is as follows:
[0020]
[0021] where P 3*2 is the centroid point and two preset acquisition points; (x,y) is the centroid point; (x+dx,y+dy) is one of the preset acquisition points; (x+dx,y+dy) is the other preset acquisition point; d is the offset.
[0022] Preferably, the preset area is 25 pixels.
[0023] Preferably, the step of inputting each point prompt into SAM for segmentation to determine the second mask of each building includes:
[0024] Input each point prompt into SAM for segmentation to determine the preliminary mask of each building;
[0025] Judge whether the edge straight lines of each preliminary mask meet the first preset condition;
[0026] When the first preset condition is not met, the preliminary mask that does not meet the first preset condition is removed;
[0027] When the first preset condition is met, the preliminary mask where the edge straight line that meets the first preset condition is located is determined as the second mask.
[0028] Preferably, the first preset condition is that the length of the edge straight line is greater than 10 pixels, and the number of edge straight lines of the preliminary mask where the edge straight line is located is greater than or equal to 2.
[0029] Preferably, the step of inputting each point prompt into SAM for segmentation to determine the preliminary mask of each building includes:
[0030] Judging whether the percentage of the pixel area of each preliminary mask in the minimum circumscribed rectangle of each preliminary mask meets the second preset condition;
[0031] When the second preset condition is not met, the preliminary mask that does not meet the second preset condition is removed;
[0032] When the second preset condition is met, the preliminary mask that meets the second preset condition is determined as the second mask.
[0033] Preferably, the second preset condition is that the percentage of the pixel area of the preliminary mask in the minimum circumscribed rectangle of each preliminary mask is greater than or equal to 30%.
[0034] Compared with the prior art, the present invention has at least the following beneficial effects:
[0035] By processing remote sensing images through U-Net to generate point prompts, then segmenting the point prompts through SAM, finally comparing to obtain a differential mask, and iterating the differential mask until the differential mask meets the preset area, the second mask is used as the final mask. Without increasing the running time, the integrity of building segmentation is effectively improved and the false detection rate is reduced, effectively avoiding the problems of incomplete mask recognition and missing details. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0037] Figure 1 It is a schematic flowchart of an embodiment of a method for extracting buildings from remote sensing images by combining U-Net and a segmentation model of the present invention;
[0038] Figure 2 This is a schematic flowchart of another embodiment of the method for extracting buildings from remote sensing images by combining the U-Net and the segmentation model of the present invention;
[0039] Figure 3 This is a schematic diagram for calculating point prompts;
[0040] Figure 4a This is Remote Sensing Image 1; Figure 4b This is the first effect diagram of generating point prompts based on Remote Sensing Image 1; Figure 4c This is the second effect diagram obtained by segmenting the first effect diagram using SAM;
[0041] Figure 5a This is Remote Sensing Image 2; Figure 5b This is the third effect diagram of the first mask after processing Remote Sensing Image 2 by U-Net; Figure 5c This is the fourth effect diagram of the second mask determined by segmenting the third effect diagram using SAM for the first time; Figure 5d This is the fifth effect diagram of the second mask determined by segmenting the third effect diagram using SAM for the second time; Figure 5e This is the sixth effect diagram of the differential mask determined based on the third effect diagram and the fourth effect diagram; Figure 5f This is the seventh effect diagram of the differential mask determined based on the third effect diagram and the fifth effect diagram;
[0042] Figure 6 This is a comparison diagram of embodiments. Among them, a is an example remote sensing image, b is the effect diagram of the ground truth of building extraction from the remote sensing image (the standard set of buildings extracted manually); c is the effect diagram of buildings extracted by U-Net; d is the effect diagram of buildings extracted by U-Net + L2; e is the effect diagram of buildings extracted by Ours (L2);
[0043] Figure 7 This is a local detail comparison diagram of embodiments. Among them, a is an example remote sensing image, b is the effect diagram of the ground truth of building extraction from the remote sensing image (the standard set of buildings extracted manually); c is the effect diagram of buildings extracted by U-Net; d is the effect diagram of buildings extracted by DeepLabv3+; e is the effect diagram of buildings extracted by U-Net + L2; f is the effect diagram of buildings extracted by Ours (L2).
[0044] The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with embodiments with reference to the accompanying drawings. Specific Embodiments
[0045] Embodiments of the present invention will be described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0046] Please refer to Figures 1 to 7 , to achieve the above object, in the first embodiment of the present invention, a method for extracting buildings from remote sensing images by combining U-Net and a segmentation model is provided, including:
[0047] Step S10, obtaining a remote sensing image, processing the remote sensing image through U-Net, and determining a first mask for each building;
[0048] Step S20, determining a point prompt for each building according to the first mask of each building (see Figures 4a to 4c );
[0049] Step S30, inputting each point prompt into SAM segmentation to determine a second mask for each building;
[0050] Step S40, determining a differential mask for each building according to each first mask and each second mask;
[0051] Step S50, judging whether the area of the differential mask meets a preset area;
[0052] Step S60, when it meets the preset area, the second mask is the final mask;
[0053] Step S70, when it does not meet the preset area, using the differential mask as the first mask, and performing the step of determining a point prompt for each building according to the first mask of each building.
[0054] By processing the remote sensing image through U-Net to generate point prompts, then segmenting the point prompts through SAM, finally obtaining a differential mask by comparison, and iterating the differential mask until the differential mask meets the preset area, the second mask is used as the final mask. Without increasing the running time, the integrity of building segmentation is effectively improved and the false detection rate is reduced, effectively avoiding the problems of incomplete mask recognition and missing details.
[0055] Specifically, U-Net is a fully convolutional deep neural network. U-Net can run quickly on the GPU, and its inference speed is faster than that of DeepLabv3+. However, it has the disadvantages of smooth edges in the segmentation results and confusing adjacent buildings, and its segmentation accuracy is slightly inferior to that of DeepLabv3+. To improve the inference speed as much as possible, this application selects the U-Net method instead of DeepLabv3+ as the building semantic extractor to generate a mask of the building for further generating point prompts. After training U-Net in the conventional way and retaining the training results of the network parameters, it can be used for segmentation, and then provide semantic prompts for EfficientViT-SAM.
[0056] Specifically, SAM is EfficientViT-SAM, which has strong generalization ability and can perform instance segmentation according to input point prompts, box prompts or text prompts. Since the data for training EfficientViT-SAM does not contain building annotations of high-resolution remote sensing images, the building cannot be directly segmented using the text prompt scheme; the point prompt function is simple and fast, but high-quality points need to be given. And U-Net has the ability to quickly extract the building mask. Therefore, this application proposes to use the point prompt segmentation function of EfficientViT-SAM for building extraction. Combining the advantages of these two methods, U-Net is used as the semantic extractor, and EfficientViT-SAM is used as the mask segmenter and combined together. By processing the segmentation results of U-Net to obtain point prompts, and then inputting the point prompts into EfficientViT-SAM to obtain high-quality segmentation results, this process is iterated until no new point prompts can be extracted from the difference map between the segmentation results of U-Net and EfficientViT-SAM. The overall process is as Figure 2 shown.
[0057] The method for extracting buildings from remote sensing images by combining U-Net and a segmentation model proposed in the second embodiment of the present invention is based on the first embodiment, and step S10 includes:
[0058] Step S11, obtaining a remote sensing image, processing the remote sensing image through U-Net, and determining the binary labels of each pixel in the remote sensing image;
[0059] Step S12, performing a denoising operation on each pixel with a binary label to eliminate the error items in each pixel with a binary label;
[0060] Step S13, processing each pixel after eliminating the error items through a connected component analysis algorithm to determine the first mask of each building.
[0061] To obtain high-quality building segmentation results using the point prompt segmentation function of EfficientViT-SAM, it is necessary to first generate point prompts for each building from the semantic segmentation results of U-Net.
[0062] Since the U-Net semantic segmentation results are essentially binary labels for each pixel indicating whether it is a building, rather than generating a mask for each building instance, it is first necessary to use a connected component analysis algorithm to obtain the mask of the building.
[0063] The Spaghetti Labeling 4-neighborhood connected component analysis is adopted to obtain as detailed a mask of the building as possible.
[0064] The segmentation results of U-Net will inevitably generate noise (i.e., error terms), resulting in the generation of incorrect prompt points. Therefore, before using the connected component analysis algorithm, it is necessary to perform denoising operations according to the preset area of the mask to exclude the influence of error terms.
[0065] The method for extracting buildings from remote sensing images by combining U-Net and a segmentation model proposed in the third embodiment of the present invention, based on the first embodiment, step S20, includes:
[0066] Step S21, according to each first mask, determine the center of gravity point of the first mask and two preset acquisition points on the long side of the minimum circumscribed rectangle of the first mask as the point prompts for each building.
[0067] Extract the center of gravity of each connected component as a point prompt.
[0068] Multiple point prompts can achieve better segmentation accuracy and stability than single point prompts. First, extract the center of gravity point of each first mask as one of the point prompts.
[0069] See Figure 3 , considering that the building mask is generally a rectangle with a large difference in length and width, calculate the minimum circumscribed rectangle of the first mask, and resample two preset acquisition points in the long side direction of the rectangle as the remaining point prompts. The center of gravity point of the first mask and the two preset acquisition points of the minimum circumscribed rectangle of the first mask, a total of three image points are used as point prompts. Denote the length of the long side of the minimum circumscribed rectangle of the first mask as l, and the angle between the long side l and the x-axis of the image coordinate system as α.
[0070] The method for extracting buildings from remote sensing images by combining U-Net and a segmentation model proposed in the fourth embodiment of the present invention, based on the third embodiment, the formula for the point prompt is as follows:
[0071]
[0072] Among them, P 3*2are the centroid point and two preset acquisition points; (x, y) is the centroid point; (x + dx, y + dy) is one of the preset acquisition points; (x + dx, y + dy) is the other preset acquisition point; d is the offset.
[0073] In the fifth embodiment of the present invention, the method for extracting buildings from remote sensing images by combining U-Net and a segmentation model, based on any one of the first to fourth embodiments, the preset area is 25 pixels.
[0074] The point prompts of each building are sequentially input into EfficientViT-SAM for refined segmentation, and finally the segmentation results are merged. However, the above strategy faces two major problems:
[0075] Limited by the segmentation accuracy of U-Net, the masks of different buildings with close distances may be connected, resulting in the point prompts generated according to the center not falling into each building, thus leading to incomplete segmentation results.
[0076] In the case of a small number of point prompts input, EfficientViT-SAM often only segments out half of the roof containing the prompt points.
[0077] See Figures 5a to 5f , so logical judgment is made according to the area of the differential mask, and steps S10 to S50 are iterated repeatedly until step S10 meets the preset area (i.e., 25 pixels) to obtain the final mask of the final structure.
[0078] In the sixth embodiment of the present invention, the method for extracting buildings from remote sensing images by combining U-Net and a segmentation model, based on any one of the first to fourth embodiments, step S30 includes:
[0079] Step S31, input each point prompt into SAM for segmentation to determine the preliminary mask of each building;
[0080] Step S32, judge whether the edge lines of each preliminary mask meet the first preset condition;
[0081] Step S33, when it does not meet the first preset condition, remove the preliminary mask that does not meet the first preset condition;
[0082] Step S34, when it meets the first preset condition, determine the preliminary mask where the edge line that meets the first preset condition is located as the second mask.
[0083] In the seventh embodiment of the present invention, the method for extracting buildings from remote sensing images by combining U-Net and a segmentation model, based on the sixth embodiment, the first preset condition is that the length of the edge line is greater than 10 pixels, and the number of edge lines of the preliminary mask where the edge line is located is greater than or equal to 2.
[0084] In the eighth embodiment of the present invention, the method for extracting buildings from remote sensing images by combining U-Net and a segmentation model, based on the sixth embodiment, step S31 includes:
[0085] Step S35: Determine whether the percentage of the pixel area of each preliminary mask in the minimum circumscribed rectangle of each preliminary mask meets the second preset condition;
[0086] Step S36: When it does not meet the second preset condition, remove the preliminary mask that does not meet the second preset condition;
[0087] Step S37: When it meets the second preset condition, determine the preliminary mask that meets the second preset condition as the second mask.
[0088] In the ninth embodiment of the present invention, the method for extracting buildings from remote sensing images by combining U-Net and a segmentation model, based on the eighth embodiment, the second preset condition is that the percentage of the pixel area of the preliminary mask in the minimum circumscribed rectangle of each preliminary mask is greater than or equal to 30%.
[0089] After the incorrect point prompt is input into EfficientViT-SAM in step S20, there will be two situations: no segmentation result or incorrect segmentation. To exclude these two incorrect segmentation results, they are excluded and filtered through the first preset condition and the second preset condition to ensure the integrity of the final mask and reduce the false detection rate.
[0090] Embodiment:
[0091] Experimental dataset and configuration:
[0092] This application selects a building dataset of a certain community as the training and validation dataset. The building dataset of a certain community is a large dataset composed of multi-source remote sensing images. The size of each image is 512 pixels × 512 pixels. For an overview and examples, see Figures 2 to 7 Among them, there are 8,189 aerial images, with a spatial resolution of about 0.075 m, covering a ground area of about 450 km in Christchurch, New Zealand 2 , including rural, residential, cultural and educational, and industrial areas. Randomly select 60% of the images in the dataset as the training set, 5% of the images as the evaluation set, and the remaining 35% of the images as the test set to evaluate the performance of the model.
[0093] The experimental hardware configuration of this application is: Intel(R) Core(TM) i7-11700 CPU @ 2.50 GHz octa-core, 32 GB of running memory, and an Nvidia GTX 1660 6 GB graphics card. The software configuration is: using the pytorch 2.1 deep learning framework under the Windows 10 system.
[0094] Accuracy evaluation indicators:
[0095] To evaluate the effect of building extraction, the Intersection over Union (IoU), Precision, and Recall are used to evaluate the detection accuracy. The formulas for calculating these three metrics through the building semantic segmentation results and the true building mask are shown in Formulas (1), (2), and (3) respectively.
[0096] Among them, Formula (1) is as follows:
[0097]
[0098] Formula (2) is as follows:
[0099]
[0100] Formula (3) is as follows:
[0101]
[0102] Among them, N TP is the number of true positive pixels, that is, the number of pixels whose true value is a building and the prediction result of the model also belongs to the building; N FN is the number of false negative pixels, that is, the number of pixels whose sample true value is a building but the model predicts it as a non-building; N FP is the number of false positive prediction results, that is, the number of pixels whose sample true value is a non-building but the model predicts it as a building.
[0103] Generally speaking, when the precision is high, the recall decreases; when the recall is high, the precision decreases.
[0104] Experiment and result analysis:
[0105] To complete the experiment, U-Net was trained on the aerial images of a certain community dataset; and for comparison, the more accurate DeepLabv3+ network was also trained. There are 5 pre-trained models available for EfficientViT-SAM according to the size and performance of the processed images. The detailed information table is as follows:
[0106]
[0107] Given that the image cropping size of a certain community building dataset is 512×512, three pre-trained models, L0, L1, and L2, were selected for comparative experiments, which are briefly denoted as Ours(L0), Ours(L1), and Ours(L2) respectively; the experiments that only use the point prompts generated by the U-Net mask and do not use iterative segmentation for input to EfficientViT-SAM are briefly denoted as U-Net+L0, U-Net+L1, and U-Net+L2. Additionally, to prove the efficiency of the method in this application, we also conducted a comparative experiment using the pre-trained model ViT-H with the best accuracy of the classic SAM (briefly denoted as Ours(ViT-H)), and the test results are shown in the following table:
[0108]
[0109]
[0110] As can be seen from the test result table, the iterative segmentation method using EfficientViT-SAM for refined segmentation has achieved a relatively large improvement in accuracy compared to directly using the U-Net method and the DeepLabv3+ method. Compared with DeepLabv3+, the method in this application using the L2 model has increased the intersection over union, precision, and recall from 75.36%, 88.35%, and 89.13% to 78.92%, 91.20%, and 89.73%. Although the U-Net+L0\L1\L2 methods do not use the iterative strategy, the segmentation accuracy has also achieved a considerable improvement. This is because the edges of the masks segmented by U-Net are usually blurred, and using EfficientViT-SAM can significantly improve the segmentation quality. The iterative segmentation method performs better than the non-iterative method on different pre-trained models, proving that the non-iterative U-Net+EfficientViT-SAM method improves the segmentation effect of buildings, but is affected by insufficient point prompts and fails to fully utilize the semantic information of the U-Net segmentation results to completely segment all buildings; while using the iterative strategy can effectively improve the utilization rate of the U-Net segmentation results and improve the final segmentation accuracy.
[0111] From the comparison of the data of U-Net and DeepLabv3+ in the test result table, it can be seen that as two classic deep segmentation networks, the segmentation performance of U-Net is slightly inferior to that of DeepLabv3+. In other words, DeepLabv3+ extracts individual buildings more completely, but the recall rate of its extraction is extremely close to that of U-Net. The advantage of the iterative strategy lies in making full use of semantic information. There is no obvious gap in the semantic extraction integrity between DeepLabv3+ and U-Net. Therefore, the accuracy difference between Ours(DeepLabv3+ / L2) and Ours(L2), which also adopt the iterative strategy, is not significant.
[0112] In terms of numerical values, the accuracy of both Ours (DeepLabv3+L2) and Ours (ViT-H) is slightly higher than that of Ours (L2), and the advantage of Ours (ViT-H) is more obvious. Ours (ViT-H) and Ours (L2) are based on the same semantic network U-Net, and the segmentation accuracy results mainly depend on the segmentation performance differences between the ViT-H and L2 pre-trained models. In most scenarios, the segmentation effect of ViT-H is better than that of L2. Therefore, Ours (ViT-H) shows the highest accuracy in the research of this application. However, when comparing the running time and memory usage, it is found that the computing power consumption of ViT-H is too high, and the running time is more than twenty times that of Ours (L2), which is not suitable for practical applications.
[0113] The segmentation effects of four methods, namely U-Net, DeepLabv3+, U-Net+L2, and Ours (L2), are as Figure 6 shown. The segmentation result of Ours (L2) is more complete than that of U-Net+L2, and the extracted incomplete buildings are more complete, which proves the effectiveness and necessity of the iterative strategy. The local details are as Figure 7 shown.
[0114] In the description of this specification, the descriptions referring to terms such as "one embodiment", "another embodiment", "other embodiments", or "the first embodiment to the Xth embodiment" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention.
[0115] In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example.
[0116] Moreover, the specific features, structures, materials, method steps, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0117] It should be noted that in, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including the element.
[0118] The serial numbers of the embodiments of the present invention described above are only for description and do not represent the advantages or disadvantages of the embodiments. Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0119] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention. All of these are within the protection scope of the present invention.
Claims
1. A remote sensing image building extraction method combining U-Net and segmentation model, characterized in that: include: Acquire a remote sensing image, process the remote sensing image through U-Net, and determine a first mask of each building; Determine the point prompt of each building according to the first mask of each building; Input each of the point prompts into SAM segmentation to determine a second mask for each building; Determining a differential mask of each building according to each of the first masks and each of the second masks; Determining whether the area of the differential mask meets a preset area; When the preset area is met, the second mask becomes the final mask; When the preset area is not met, the differential mask is used as the first mask, and the step of determining the point prompt of each building according to the first mask of each building is performed.
2. The remote sensing image building extraction method combining U-Net and segmentation model as claimed in claim 1, characterized in that: The step of acquiring a remote sensing image, processing the remote sensing image through U-Net, and determining a first mask of each building includes: The remote sensing image is acquired, and the remote sensing image is processed by U-Net to determine a binary label of each pixel in the remote sensing image; Performing a denoising operation on each of the pixels marked with the binary label to eliminate error items in each of the pixels marked with the binary label; The first mask of each building is determined by processing each pixel after eliminating error items through a connected component analysis algorithm.
3. The remote sensing image building extraction method combining U-Net and segmentation model as claimed in claim 1, characterized in that: The step of determining the point prompt of each building according to the first mask of each building comprises: According to each of the first masks, the center of gravity of the first mask and two preset collection points of the long side of the minimum circumscribed rectangle of the first mask are determined as the point prompts of each building.
4. The remote sensing image building extraction method combining U-Net and segmentation model as claimed in claim 3, characterized in that: The formula for the point prompt is as follows: Among them, P 3*2 is the center of gravity and two preset collection points; (x, y) is the center of gravity; (x+dx, y+dy) is one of the preset collection points; (x+dx, y+dy) is the other preset collection point; d is the offset.
5. The remote sensing image building extraction method combining U-Net and segmentation model as described in any one of claims 1 to 4, characterized in that: The preset area is 25 pixels.
6. The remote sensing image building extraction method combining U-Net and segmentation model as described in any one of claims 1 to 4, characterized in that: The step of inputting each of the point prompts into SAM for segmentation and determining the second mask of each building comprises: The point prompts are input into SAM for segmentation to determine the preliminary mask of each building; Determining whether the edge straight lines of each of the preliminary masks meet a first preset condition; When the first preset condition is not met, removing the preliminary mask that does not meet the first preset condition; When the first preset condition is met, the preliminary mask where the edge straight line meeting the first preset condition is located is determined as the second mask.
7. The remote sensing image building extraction method combining U-Net and segmentation model as claimed in claim 6, characterized in that: The first preset condition is that the length of the edge straight line is greater than 10 pixels, and the number of the edge straight lines in the preliminary mask where the edge straight line is located is greater than or equal to 2.
8. The remote sensing image building extraction method combining U-Net and segmentation model as claimed in claim 6, characterized in that: The step of inputting each of the point prompts into SAM for segmentation and determining a preliminary mask of each building comprises: Determining whether the percentage of the pixel area of each of the preliminary masks to the minimum circumscribed rectangle of each of the preliminary masks meets a second preset condition; When the second preset condition is not met, removing the preliminary mask that does not meet the second preset condition; When the second preset condition is met, the preliminary mask meeting the second preset condition is determined as the second mask.
9. The remote sensing image building extraction method combining U-Net and segmentation model as claimed in claim 8, characterized in that: The second preset condition is that the percentage of the pixel area of the preliminary mask to the minimum circumscribed rectangle of each preliminary mask is greater than or equal to 30%.
Citation Information
Patent Citations
Remote sensing image building instance mask extraction method and system, medium and equipment
CN112991301A
Land parcel segmentation method, device and equipment and storage medium
CN118470316A
Semantic segmentation method and system for remote sensing image multi-level mask classification optimization
CN119762779A
Mask-guided image mosaic line generation method and apparatus, computer device, and storage medium
WO2024016368A1