Illegal building high-precision detection method based on remote sensing image
Through multi-stage processing flow and multi-source data fusion strategy, the problem of environmental noise interference in remote sensing image change detection is solved, and high-precision and robust illegal building detection is achieved.
Patent Information
- Application Number
- CN202510437244.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
AI Technical Summary
The existing remote sensing image change detection technology is insufficient in dealing with multi-source environmental noise such as cloud shading, shadow superposition and mist interference, which makes it difficult to meet the high-precision requirements for detection accuracy and accuracy.
A multi-stage processing process is adopted, including data preprocessing (remote sensing image registration, mist removal, cloud and shadow extraction), change detection (based on deep neural networks), and data postprocessing (noise removal, boundary optimization, cloud and shadow filtering), combined with multi-source data fusion strategy, improve detection accuracy and robustness.
Significantly improve detection accuracy (F1-Score reaches 92.7%), effectively suppress environmental noise interference, and support small-target illegal buildings detection recall rate to reach 89%.
Smart Images

Figure CN120298916A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing, and specifically relates to a high-precision detection method for illegal buildings based on multi-temporal high-resolution remote sensing images, which is particularly suitable for urban governance and land supervision scenarios. Background Art
[0002] With the acceleration of urbanization, illegal building detection has become an important requirement for urban governance and land supervision. Traditional manual inspection methods rely on on-site investigations and visual interpretation of images, which have problems such as low efficiency, limited coverage, and strong subjectivity. Especially in high-density building areas or complex geographical environments, the missed detection rate increases significantly. In recent years, remote sensing image change detection technology based on deep learning has gradually replaced manual methods, but existing technologies still cannot meet the actual needs of high-precision and automated detection.
[0003] Current technologies can be divided into two categories: the first is the manual-dominated method, which manually compares multi-temporal remote sensing images or field measurement data, which is time-consuming and laborious, and the accuracy is greatly affected by the operator's experience. For example, in scenes with cloud occlusion or mist interference, it is difficult for humans to accurately identify changes in building outlines, resulting in a high rate of misjudgment. The second is a change detection method based on deep neural networks, such as automated systems using network architectures such as UNet. Although it improves detection efficiency, its technical implementation has significant defects: the existing methods based on deep neural networks are not robust enough when dealing with multi-source environmental noise such as cloud occlusion, shadow superposition, and mist interference, and the post-processing link has defects in the filtering mechanism of pseudo-change areas, resulting in high false detection and missed detection rates in the detection results, which is difficult to meet the engineering needs of high-precision semi-automatic detection.
[0004] It should be noted that the information disclosed in this background technology section is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or suggestion in any form that the information constitutes the prior art known to those skilled in the art. Summary of the invention
[0005] The purpose of the present invention is to provide a high-precision detection method for illegal buildings based on remote sensing images, so as to solve the problem that the existing methods are insufficiently robust when dealing with multi-source environmental noise such as cloud occlusion, shadow superposition and mist interference.
[0006] To solve the above problems, the present invention designs a high-precision detection scheme for illegal buildings based on remote sensing images according to the characteristics of remote sensing images, which includes the following steps:
[0007] S1 Data preprocessing: perform registration, mist removal, cloud detection and extraction, shadow extraction, and data augmentation operations on multi-temporal remote sensing images to generate image data without mist interference;
[0008] S2 Change Detection: Detect the changed areas in the preprocessed image based on a deep neural network model, and generate a binary mask of the changed areas;
[0009] S3 Data Post-processing: Remove noise, optimize the boundaries, filter clouds and shadows from the mask of the changed areas, and output the final illegal building detection result by combining a multi-source data fusion strategy.
[0010] Further, in the haze elimination in step S1, the dark channel dehazing algorithm is used to estimate the atmospheric light value and transmittance through the dark channel prior, and the transmittance efficiency map is optimized by combining guided filtering to generate the dehazed image.
[0011] Further, in the cloud detection in step S1, a semantic segmentation network based on PSPNet is used to generate a cloud segmentation model by training a cloud standard data set, and the trained cloud segmentation model is used to process the image to generate a binary cloud mask of the image.
[0012] Further, the shadow extraction in step S1 is based on HSV color space threshold segmentation and morphological post-processing to generate a binary shadow mask.
[0013] Further, in step S2, a change detection network is used for change detection to obtain a 0 / 1 binary mask layer representation of the changed areas, where 0 represents the unchanged areas and 1 represents the changed areas; the change detection network can be one of SNUNet, FC-EF, SegNet, FC-Siam-Diff, FC-Siam-Conv, UNet++_MSOF, IFN, and BIT.
[0014] Further, the noise removal in step S3 includes morphological erosion and dilation operations to filter out noise points with an area smaller than a set threshold and fill the hole areas in the mask.
[0015] Further, the boundary optimization in step S3 adopts an image slice center retention strategy, only retains the detection results in the central area of the slice, and fuses the boundary areas through an overlapping prediction and voting mechanism.
[0016] Further, the multi-source data fusion strategy in step S3 includes:
[0017] (1) Based on a multi-model voting mechanism, integrate at least two different change detection networks and output results;
[0018] (2) Dynamically adjust the decision threshold and adaptively select the threshold range according to the complexity of the image area.
[0019] Further, the method further includes a time series verification step: Exclude the legal change areas of registered buildings by comparing with the historical compliance building database.
[0020] Furthermore, the method further includes a data augmentation strategy, including color perturbation, geometric transformation, and occlusion simulation operations, for enhancing the generalization ability of the model.
[0021] The present invention adopts the above technical solutions. Compared with the prior art, the beneficial effects are as follows:
[0022] 1. The detection accuracy is significantly improved (F1-Score reaches 92.7%);
[0023] 2. It has strong environmental robustness and effectively suppresses the interference of clouds, shadows, and haze;
[0024] 3. It supports the detection of small target violations (recall rate for pixel area < 100 > 89%). BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 Shows a flowchart of a high-precision detection method for illegal buildings based on remote sensing images;
[0026] Figure 2 Shows a schematic diagram of remote sensing image registration based on the SIFT feature matching algorithm;
[0027] Figure 3 Shows the effect diagram of haze removal from remote sensing images;
[0028] Figure 4 Shows the network structure diagram of PSPNet;
[0029] Figure 5 Shows the effect diagram of cloud extraction from remote sensing images;
[0030] Figure 6 Shows the effect diagram of shadow extraction from remote sensing images;
[0031] Figure 7 Shows the comparison diagram of noise removal. The blue area in the left figure is noise points, the yellow area is holes, and the right figure is the processed result diagram;
[0032] Figure 8 Shows the network model structure diagram of SNUNet. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following further describes the technical features and advantages of the present invention in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the protection scope of the present invention more clearly defined.
[0034] In view of the characteristics of remote sensing images, the embodiments of the present invention design a detection method for illegal buildings based on remote sensing images. Please refer to Figure 1, is the flowchart of the method. The basic idea of the method is to solve the problem of multi-source environmental noise interference and improve the detection accuracy and robustness by synergistically optimizing the data preprocessing, change detection, and post-processing processes. The method can be divided into three stages: data preprocessing, change detection, and data post-processing from the processing flow.
[0035] In the data preprocessing stage, remote sensing image registration, haze removal, cloud detection and extraction, shadow extraction, and data augmentation operations are performed; in the change detection stage, based on the change detection network, the training and inference of the illegal building change detection model are carried out by tuning the network model parameters; in the data post-processing stage, the detection accuracy is further improved through various means such as noise removal, boundary removal, cloud filtering, shadow filtering, and change fusion. Next, a specific remote sensing image is used to describe the method in detail through detailed steps.
[0036] Step S1, Data Preprocessing
[0037] Data preprocessing mainly considers the interference factors existing in the data itself, and makes the data sent to the model be consistent and unbiased by detecting and filtering the interference factors. The interference detections used in the embodiments of the present invention include: (1) registration of remote sensing images; (2) haze removal; (3) cloud detection; (4) data augmentation (Color Jitter, rotation, mirroring, etc.).
[0038] (1) Registration of remote sensing images: The accurate registration of remote sensing images is the basis for the implementation of the change detection process. The existence of registration errors will cause difficult-to-eliminate pseudo-changes in the change detection process, affecting the accuracy of change detection. Registration errors are also a major error source in the entire change detection process. In the embodiments of the present invention, the SIFT (Scale Invariant Feature Transform) feature matching algorithm is used to register remote sensing images, and the matching effect is as Figure 2 shown.
[0039] The core goal of SIFT is to extract local feature points and their descriptors with scale and rotation invariance from the image, and use these features to achieve image matching. The detailed steps of the SIFT feature matching algorithm are as follows:
[0040] I. Scale Space Extreme Value Detection
[0041] 1. Construct a Gaussian pyramid
[0042] Perform multi-scale Gaussian convolution on the input image to generate Gaussian blurred images of different scales. The formula is expressed as:
[0043] L(x, y, σ) = I(x, y) * G(x, y, σ)
[0044] Where, is a Gaussian kernel, and σ controls the scale parameter. Multiple sets (Octave) of images are generated by downsampling layer by layer (such as reducing by a factor of 2), forming a pyramid structure.
[0045] 2. Generate the Difference of Gaussian (DoG) pyramid
[0046] Subtract adjacent scale Gaussian images to obtain the DoG image:
[0047] D(x, y, σ) = L(x, y, kσ) - L(x, y, σ)
[0048] The DoG pyramid is used to detect local extreme points. The extreme points need to be compared with 26 points in their neighborhood in three-dimensional space (scale, x, y) (adjacent points in the current layer and the upper and lower layers).
[0049] II. Precise localization of key points
[0050] 1. Eliminate low-contrast points
[0051] Fit the extreme points in the DoG space with a three-dimensional quadratic function to calculate their precise positions and contrasts. If the contrast is lower than the threshold (such as 0.03), it is determined as an unstable point and eliminated.
[0052] 2. Eliminate edge responses
[0053] Use the Hessian matrix to calculate the curvature of the key points. If the edge response is too strong (the ratio of the principal curvature to the secondary curvature exceeds the threshold, such as 10:1), then eliminate this point. The formula is:
[0054]
[0055] where r is the threshold ratio.
[0056] III. Direction assignment
[0057] 1. Calculate the gradient magnitude and direction
[0058] On the Gaussian scale image where the key point is located, calculate the gradient magnitude m(x, y) and direction θ(x, y) of the neighborhood pixels:
[0059]
[0060] 2. Generate the direction histogram
[0061] Divide 360° into 36 intervals (one bin per 10°), count the gradient directions of the pixels in the neighborhood of the key point (radius 3×1.5σ), and generate a histogram with weighted magnitudes (Gaussian weighted weights). Select the peak of the histogram as the main direction. If the peak of the secondary direction exceeds 80% of the main direction, then retain multiple directions.
[0062] IV. Generation of Feature Descriptors
[0063] 1. Division of Neighborhood Sub-regions
[0064] With the key point as the center, divide the 16×16 neighborhood rotated to the main direction into 4×4 sub-regions (each sub-region is 4×4 pixels).
[0065] 2. Calculation of Gradient Direction Histogram
[0066] Within each sub-region, calculate the gradient histogram in 8 directions (one bin for every 45°), generating a total of 4×4×8 = 128-dimensional vector. The gradient direction needs to be rotationally corrected according to the main direction of the key point to ensure rotational invariance.
[0067] 3. Normalization
[0068] Normalize the 128-dimensional vector and limit the maximum value of each direction component (such as 0.2) to enhance the robustness to illumination changes.
[0069] V. Feature Matching Strategy
[0070] 1. Nearest Neighbor Search
[0071] Use the K-D tree (K-Dimensional Tree) or brute-force matching method to calculate the Euclidean distance between the feature descriptors of two images and find the nearest neighbor (minimum distance) and the second-nearest neighbor (second-minimum distance).
[0072] 2. Nearest Neighbor Distance Ratio (NNDR) Screening
[0073] If the ratio of the nearest neighbor distance to the second-nearest neighbor distance is less than the threshold (such as 0.8), it is determined as a valid match; otherwise, it is regarded as a false match.
[0074] 3. Optimization by RANSAC Algorithm
[0075] Estimate the transformation matrix (such as affine transformation or homography matrix) through the random sample consensus algorithm, and eliminate the false matching point pairs that do not conform to geometric constraints.
[0076] This process ensures high-precision matching of remote sensing images through strict extreme value screening, direction normalization, and geometric verification.
[0077] (2) Haze Removal: For the haze phenomenon commonly existing in remote sensing images, the dark channel dehazing method is used to remove the haze, and the removal effect is as Figure 3 shown.
[0078] The basic principle of the dark channel dehazing method is as follows:
[0079] The Dark Channel Prior Dehazing Algorithm is based on the statistical laws of natural haze-free images: in local patches of non-sky regions, there is at least one color channel where the pixel values approach zero (i.e., the "dark channel"). The presence of haze disrupts this property, and by estimating the atmospheric light intensity and transmittance, the haze-free image can be reversely restored. Its core formula is the atmospheric scattering model:
[0080] I(x) = J(x)t(x) + A(1 - t(x))
[0081] Where:
[0082] I(x): Hazy image;
[0083] J(x): Haze-free image;
[0084] A: Global atmospheric light intensity;
[0085] t(x): Transmittance (describing the degree of light attenuation in haze).
[0086] The specific steps are as follows:
[0087] 1. Dark channel map calculation
[0088] For each pixel point of the input remote sensing image, take the minimum value of the RGB three channels to generate an initial dark channel map:
[0089]
[0090] Eliminate noise interference through minimum value filtering (such as a 15×15 window) to obtain a smooth dark channel map.
[0091] 2. Atmospheric light intensity estimation
[0092] Select the top 0.1% of pixel points with the highest brightness from the dark channel map, and the average value of the pixels with the highest brightness in the original image corresponding to them is used as the global atmosphere A.
[0093] 3. Transmittance estimation and optimization
[0094] According to the dark channel prior assumption, the initial estimate of the transmittance t(x) is:
[0095]
[0096] Where ω is an adjustment factor (usually taken as 0.95), which controls the dehazing intensity. Use guided filtering or Gaussian filtering to optimize the transmittance, eliminate block effects and retain edge details.
[0097] 4. Haze-free image restoration
[0098] Invert the haze-free image according to the atmospheric scattering model:
[0099]
[0100] Where t0 is the lower threshold of transmittance (such as 0.1) to avoid noise amplification caused by too small denominator.
[0101] 5. Post-processing and optimization
[0102] Perform contrast stretching or histogram equalization on the restored image to enhance details;
[0103] Aiming at the uneven distribution of haze in remote sensing images, the transmittance can be adaptively adjusted in combination with regional segmentation (such as dividing vegetation areas by NDVI).
[0104] (3) Cloud extraction: In change detection, the presence of clouds seriously affects the effect of change detection. It is very necessary to extract clouds before change detection. In the embodiment of the present invention, PSPNet is used for cloud detection and extraction. PSPNet is a publicly available network structure (proposed by Pyramid Scene Parsing Network), and the structure diagram is as Figure 4 shown. The extracted clouds are represented by a 0 / 1 binary mask layer (0 represents non-clouds, 1 represents clouds), and the extraction effect is as Figure 5 shown.
[0105] (4) Shadow extraction: For the shadows of buildings, a method based on threshold segmentation in the HSV color space and morphological post-processing is used for shadow detection and extraction. The extracted shadow layer is represented by a 0 / 1 binary mask layer as Figure 6 shown (0 represents shadow, 1 represents non-shadow).
[0106] The basic principle of the method based on threshold segmentation in the HSV color space and morphological post-processing is as follows:
[0107] 1. Characteristics of the HSV color space
[0108] Due to insufficient light, the shadow area shows the following characteristics in the HSV (Hue-Saturation-Value) model:
[0109] Hue (H): In the shadow area, due to the dominance of blue-violet light in the scattered light, the H value tends to be at a high angle (such as 180° - 300°);
[0110] Saturation (S): In the shadow area, due to the single wavelength of the scattered light, the S value is relatively high;
[0111] Value (V): The brightness of the occluded area is significantly lower than that of the non-shadow area.
[0112] 2. Threshold segmentation and morphological optimization
[0113] Set the threshold range using the physical characteristics of the HSV three channels, initially extract the shadow area, and then eliminate noise and correct the boundary through morphological operations (such as erosion and dilation) to improve the detection accuracy.
[0114] The specific process of this method is as follows:
[0115] 1. RGB to HSV color space
[0116] Convert the original remote sensing image from the RGB space to the HSV space, and separate the three channels of hue (H), saturation (S), and value (V).
[0117] 2. Multi-channel threshold segmentation
[0118] Set the HSV threshold for the shadow area:
[0119] Range of H: Adjust according to the scene (e.g., 180° ≤ H ≤ 300°);
[0120] Threshold of S: Higher than the mean of the background area (e.g., S ≥ 0.3);
[0121] Threshold of V: Lower than the brightness of the non-shadow area (e.g., V ≤ 0.4).
[0122] Filter the pixels that simultaneously meet the three conditions through logical AND operation to generate a preliminary shadow binary mask.
[0123] 3. Morphological post-processing
[0124] Opening operation: Erode first and then dilate to eliminate isolated noise points and small-area misdetection areas;
[0125] Closing operation: Dilate first and then erode to fill the holes inside the shadow and smooth the boundary.
[0126] 4. Generate a binary mask
[0127] Mark the processed shadow area as 0 and the non-shadow area as 1, and output the 0 / 1 binary mask layer.
[0128] (5) Data augmentation: Provide various data augmentation methods (such as Color Jitter, rotation, mirroring, etc.) for model training, and the data augmentation information is used for model training and change fusion in data post-processing.
[0129] Step S2. Change detection
[0130] Advanced change detection networks such as deep neural networks FC-EF, SegNet, FC-Siam-Diff, FC-Siam-Conv, Unet++_MSOF, IFN, SNUNet, and BIT are used for change detection to obtain a 0 / 1 binary mask layer identification of the changed area (0 represents the unchanged area, and 1 represents the changed area).
[0131] Step S3, Data post-processing
[0132] In the data post-processing stage, the results after algorithm processing are filtered and fused, mainly including noise removal, boundary removal, cloud filtering, shadow filtering, and change fusion.
[0133] (1) Noise removal (noise points and holes): Using erosion and dilation processing, according to the conventional area calculation of buildings, small-area noise points are filtered out. This can filter out most of the miscellaneous noise points and fill the holes existing in the results. Figure 7 The comparison diagram of noise removal is shown. In the left figure, the blue area is noise points, and the yellow area is holes. The right figure is the result after processing.
[0134] (2) Boundary removal: Buildings located at the image edge are divided into two parts due to image slicing, which easily leads to the loss of illegal building detection. Therefore, in the embodiments of the present invention, only 80% of the pixels in the middle of the predicted target are retained, the edge images are removed, and 50% overlapping prediction is set to ensure that each building can exist in the center of the image.
[0135] (3) Cloud filtering: According to the cloud mask information obtained in the data preprocessing step, the change results are filtered. The specific method is to remove the change detection results in the cloud area.
[0136] (4) Shadow filtering: According to the shadow mask information obtained in the data preprocessing step, the change results are filtered. The specific method is to remove the change detection results in the shadow area.
[0137] (5) Change fusion: In the data preprocessing step, various transformations are performed on the remote sensing image, such as ColorJitter, rotation, and mirroring. Here, the change results of the images under different transformation conditions are uniformly evaluated. Pixels with a final result greater than the threshold (set to 0.5 in the embodiments of the present invention) are determined as change points, and vice versa, they are determined as non-change points.
[0138] The method of the embodiments of the present invention realizes the full-process optimization from noise suppression to high-precision illegal building detection by eliminating environmental interference, enhancing the model feature learning ability, and performing refined post-processing.
[0139] The application scenarios of the embodiments of the present invention include, but are not limited to, fields such as urban planning supervision, land law enforcement, and disaster emergency assessment that require rapid identification of surface changes and the like.
[0140] The technical solution of the present invention will be described in detail below with a detailed implementation plan.
[0141] Collaborative optimization implementation in the S1 data preprocessing stage
[0142] S11 Multi-temporal remote sensing image registration
[0143] Algorithm selection: The SIFT (Scale-Invariant Feature Transform) algorithm is used for feature point matching, and the RANSAC (Random Sample Consensus) algorithm is combined to eliminate mismatched points.
[0144] The RANSAC algorithm is a mature algorithm, and its basic principle and process (used for eliminating mismatches in SIFT feature matching) are as follows:
[0145] Basic principle:
[0146] RANSAC (Random Sample Consensus) is a robust model parameter estimation method. Its core idea is: Fit the model by randomly sampling the minimum data set multiple times, and count the inliers (Inliers) that satisfy the model constraints. Finally, select the model with the most inliers as the optimal solution. This algorithm can still maintain stability for a high proportion of mismatches (such as more than 50%), and is suitable for estimating geometric transformation parameters after SIFT matching.
[0147] The specific steps are as follows (combined with the SIFT feature matching scenario):
[0148] 1. Input data preparation
[0149] The input is the initial set of matching point pairs P = {p i =(x i , x' i )} generated by the SIFT algorithm, where x i and x' i are the coordinates of the matching points in the two images respectively.
[0150] 2. Parameter initialization
[0151] Set the maximum number of iterations N, the inlier determination threshold ∈ (such as 2 - 5 pixels), and the minimum number of samples s required for the model (such as 3 pairs of points for affine transformation and 4 pairs of points for homography matrix).
[0152] 3. Iterative optimization
[0153] Step 1: Random sampling
[0154] Randomly select s pairs of matching points from P as hypothesized inliers.
[0155] Step 2: Model estimation
[0156] Calculate the geometric transformation model H (such as a homography matrix or an affine transformation matrix) based on the sampled points.
[0157] Step 3: Inlier screening
[0158] Use H to calculate the projection error e = ||x′ i - Hx i || of all matching points. If e i < ∈, then it is determined as an inlier, and count the current number of inliers n inlier .
[0159] Step 4: Dynamically update the number of iterations
[0160] According to the current inlier ratio ρ = n inlier / |P|, adjust the maximum number of iterations:
[0161]
[0162] where η is the confidence level (such as 0.99).
[0163] 4. Optimal model output
[0164] After the iteration ends, select the model H with the largest number of inliers best , and re - estimate the precise model using all inliers by the least - squares method.
[0165] 5. Rejecting false matches
[0166] Based on H best screen the final inliers, reject the outliers that do not meet the threshold, and output the refined matching point pairs. Parameter configuration:
[0167] SIFT feature point detection threshold: The contrast threshold is set to 0.04, and the edge threshold is set to 10;
[0168] RANSAC iteration times: 2000 times, the inlier threshold is set to 3 pixels;
[0169] Registration model: Select the affine transformation model, and the optimization objective function is to minimize the reprojection error.
[0170] Implementation effect: The root - mean - square error (RMSE) of the registered image ≤ 1.5 pixels, meeting the spatial consistency requirements for change detection.
[0171] S12 Mist elimination and cloud / shadow detection
[0172] Dark Channel Defogging Algorithm:
[0173] Dark channel window size: 15×15 pixels;
[0174] Atmospheric light estimation: Select the average value of the top 0.1% of the pixels in terms of brightness in the image;
[0175] Transmittance optimization: Use Guided Filter to smooth the transmittance map, with the filter radius set to 60 pixels and the regularization coefficient ε = 1e-6.
[0176] Cloud Detection Network (PSPNet):
[0177] Network architecture: PSPNet (Pyramid Scene Parsing Network), input size 512*512, output is a binary mask image (cloud / non-cloud) of the same size and with 1 channel;
[0178] Training data: Use the L8 SPARCS cloud annotation dataset, set the batch size to 16, the initial learning rate to 0.001, and adopt cosine annealing scheduling.
[0179] Shadow extraction method:
[0180] Threshold segmentation based on the HSV color space: H∈[0, 180], S∈[0.2, 0.6], V∈[0.1, 0.5];
[0181] Morphological post-processing: Use closing operation (5×5 rectangular kernel) to connect broken regions, and filter by area threshold (shadow regions smaller than 50 pixels are regarded as noise).
[0182] S13 Data Augmentation Strategy
[0183] Augmentation operation combination:
[0184] Color perturbation (ColorJitter): Brightness adjustment ±20%, contrast ±15%, saturation ±10%;
[0185] Geometric transformation: Random rotation (-30° to +30°), horizontal / vertical mirroring (probability 50%);
[0186] Occlusion simulation: Use GAN to generate cloud occlusion blocks (size 64×64 to 128×128, coverage area 5% to 15%).
[0187] Implementation process: Dynamically apply augmentation during the training phase, and randomly select 3 operation combinations for each batch.
[0188] Dynamic Tuning Implementation of the S2 Change Detection Network
[0189] S21 Network Architecture Selection and Improvement
[0190] Baseline Network: SNUNet (Siam-NestedUNet) is selected as the basic network, and its dual-encoder structure supports multi-temporal feature fusion.
[0191] SNUNet is a publicly available network model (proposed in SNUNet-CD: A Densely Connected Siamese Network for Change Detection of VHR Images), and the model structure is as Figure 8 shown.
[0192] Edge Gradient Perception Loss: On the basis of the cross-entropy loss, an edge gradient loss term extracted by the Sobel operator is added, and the weight coefficient λ = 0.3;
[0193] Multi-scale Feature Fusion: A Feature Pyramid Network (FPN) is introduced in the decoder to fuse feature maps with 1 / 4, 1 / 2, and original resolutions.
[0194] The feature fusion mechanism of SNUNet can be expressed by the red and gray connection lines in the above network structure diagram (left).
[0195] S22 Model Training and Inference
[0196] Training Parameters:
[0197] Optimizer: AdamW, initial learning rate 3e-4, weight decay 1e-4;
[0198] Learning Rate Scheduling: Linear warmup for 500 steps, followed by cosine annealing;
[0199] Training Epochs: 200 Epochs, with an early stopping strategy (terminate when the validation set loss does not decrease for 10 rounds).
[0200] Inference Acceleration:
[0201] The model is quantized using TensorRT (FP16 precision), and the inference speed is increased by 2.3 times;
[0202] The input image is divided into blocks for inference (Tile Size 1024×1024), and the overlapping regions are fused by voting.
[0203] S23 Performance Verification
[0204] Test Set Metrics: On the WHU-CD dataset, the improved SNUNet achieves an F1-Score of 92.7%;
[0205] Small target detection: The recall rate for illegal construction with a pixel area < 100 (such as rooftop additions) is increased to 89.5%.
[0206] Implementation of cascaded filtering for S3 data post - processing
[0207] S31 Noise removal and boundary optimization
[0208] Morphological filtering:
[0209] Erosion operation: A 3×3 cross - shaped kernel, iterated 2 times, to eliminate noise points with an area < 50 pixels;
[0210] Dilation operation: A 3×3 rectangular kernel, iterated 1 time, to fill the hole areas.
[0211] Boundary retention strategy:
[0212] Slice center retention area: Only retain 80% of the area at the center of the image slice (for example, in a 1024×1024 slice, the central 819×819 area is valid);
[0213] Overlapping prediction: The slice sliding step is set to 512 pixels (50% overlap), and the results of the overlapping areas are fused through a voting mechanism.
[0214] 3.2 Multi - level pseudo - change filtering
[0215] Primary filtering (mask elimination):
[0216] Cloud area: Invert the cloud mask generated by pre - processing, and then perform a pixel - by - pixel logical AND with the change detection result to directly eliminate false detections in the area covered by the cloud mask;
[0217] Shadow area: Perform a pixel - by - pixel logical AND between the shadow mask generated by pre - processing and the change detection result to directly eliminate false detections in the area covered by the shadow mask;
[0218] Secondary filtering (morphological optimization):
[0219] Region growing algorithm: Select change regions with a seed point area > 200 pixels, and set the similarity threshold to 0.7 (based on NDVI difference);
[0220] Tertiary filtering (temporal verification):
[0221] Compliant building database matching: Compare the historical compliant building outlines through the GIS system to exclude legal changes of registered buildings.
[0222] 3.3 Change fusion and output
[0223] Multi - model voting mechanism:
[0224] Integrated model: Detection results of three networks, SNUNet, BIT, and IFN;
[0225] Voting rule: Only the pixels determined to have changed by at least two models are retained;
[0226] Threshold optimization:
[0227] Dynamic threshold adjustment: Different thresholds (0.4 - 0.6) are set according to the complexity of the image area (such as urban / suburban areas) and automatically selected through Bayesian optimization.
[0228] Example 2: Detection of illegal buildings in a certain urban area
[0229] Input data: Two - phase remote sensing images with a resolution of 0.5 meters;
[0230] Pre - processing: After registration, haze elimination is carried out, and the cloud mask coverage rate reaches 98%;
[0231] Change detection: The SNUNet outputs the changed area;
[0232] Post - processing: After filtering out cloud shadows, legal changes are excluded through temporal verification, and finally the locations and areas of illegal buildings are output.
[0233] Example 3: Detection of additional buildings on roofs in mountainous areas
[0234] Data augmentation: Adding rotation and occlusion simulation to improve the sensitivity of the model to small targets;
[0235] Dynamic threshold: Set the threshold to 0.4 according to the terrain complexity, and the recall rate is increased by 12%.
[0236] In the description of the embodiments of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "center", "top", "bottom", "top part", "bottom part", "inner", "outer", "inner side", "outer side", etc. are based on the orientation or positional relationships shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. Among them, the "inner side" refers to the internal or enclosed area or space. The "periphery" refers to the area around a specific component or specific area.
[0237] In the description of the embodiments of the present invention, the terms "first", "second", "third", "fourth" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", "third", "fourth" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.
[0238] In the description of the embodiments of the present invention, it should be noted that, unless otherwise clearly specified and defined, the terms "installation", "connection", "linkage", and "assembly" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0239] In the description of the embodiments of the present invention, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.
[0240] In the description of the embodiments of the present invention, it should be understood that "-" and "~" represent the range between two numerical values, and this range includes the endpoints. For example, "A - B" represents a range greater than or equal to A and less than or equal to B. "A ~ B" represents a range greater than or equal to A and less than or equal to B.
[0241] In the description of the embodiments of the present invention, the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0242] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A high-precision detection method for illegal buildings based on remote sensing images, characterized in that, It includes the following steps: S1 Data preprocessing: Perform registration, haze removal, cloud detection and extraction, shadow extraction, and data augmentation operations on multi-temporal remote sensing images to generate image data with haze interference removed; S2 Change detection: Detect change regions in the preprocessed images based on a deep neural network model to generate a binary mask of the change regions; S3 Data postprocessing: Remove noise, optimize the boundaries, filter clouds and shadows from the mask of the change regions, and output the final illegal building detection results by combining multi-source data fusion strategies.
2. The method according to claim 1, characterized in that For haze removal in step S1, the dark channel dehazing algorithm is used to estimate the atmospheric light value and transmittance through the dark channel prior, and the transmittance efficiency map is optimized by guided filtering to generate the dehazed image.
3. The method according to claim 1, wherein For cloud detection in step S1, a semantic segmentation network based on PSPNet is used to generate a cloud segmentation model by training a cloud standard data set, and the trained cloud segmentation model is used to process the images to generate a binary cloud mask of the images.
4. The method according to claim 1, characterized in that, For shadow extraction in step S1, it is based on HSV color space threshold segmentation and morphological postprocessing to generate a binary shadow mask.
5. The method according to claim 1, characterized in that In step S2, a change detection network is used for change detection to obtain a 0 / 1 binary mask layer representation of the change regions, where 0 represents the unchanged regions and 1 represents the change regions; the change detection network can be one of SNUNet, FC-EF, SegNet, FC-Siam-Diff, FC-Siam-Conv, UNet++_MSOF, IFN, and BIT.
6. The method according to claim 1, characterized in that The noise removal in step S3 includes morphological erosion and dilation operations to filter out noise points with an area smaller than a set threshold and fill the hole regions in the mask.
7. The method according to claim 1, characterized in that, The boundary optimization in step S3 adopts an image slice center retention strategy, only retaining the detection results in the central region of the slice, and fusing the boundary regions through an overlapping prediction and voting mechanism.
8. The method according to claim 1, wherein The multi-source data fusion strategy in step S3 includes: (1) Based on a multi-model voting mechanism, integrate at least two different change detection networks and output results; (2) Dynamically adjust the decision threshold and adaptively select the threshold range according to the complexity of the image region.
9. The method according to claim 1, characterized in that, The method further includes a time series verification step: By comparing with the historical compliant building database, exclude the legal change regions of the registered buildings.
10. The method according to any one of claims 1-9, characterized in that, The method further includes a data augmentation strategy, including color perturbation, geometric transformation, and occlusion simulation operations, to improve the generalization ability of the model.
Citation Information
Cited By
Moire attack method based on mathematical simulation and gradient optimization
CN120580162A