A ground surface deformation monitoring and classification method based on InSAR and deep learning
By combining InSAR with deep learning technology and using random forests and ViT networks for SAR image classification, the problem of InSAR deformation monitoring being unable to classify in existing technologies has been solved, achieving high-precision deformation monitoring and classification, and supporting geological disaster early warning and urban planning.
Patent Information
- Application Number
- CN202510140352.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-02-08
AI Technical Summary
Existing InSAR deformation monitoring methods cannot perform effective classification analysis after detecting surface deformation, and the classification accuracy is limited.
By combining InSAR and deep learning technologies, multiple time-series SAR images and DEM data are acquired, and registration, interferometry, and DEM differencing are performed. PS pixels are selected using amplitude deviation and coherence coefficient thresholds, a spatial network is constructed, model parameters are calculated, deformation information is recovered, and random forest and ViT networks are used for image classification to improve classification accuracy.
It significantly improves the classification accuracy and efficiency of SAR images, enabling accurate identification of surface deformation types and providing a scientific basis for geological disaster early warning, urban planning, and resource management.
Smart Images

Figure CN120044524B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing and information technology, in particular to a ground surface deformation monitoring and classification method based on InSAR and deep learning. BACKGROUND
[0002] The ground surface deformation monitoring and classification method based on InSAR and deep learning is an important research direction in the field of remote sensing. InSAR technology measures the small phase difference between two radar signals to infer the height change of the ground surface. Its principle is to use a SAR system to emit microwave signals to the earth's surface and receive the return signals, and through phase interference processing, a phase difference image reflecting the height change of the ground surface is obtained, and then the ground surface deformation information is analyzed. InSAR technology has been widely used in earthquake monitoring, mineral resource development, urban construction and other fields. However, it also has some limitations, such as poor radar interference results in some geological environments (such as areas with severe vegetation coverage, wetlands, snow-covered areas, etc.), difficulty in phase delay unwrapping, and sensitivity to weather changes.
[0003] Deep learning technology uses models such as Convolutional Neural Networks (CNN), Fully Convolutional Networks, U-Net Networks, and Vision Transformer (ViT) networks to automatically analyze ground surface features from remote sensing data obtained from satellite, unmanned aerial vehicle and other earth observation platforms, and identify changes in terrain. This technology is widely used in urban planning, environmental monitoring, disaster warning, land use change analysis and other fields.
[0004] Existing work on InSAR deformation classification and analysis mostly uses machine learning models, with limited classification accuracy. Moreover, existing InSAR deformation monitoring methods cannot classify and analyze deformation after detecting ground surface deformation. SUMMARY
[0005] To overcome the deficiencies of the prior art, the present application provides a ground surface deformation monitoring and classification method based on InSAR and deep learning, which improves the classification accuracy of existing SAR image classification methods and solves the problem that existing InSAR deformation monitoring methods cannot classify and analyze deformation after detecting ground surface deformation.
[0006] The technical scheme provided by the present application is a ground surface deformation monitoring and classification method based on InSAR and deep learning, comprising the following steps:
[0007] (1) Collect multi-scene time series SAR images and DEM data in the monitoring area;
[0008] (2) Register, interfere and DEM difference of multi-scene time series SAR images;
[0009] (3) Select PS pixels by using amplitude dispersion and coherence coefficient threshold;
[0010] (4) Link PS pixels to construct a spatial network
[0011] The PS spatial network is constructed by using TIN, and the Delaunay triangulation method is applied to the spatial distribution of PS points to form a stable triangular grid structure;
[0012] (5) Solve the model parameters of the arc segment of the spatial network
[0013] On the PS spatial network, the phase difference on the arc segment of two adjacent coherent points in the interferogram is calculated, and the absolute value of the complex overall coherence coefficient is taken as the parameter estimation The integral on the arc segment is solved to all PS monitoring points by selecting a stable reference point, and the deformation rate and terrain residual of the monitoring area are obtained;
[0014] (6) Remove the atmospheric error influence from the residual phase to restore the deformation information of the PS monitoring points;
[0015] (7) Use random forest to classify the registered SAR image once
[0016] Filter each registered SAR image to reduce noise, manually annotate samples based on intensity images, classify SAR images according to actual needs and classifier performance, train the model using annotated samples and intensity features, randomly select a part of the samples as the training set, and the remaining labeled samples as the validation set to verify the model accuracy, and use the trained random forest model to predict and classify the entire monitoring area;
[0017] (8) Use the first classification sample and ViT network to classify the SAR image twice
[0018] Input the SAR image block into the ViT network backbone to construct the feature expression, and then use the Softmax classifier with cross-entropy loss function for supervised classification. The label of the supervised classification comes from the first classification result generated by the random forest. Through the above processing, the SAR image can be further subdivided, and the accuracy of the first classification can be improved;
[0019] (9) Based on the multi-classification result, classify, count and analyze the causes of deformation.
[0020] Further, the step (1) utilizes SAR satellite platform to download SAR image data of the required area, arranges the acquired multi-scene SAR image data according to time sequence, ensures that each image has clear time stamp and geographic position information, and then performs preprocessing work of radiation correction, geometric correction and multi-view processing on the acquired SAR image data.
[0021] Further, the DEM data acquired in the step (1) is arranged according to area, and then format conversion and precision evaluation are performed on the DEM data, so as to ensure that the DEM data matches the geographic information of the SAR image.
[0022] Further, the specific steps of the step (2) are as follows:
[0023] Image registration: the registration of the SAR image includes coarse registration and fine registration, the coarse registration utilizes satellite orbit parameters, InSAR geometry and SAR positioning equation to obtain the homonymic point of the center point of the reference image in the image to be registered, the fine registration is based on local pixel offset estimation of cross-correlation function, so as to achieve higher registration accuracy, and finally the image to be registered is subjected to coordinate transformation and pixel interpolation resampling;
[0024] Interference processing: the registered image is subjected to interference processing to generate an interference graph, and the interference graph is subjected to filtering processing to remove the influence of noise and incoherent area;
[0025] DEM difference: external digital elevation model data is acquired as reference terrain information, the external DEM data is utilized to remove the influence of terrain phase from the interference graph to obtain a differential interference graph reflecting only the ground surface deformation.
[0026] Further, in the step (3), on the basis of SAR image registration, a double-threshold method of amplitude deviation and coherence coefficient is adopted to select PS pixels,
[0027] Under the condition of high signal-to-noise ratio, the phase deviation σ φ of the pixel is estimated by the amplitude deviation D A , and the amplitude deviation is expressed by the following formula:
[0028]
[0029] Wherein, m A and σ A respectively represent the amplitude time average value and amplitude standard deviation of the pixel, when the value of the amplitude deviation index of the pixel point is lower than a certain threshold value, the pixel point is selected as a PS candidate point.
[0030] Coherence coefficient threshold: the coherence coefficient threshold can further remove water points, wetlands, ice and snow low-coherence points, and the complex coherence of two complex signals S1 and S2 with zero mean is defined as:
[0031]
[0032] E(x) represents the expected value, the coherence amplitude within the sampling window distance n and azimuth m The maximum likelihood estimation of the coherence amplitude is expressed as:
[0033]
[0034] where (i, j) represents the pixel coordinates, the coherence amplitude of each pixel in the selected interferogram is estimated after removing the terrain phase and flat phase, and the average coherence map is generated by the following formula:
[0035]
[0036] where N represents the number of interferograms, the coherence amplitude of the i-th interferogram, and the pixels exceeding the set average coherence threshold are selected as PS candidate points.
[0037] Further, in the step (5):
[0038] In the PS space network, assuming that the two adjacent coherent points in the i-th interferogram are x and y, the phase difference on the arc segment connecting them is expressed as:
[0039]
[0040] where, is the wrapping operation; Δv(x, y) and Δz(x, y) are the differences in deformation rate and DEM residual between the coherent points x and y, respectively; B ⊥ is the vertical baseline; t j is the time baseline; λ is the radar wavelength; R and θ are the radar slant range and incidence angle, respectively; then represents the difference in residual phase, since the atmospheric characteristics, nonlinear deformation, etc. of adjacent coherent points generally have no significant difference, for the solution space search method, it is assumed that:
[0041]
[0042] At this time, the absolute value of the complex overall coherence coefficient can be used as a reliable standard for parameter estimation:
[0043]
[0044] where M is the number of interferograms, e j is in complex form, The value range is [0, 1]. Finally, a stable reference point is selected to integrate the solution on the arc segment onto all PS monitoring points to obtain the deformation rate and topographic residual results of the monitored area.
[0045] Furthermore, in step (6), linear deformation and topographic components are removed from the interferometric phase. Based on this, the residual phase is spatiotemporally filtered to remove atmospheric errors. Finally, the linear deformation and nonlinear deformation are combined to obtain the final deformation time series result.
[0046] Furthermore, in step (7), the random forest is a decision tree method based on the guided clustering algorithm. Given a dataset X = x1, ..., xn, the Bagging method repeatedly samples from the training set with replacement k times, and then trains the tree model on these samples. After training, the predicted value for the unknown sample x is obtained. This is achieved by averaging the predictions of all individual regression trees on the training sample x′:
[0047]
[0048] Where f represents the prediction model.
[0049] The standard deviation σ of the predictions of all individual regression trees on x′ is used as an estimate of the uncertainty of the prediction:
[0050]
[0051] Furthermore, the ViT network structure in step (8) includes a data preprocessing layer, a self-attention layer, and a multi-sensor layer.
[0052] Data preprocessing layer: First, an m×m image patch is input into the model, and the image patch is uniformly divided into N... 2 Each part is expanded according to its spatial dimension to form a (m / N)*(m / N) dimensional vector. All vectors are then superimposed. The input will be from R... m×m become Specifically, when the value of N equals m, each partition of the original image patch will contain only one pixel, and then a learnable matrix will be generated. For linear projection, the matrix product is called the pixel embedding. The above operation is consistent with the fully connected layer of a neural network. After obtaining the pixel embedding, a randomly initialized vector, called the class embedding, is prepared. Finally, the two are stacked to form a new pixel embedding, which is used as the input to the next layer, with a size of [missing value]. Where 1 represents class embedding, m 2 It is the old pixel embedding. During the learning process, class embedding is used to integrate the information of the rest, and finally represent the original SAR image patch.
[0053] Self-attention layer: first utilize learnable matrices Embedding the pixels into three dimensionally identical variables (q, k, v) and then extracting information and constructing feature representations by taking and utilizing attention, in this process, a weight matrix A can be generated based on the pairwise similarity between two sequence elements, the i-th element of which is written as
[0054]
[0055] Here N' = 1 + m 2 , The similarity is obtained by calculating the dot product of q and k, and the scaling factor is used in actual calculation, and the output is mapped to (0, 1) using the Softmax activation to generate weights, the above equation determines the importance of the i-th position by evaluating the correlation between q i and k, and the output SA is obtained in the following way:
[0056]
[0057] Where the i-th element of SA is the weighted sum of v and the importance of the i-th position, this operation releases the attention obtained from (10) and constructs a new feature representation based on the mapped pixel embedding;
[0058] MLP layer: including two fully connected layers and a GeLU activation function, used to receive the output from the attention layer, in this stage, the model will map the learned features to the class probability space and further process it, input the Softmax classifier to generate the final image classification prediction.
[0059] Further, the step (9) is to count the occurrence frequency, degree and object type involved of each type of deformation, analyze the causes of each type of deformation, including the size, direction, duration of external force, and the material properties, structural characteristics of the object itself.
[0060] The beneficial effects of the present application are:
[0061] (1) The present application provides a method for secondary classification of SAR images by combining random forest and ViT network, which can significantly improve the classification accuracy and efficiency. The combination of random forest and ViT network for secondary classification of SAR images can make full use of the advantages of both. First, random forest can extract rich features from SAR images, which can be used as input for ViT network. Second, ViT network can perform global and deep processing on these features, further improving the accuracy of classification. In addition, due to the strong feature extraction and representation ability of ViT network, it can also help random forest to better handle high-dimensional and complex data, thereby improving the overall classification efficiency. In summary, compared with existing SAR image classification methods, the present application can significantly improve the accuracy and efficiency of classification, providing a more accurate and reliable basis for subsequent applications of SAR images.
[0062] (2) The present application can classify and analyze InSAR deformation by using land classification results. Land classification results provide detailed information about surface cover types, which is crucial for understanding InSAR deformation monitoring results. Different land use types (such as urban land, agricultural land, forest, etc.) have different effects on surface deformation, so combining land classification results can more accurately analyze the factors causing InSAR monitored deformation. In addition, InSAR deformation analysis combined with land classification results can provide more scientific basis for geological disaster warning, urban planning, resource management, etc. By identifying deformation patterns and trends under different land use types, potential geological disaster risks can be detected in a timely manner, providing decision support for disaster prevention and mitigation. BRIEF DESCRIPTION OF DRAWINGS
[0063] Figure 1 is a technical scheme flow chart of surface deformation monitoring and classification method based on InSAR and deep learning;
[0064] Figure 2 is a primary classification method diagram based on random forest;
[0065] Figure 3 is a secondary classification method diagram based on ViT network. DETAILED DESCRIPTION
[0066] The technical solutions in the embodiments of the present application will be described in detail below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0067] As Figure 1As shown, the ground deformation monitoring and classification method based on InSAR and deep learning specifically includes the following steps:
[0068] Step 1, collect multi-scene time series SAR images and DEM data of the monitoring area
[0069] SAR data acquisition: Use SAR satellite platforms (such as Sentinel-1, ALOS-2 and TerraSAR-X, etc.) to download SAR image data of the required area. Open SAR satellite images can be downloaded through specialized commercial websites or free platforms. Ensure that the downloaded SAR image data covers the monitoring area and has time series characteristics, that is, contains image data at different time points. Organize the acquired multi-scene SAR image data according to the time sequence, ensure that each image has a clear timestamp and geographic location information. Perform radiation correction, geometric correction, and multi-view processing on the acquired SAR image data for preprocessing to improve image quality and reduce noise.
[0070] DEM data acquisition: DEM data can be downloaded for free from websites, or the DEM acquisition tool of SARscape and other radar image processing software can be used to automatically download the DEM of the corresponding area. Similarly, the acquired DEM data is also organized according to the area for subsequent processing and analysis. Perform necessary format conversion and precision evaluation on the DEM data to ensure that it matches the SAR image geographic information.
[0071] Step 2: Register, interfere and DEM difference the multi-scene time series SAR images
[0072] Image registration: SAR image registration methods include coarse registration and fine registration. Coarse registration: Identify a small number of homonymous points (usually automatically selected), and use satellite orbit parameters, InSAR geometry and SAR positioning equations to obtain the homonymous points of the reference image center point in the image to be registered. Fine registration: Use a more detailed method, such as local pixel offset estimation based on cross-correlation function, to achieve higher registration accuracy. Finally, perform coordinate transformation and pixel interpolation resampling on the image to be registered.
[0073] Interference processing: Perform interference processing on the registered images to generate an interference graph that reflects the phase difference caused by surface deformation and terrain information. Filter the interference graph to remove noise and the effects of incoherent areas, improving the quality of the interference graph.
[0074] DEM difference: Obtain external digital elevation model data as reference terrain information. Use external DEM data to remove the terrain phase effect from the interference graph to obtain a differential interference graph that only reflects surface deformation.
[0075] Step 3: Select PS pixels by amplitude deviation and coherence coefficient threshold
[0076] On the basis of SAR image registration, the application selects PS pixels by using the amplitude deviation and the coherence coefficient double threshold method.
[0077] Amplitude deviation index: Since the PS technique directly processes a series of single master image interference pairs without considering the baseline limitation, the interferogram is seriously affected by the spatial decorrelation factor, therefore, the criterion based on spatial coherence estimation can only select points with high coherence, and cannot select points with high phase quality. The amplitude deviation index is based on the statistical characteristic relationship between the amplitude deviation and the phase standard deviation in the time sequence, that is, when the phase standard deviation is less than 0.25 rad, in the case of high signal-to-noise ratio, the amplitude deviation is approximately equal to the phase standard deviation, so that the amplitude deviation index can be used to select PS pixels.
[0078] And according to the above, under the condition of high signal-to-noise ratio, the phase deviation σ φ of the pixel can be estimated by the amplitude deviation D A . The amplitude deviation can be expressed by the following formula:
[0079]
[0080] Wherein, m A and σ A respectively represent the amplitude time average (mathematical expectation) and the amplitude standard deviation of the pixel. For a sufficient number of SAR images, when the value of the amplitude deviation index of the pixel point is lower than a certain threshold (such as 0.3), the pixel point is selected as a PS candidate point.
[0081] Coherence coefficient threshold: the criterion based on spatial coherence estimation can select points with high coherence, and the coherence coefficient threshold can further remove water points and low-coherence points such as wetlands, ice and snow. The complex coherence of two complex signals S1 and S2 with zero mean is defined as:
[0082]
[0083] In the formula, E(x) represents the expectation value. The above formula is based on the ergodic hypothesis, therefore, in the sampling window n x m (range x azimuth), the maximum likelihood estimation of the coherence amplitude can be expressed as:
[0084]
[0085] Wherein, (i, j) represents the pixel coordinates. After removing the terrain phase and flat ground phase, the coherence amplitude of each pixel in the selected interferogram is estimated, and the average coherence map is generated by the following formula:
[0086]
[0087] where N represents the number of interferograms. represents the coherence amplitude of the i-th interferogram, and the pixels exceeding the set average coherence threshold are selected as PS candidates.
[0088] In actual operation, the amplitude dispersion threshold and the coherence coefficient threshold are usually set, and only the pixels that meet both thresholds are selected as PS pixels.
[0089] In addition, in order to improve the accuracy and reliability of PS pixel selection, other methods such as phase analysis method or multi-threshold screening method can also be combined to adapt to the PS selection requirements in different surface environments.
[0090] Step 4: Linking PS pixels to construct spatial network
[0091] After selecting high-quality PS pixels, Triangulated Irregular Network (TIN) is used to construct PS spatial network. This network is usually composed of multiple PS pixels through certain connection relationships, which can be determined based on the spatial distance between pixels, phase relationship and other factors. By constructing such a spatial network, the spatial distribution characteristics of surface deformation and the interaction relationship between different deformation regions can be further analyzed.
[0092] The process of constructing PS spatial network using TIN is mainly to apply the Delaunay triangulation method to the spatial distribution of PS points to form a stable triangular grid structure. In PS-InSAR technology, the selected PS points have certain distribution characteristics in geographical space, and in order to more effectively analyze the deformation information of these points, it is necessary to construct the spatial network between them. Delaunay triangulation is a commonly used triangulation method, which follows the "minimum angle maximum" and "empty circumscribed circle" criteria for triangulation. In this process, the algorithm will automatically select the optimal triangular connection mode according to the spatial position of the PS points to meet the criteria of Delaunay triangulation. Through Delaunay triangulation, PS points are connected into a series of triangles to form a stable triangular grid structure, i.e. PS network. This network can be used for subsequent deformation monitoring and analysis.
[0093] Step 5: Solving model parameters of spatial network segments
[0094] The arc segment deformation parameters and terrain error parameters of the aforementioned spatial network are estimated. Since the interferogram phase is entangled, the obtained phase observation values only represent the principal part of the true values. Phase unwrapping is required to recover integer ambiguity. The key to multi-temporal InSAR parameter solving is how to solve or avoid the phase ambiguity problem. In PS-InSAR technology, the commonly used inversion model is a linear model, which assumes that the deformation of the PS point changes linearly with time. During the inversion process, the fitting relationship between the multi-temporal interferogram phase and the model parameters is utilized. Optimization algorithms (such as least squares method, solution space search method, etc.) are used to estimate the model parameters, and the accuracy of the solved model parameters is evaluated to ensure the reliability and accuracy of the results.
[0095] In the PS space network, assuming that two neighboring coherent points are x and y in the i-th interferogram, the phase difference on the arc connecting them is... It can be represented as:
[0096]
[0097] in, For the winding operation; Δv(x,y) and Δz(x,y) are the difference in deformation rate between coherent points x and y and the difference in DEM residual, respectively; B ⊥ For vertical baseline; t j λ is the time baseline; λ is the radar wavelength; R and θ are the radar slant range and incident angle, respectively; This represents the residual phase difference. Since the atmospheric characteristics and nonlinear deformations of nearby coherent points are generally not significantly different, for the solution space search method, we can assume:
[0098]
[0099] At this point, the absolute value of the complex total coherence coefficient can be... As a reliable standard for parameter estimation:
[0100]
[0101] Where M is the number of interferograms, e j This is a plural form. The value range of is [0, 1]. To ensure the accuracy of the solution, it is usually... Arc segments below a certain threshold are discarded; a threshold of 0.7 is recommended. Finally, a stable reference point is selected, and the solution on the arc segment is integrated over all PS monitoring points to obtain the deformation rate and topographic residual results for the monitored area.
[0102] Step 6: Restore the deformation information of the PS monitoring points
[0103] After estimating the model parameters, the deformation sequence of the PS monitoring points needs to be recovered. This involves applying these parameters to the original interferometric phase data, extracting deformation information through phase unwrapping and time series analysis methods. The main content includes removing linear deformation and terrain components from the interferometric phase, on this basis, spatial and temporal filtering of the residual phase is performed to remove atmospheric errors, and finally the linear deformation and nonlinear deformation are combined to obtain the final deformation time series results.
[0104] Temporal and spatial filtering is a commonly used atmospheric correction method. By analyzing the interferometric phase of multiple images, the statistical characteristics of atmospheric errors in time and space are used to estimate and correct them. In addition, external independent data can be used to assist in the removal of atmospheric phase. External meteorological data (such as GNSS, MERIS, MODIS, etc.) can be introduced to calculate atmospheric delay. These external data can provide information about the state of the atmosphere, helping to more accurately estimate and remove atmospheric phase delays in InSAR interferometric phase maps. Advanced techniques such as conditional generative adversarial neural networks can also be used to build atmospheric phase removal models. These methods estimate the size of the phase disturbance by analyzing the propagation path and time delay of the interferometric signal in the atmosphere and make the corresponding correction.
[0105] In summary, through advanced atmospheric phase removal models and InSAR atmospheric correction software, etc. methods, the influence of the atmosphere can be effectively removed, and the deformation information of the PS monitoring points can be recovered.
[0106] Step 7: Use random forest to classify the registered SAR image
[0107] Before image classification, a filter is used to denoise the registered SAR image to reduce the impact of noise on the detection results. Then, for each filtered SAR image, a random forest is used for initial classification, i.e. coarse classification, to divide the image into different categories. Random forest is a supervised machine learning algorithm, and supervised learning is a subcategory of machine learning. This type of learning relies on classification labels to generate a function (model) to identify different categories in the image. There are two types of classification problems, binary classification and multi-classification.
[0108] In order to classify the InSAR deformation, the land classification results of the monitoring area need to be obtained, and the present application uses a random forest to classify the registered SAR image.
[0109] Reference Figure 2 Random forest is a decision tree method based on bootstrap aggregating (Bagging), the principle is as follows: given a data set D of size n, Bagging algorithm uniformly and with replacement (using bootstrap sampling method) selects m subsets Di As training sets. Using classification, regression, etc. algorithm on these m training sets, m models can be obtained, and Bagging results can be obtained by averaging, majority voting, etc. That is:
[0110] Given a data set X = x1,..., xn, the Bagging method repeatedly samples from the training set with replacement k times, and then trains tree models on these samples. After training, the prediction value of unknown sample x is The prediction of all individual regression trees on training sample x' can be achieved by averaging:
[0111]
[0112] Where f represents the prediction model. This Bagging method reduces the variance without increasing the bias, thus bringing better performance. However, simply training multiple tree models on the same data set will produce strongly correlated tree models (even identical tree models). Bagging uses Bootstrap self-sampling method to reduce the correlation between tree models by generating different training sets.
[0113] In addition, the standard deviation σ of the prediction of all individual regression trees on x' can be used as an estimate of the uncertainty of the prediction:
[0114]
[0115] The number of trees is a free parameter, usually several hundred to several thousand trees, depending on the size and nature of the training set. Using cross-validation, or by observing the prediction error, the optimal number of trees can be found.
[0116] In actual random forest classification, first filter the registered SAR image of each scene to reduce the influence of noise on classification. On this basis, based on the intensity image, the sample is labeled manually and empirically. According to the actual demand and the performance of the classifier, the SAR image can be labeled into three categories, such as building area, farmland, grassland, bare land and water area. Use the labeled samples and intensity features to train the model, randomly select 70% of the samples as the training set, and the remaining 30% of the labeled samples as the validation set to verify the model accuracy. Finally, use the trained random forest model to predict and classify the entire monitoring area.
[0117] Step 8: using a classification sample and a ViT network to perform secondary classification on the SAR image
[0118] On the basis of the first classification, the SAR image is classified by the ViT network for the second time, further improving the classification category and accuracy.
[0119] ReferenceFigure 3 The ViT network model divides the input SAR image into a series of image patches and maps each patch into a vector. These vectors represent the input sequence of the ViT model. Next, the ViT model models the relationship between different image blocks in the image through a series of multi-head self-attention layers and fully connected layers, and finally generates a global feature vector for image multi-target classification. Similar to the CNN convolution architecture, the ViT network structure can also be reconstructed layer by layer, including data preprocessing layers, self-attention layers, and multilayer perceptron (MLP) layers.
[0120] Data preprocessing layer: first input an m x m image block into the model, and evenly divide the image block into N 2 m×m Specifically, when the value of N is equal to m, each partition of the original image block will only contain one pixel, then a learnable matrix is used for linear projection, and the matrix product is called pixel embedding. The above operation is consistent with the fully connected layer of the neural network, except that it lacks a nonlinear activation function, which can be regarded as a simple feature extraction. After obtaining the pixel embedding, a randomly initialized vector called class embedding is prepared, and finally the two are stacked to form a new pixel embedding as the input of the next layer, whose size is where 1 represents the class embedding, m 2 is the old pixel embedding. During the learning process, the class embedding is used to integrate the information of the remaining parts, and finally represents the original SAR image block;
[0121] Self-attention layer: similar to the role of convolution in CNN neural networks, ViT features are mainly constructed through self-attention mechanisms. By learning the importance of each pixel relative to other pixels, the interaction between them is understood.
[0122] First, use a learnable matrix to linearly map the pixel embedding to three variables of the same dimension (q, k, v). Then, by obtaining and using attention to extract information and construct feature expression, in this process, based on the pairwise similarity between two sequence elements, a weight matrix A can be generated, and the ith element of A is written as
[0123]
[0124] Here N' = 1 + m 2 , The similarity is obtained by calculating the dot product of q and k, which is related to the cosine similarity, and a scaling factor is used in actual calculation And use Softmax activation to map the output to (0, 1) to generate weights, the above equation determines the importance of the ith position by evaluating q i And k correlation, the output SA is obtained by:
[0125]
[0126] Where the ith element of SA is the weighted sum of v and the importance of the ith position, this operation releases the attention obtained from (10) and constructs a new feature representation based on the mapped pixel embedding; similar to the concept of kernel number in convolution, self-attention can also be calculated multiple times to improve model capacity, which is called multi-head self-attention mechanism.
[0127] MLP layer: including two fully connected layers and a GeLU activation function, used to receive the output from the attention layer, at this stage, the model will learn to map the features to the class probability space and further process it, input the Softmax classifier to generate the final image classification prediction.
[0128] In actual processing, first input the SAR image block into the ViT network backbone to construct the feature expression, and then use the Softmax classifier with cross-entropy loss function for supervised classification. The label of this supervised classification comes from the classification result produced by the random forest, through the above processing, the SAR image can be further divided into four categories, such as building area, bare land area, farmland and grassland area and water area, the secondary classification effectively improves the accuracy of the first classification.
[0129] Step 9: classification, statistics and cause analysis of deformation based on multi-classification results
[0130] After classifying the deformation, statistics of each type of deformation need to be conducted. This includes the frequency, degree and object types involved in each type of deformation. Statistical tables, pictographic statistics and other forms can be used to present the results of classification statistics, so as to more intuitively understand the distribution characteristics and internal relations of deformation.
[0131] By superimposing the deformation data on the land multi-classification results, it can be determined through comparative analysis which land use types correspond to the deformation area, so as to count the area of different deformation types, analyze the frequency and distribution of different deformation types in the region, and understand the spatial distribution characteristics of deformation. In addition, quantitative analysis can also be performed on the deformation displacement time series data set to count parameters such as the average deformation rate and the maximum deformation magnitude of different deformation types, so as to understand the degree of deformation. Combined with various factors such as geological structure, climate change, human activities, etc., the causes of surface deformation can be comprehensively analyzed. For example, climate change and hydrological change are comprehensively analyzed.
[0132] On the basis of classification and statistics of deformation, cause analysis can be carried out. The causes of various types of deformation are analyzed, including the size, direction, duration of external force, as well as the material properties and structural characteristics of the object itself. Through cause analysis, the generation mechanism of deformation can be understood in depth, and scientific basis can be provided for the prevention and control of deformation. For example, if the deformation area is mainly located in farmland area, it can be preliminarily judged that the deformation may be related to agricultural activities; if the deformation area is located in urban construction land, it may be related to urban expansion, underground engineering activities, etc.
[0133] Although embodiments of the present application have been shown and described, it will be understood by those having ordinary skill in the art that various changes, modifications, alternatives and variations can be made thereto without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A ground surface deformation monitoring and classification method based on InSAR and deep learning, characterized in that, The method comprises the following steps: (1) Collecting multi-scene time-series SAR images and DEM data of the monitoring area; (2) Registering, interfering and DEM differentiating the multi-scene time-series SAR images; (3) Selecting PS pixels by using amplitude deviation and coherence coefficient threshold values; (4) Linking PS pixels to construct a spatial network The PS spatial network is constructed by using a TIN, and the Delaunay triangulation method is applied to the spatial distribution of the PS points to form a stable triangular grid structure; (5) Solving model parameters of the spatial network arc segments On PS-space network, the phase difference of the arc segment between two adjacent coherent points in the interferogram is calculated to the absolute value of the complex total coherent coefficient As parameter estimation, a stable reference point is selected to de-integrate the arc segment to all PS monitoring points, obtaining the deformation rate and terrain residual results of the monitoring area. (6) Removing the influence of atmospheric errors from the residual phase to restore the deformation information of the PS monitoring points; (7) Using a random forest to perform primary classification on the registered SAR images Each registered SAR image is filtered to reduce noise, and samples are labeled based on intensity images according to human experience. According to actual requirements and the performance of the classifier, the SAR images are labeled and classified. The labeled samples and intensity features are used for model training. A part of the samples are randomly selected as a training set, and the remaining labeled samples are used as a verification set to verify the accuracy of the model. The trained random forest model is used to predict and classify the entire monitoring area; (8) Using primary classification samples and a ViT network to perform secondary classification on the SAR images The SAR image blocks are input into the ViT network backbone to construct feature expression, and then a Softmax classifier with a cross-entropy loss function is used for supervised classification. The label of the supervised classification is derived from the primary classification result of the random forest; (9) Classifying, counting and analyzing the causes of deformation based on the multi-classification results.
2. The InSAR and deep learning-based ground surface deformation monitoring and classification method according to claim 1, characterized in that, In step (1), SAR image data of the required area is downloaded from a SAR satellite platform. The obtained multi-scene SAR image data is arranged in time sequence to ensure that each image has clear time stamp and geographic location information. Then, the SAR image data is preprocessed by radiation correction, geometric correction and multi-view processing.
3. The InSAR and deep learning-based ground surface deformation monitoring and classification method according to claim 1, characterized in that, In step (1), the obtained DEM data is arranged according to the region, and then the format of the DEM data is converted and the accuracy is evaluated to ensure that it matches the geographic information of the SAR images.
4. The InSAR and deep learning-based ground surface deformation monitoring and classification method according to claim 1, characterized in that, The specific steps of step (2) are as follows: Image registration: The registration of the SAR images includes coarse registration and fine registration. The coarse registration uses satellite orbit parameters, InSAR geometry and SAR positioning equations to obtain the homonymic points of the reference image center point in the image to be registered. The fine registration is based on local pixel offset estimation of the cross-correlation function to achieve higher registration accuracy. Finally, the image to be registered is subjected to coordinate transformation and pixel interpolation resampling; Interferometric processing: The registered images are subjected to interferometric processing to generate an interferogram. The interferogram is subjected to filtering processing to remove the influence of noise and incoherent areas; DEM differentiation: Obtain external digital elevation model data as reference terrain information. The external DEM data is used to remove the influence of terrain phase from the interferogram to obtain a differential interferogram reflecting only the ground deformation.
5. The InSAR and deep learning-based ground surface deformation monitoring and classification method according to claim 1, characterized in that, In step (3), the amplitude deviation and coherence coefficient double threshold method is used to select PS pixels based on SAR image registration, Under high signal-to-noise ratio conditions, the phase dispersion σ φ is approximated by the amplitude dispersion D A which is expressed by the following equation: where m A and σ A respectively represent the amplitude time average and amplitude standard deviation of the pixel, and the pixel point amplitude deviation index value is lower than a certain threshold, it is selected as a PS candidate point. Coherence coefficient threshold: the coherence coefficient threshold can further remove water area points and wetlands, ice and snow low coherence points, the complex coherence of two mean values of zero complex signals S1 and S2 is defined as: In the formula, E(x) represents the expected value, which is the coherence amplitude within the sampling window in the distance direction n × azimuth direction m. The maximum likelihood estimate is expressed as: Where (i, j) represents the pixel coordinates, after removing the terrain phase and flat ground phase, the coherence amplitude of each pixel in the selected interferogram is estimated, and the average coherence map is generated by the following formula: where N represents the number of interferograms, represents the coherence amplitude of the i-th interferogram, and the pixels exceeding the set average coherence threshold are selected as PS candidates.
6. The InSAR and deep learning-based ground surface deformation monitoring and classification method according to claim 1, characterized in that, In the step (5): On the PS space network, assume that two adjacent coherent points in the jth interference pattern are x and y, and the phase difference on the arc segment connecting them is is represented as: where, is the wrapping operation; Δv(x, y) and Δz(x, y) are the differences in the deformation rate and DEM residual between the adjacent points x and y, respectively; B ⊥ is the vertical baseline; t j is the time baseline; λ is the radar wavelength; r and θ are the radar slant range and incidence angle, respectively; which represents the difference in the residual phase. Since the atmospheric characteristics, nonlinear deformation, etc. of the adjacent points generally have no significant difference, for the solution space search method, it is assumed that: The absolute value of the complex population coherence coefficient at this time As a reliable criterion for parameter estimation: where M is the number of interferograms, e j is a complex representation, The value range of is [0, 1], finally, a stable reference point is selected to integrate the arc segment on all PS monitoring points to obtain the deformation rate and terrain residual results of the monitoring area.
7. The InSAR and deep learning-based ground surface deformation monitoring and classification method according to claim 1, characterized in that, In the step (6), the linear deformation and terrain component are removed from the interference phase, on this basis, the residual phase is filtered in space and time to remove atmospheric errors, and finally the final deformation time sequence result is obtained by combining the linear deformation and the nonlinear deformation.
8. The InSAR and deep learning-based ground surface deformation monitoring and classification method according to claim 1, characterized in that, In step (7), the random forest is a decision tree method based on the guided clustering algorithm. Given a dataset X = x1, ..., xn, the Bagging method repeatedly samples from the training set with replacement k times, and then trains the tree model on these samples. After training, the predicted value for the unknown sample x is obtained. This is achieved by averaging the predictions of all individual regression trees on the training sample x′: Where f represents the prediction model, The standard deviation σ of the predictions of all individual regression trees on x' as an estimate of the uncertainty of the prediction:
9. The InSAR and deep learning-based ground surface deformation monitoring and classification method according to claim 1, characterized in that, In the step (8), the ViT network structure includes a data preprocessing layer, a self-attention layer and a multi-perception layer, Data pre-processing layer: Firstly, an m x m image block is input into the model, and the image block is evenly divided into N 2 parts, each part is expanded according to the spatial dimension to form a (m / N)*(m / N) dimensional vector, all vectors are stacked, and the input will change from R m×m to Specifically, when the value of N is equal to m, each partition of the original image block will only contain one pixel, then a learnable matrix is used for linear projection, the matrix product is called pixel embedding, the above operation is consistent with the fully connected layer of the neural network, after obtaining the pixel embedding, a randomly initialized vector is prepared, called class embedding, finally, the two are stacked to form a new pixel embedding as the input of the next layer, the size is where 1 represents the class embedding, m 2 is the old pixel embedding, during the learning process, the class embedding is used to integrate the information of the remaining parts, and finally represents the original SAR image block; Self-attention layer: First, a learnable matrix The pixels are embedded into a linear mapping into three dimensional variables (q, k, v). Then, information is extracted and feature representations are constructed by taking and utilizing attention. In this process, a weight matrix A can be generated based on the pairwise similarity between two sequence elements, whose i-th element is written as Here N' = 1 + m 2 , The similarity is obtained by computing the dot product of q and k, in practice a scaling factor is used and the output is mapped to (0, 1) using a Softmax activation to generate weights, the above equation determines the importance of the ith position by evaluating the relevance of q i and k, the output SA is obtained by the following way: Where the i-th element of SA is the weighted sum of the importance of v and the i-th position, this operation releases the attention obtained from (10) and constructs a new feature representation based on the mapped pixel embedding; MLP layer: including two fully connected layers and a GeLU activation function, used to receive the output from the attention layer, at this stage, the model will map the learned features to the class probability space and further process it, input the Softmax classifier to generate the final image classification prediction.
10. The InSAR and deep learning based ground surface deformation monitoring and classification method according to claim 1, characterized in that, In the step (9), the occurrence frequency, degree and object type involved of each type of deformation are counted, and the causes leading to the occurrence of each type of deformation are analyzed, including the size, direction, duration of external force, and the material properties, structural characteristics of the object itself.
Citation Information
Patent Citations
Method for detecting target changes of SAR images based on direction information measure
CN101634705A
Water body extraction method based on pixel-level SAR (synthetic aperture radar) image time sequence similarity analysis
CN103440489A