SAR homogeneous pixel selection method and device based on ViT and polarization decomposition

By performing Freeman-Durden decomposition and ViT network processing on multi-view holistic polarized SAR images, combining slope angle, slope angle and geometric distortion factors, homogeneous pixel points covering mountainous areas with dense vegetation are screened out, solving the problem of low discrimination accuracy in the existing technology and improving the deformation monitoring accuracy and stability of InSAR technology.

CN120491068APending Publication Date: 2025-08-15湖南省地质灾害调查监测所(湖南省地质灾害应急救援技术中心) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510683797.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In mountainous areas covered with dense vegetation, it is difficult for the existing technology to effectively select stable homogeneous pixel points, resulting in low accuracy in InSAR technology in deformation monitoring, especially in the context of high noise, which can easily cause loss and misjudgment of deformation information.

Method used

Using ViT and polarization decomposition methods, Freeman-Durden decomposition is used to decompose multi-view holographic SAR images to obtain surface scattering power, secondary scattering power and bulk scattering power. Combined with slope angle, slope angle and geometric distortion factors, ViT network is used to classify land objects and screen out homogeneous pixel points.

Benefits of technology

It improves the accuracy of discrimination of homogeneous pixel points in mountainous areas covered by dense vegetation, improves the accuracy and stability of InSAR technology in mountainous landslide monitoring, and enhances the accuracy and classification stability of polarization parameter estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491068A_ABST
    Figure CN120491068A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of surface deformation monitoring, and provides an SAR homogeneous pixel selection method and device based on ViT and polarization decomposition, and the method comprises the steps: obtaining a multi-scene complete polarization SAR image of a target region, and carrying out the registration of the multi-scene complete polarization SAR image; freeman-Durden decomposition is carried out on the complete polarization SAR image obtained after registration is carried out on the complete polarization SAR image obtained after registration is carried out on the complete polarization SAR image; based on a decomposition result, obtaining a time sequence average Freeman-Durden three-component scattering feature; the time sequence average Freeman-Durden three-component scattering features are input into a ViT network to be processed, and a ground feature classification image of the target area is obtained; calculating a slope angle, a slope direction angle and a geometric distortion factor of each pixel point in the main image; and determining homogeneous pixel points in the ground feature classification image according to the slope angle, the slope direction angle and the geometric distortion factor. The method can improve the discrimination accuracy of the homogeneous pixel points in the dense vegetation covered mountainous area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of surface deformation monitoring, and in particular relates to a SAR homogeneous pixel selection method and device based on ViT and polarization decomposition. Background Art

[0002] Interferometric Synthetic Aperture Radar (InSAR) technology, by analyzing time-series SAR imagery, can obtain long-term, high-spatial-resolution surface deformation data for a study area, providing crucial data support for the interpretation, analysis, and prevention of related geological hazards. Compared to areas with low vegetation cover and relatively flat terrain, mountainous areas, influenced by their unique climate and environmental conditions, feature diverse vegetation types and highly undulating terrain. The decoherence caused by dense vegetation cover severely limits the applicability of traditional InSAR technology in high-noise, high-dynamic environments. This makes it difficult to maintain monitoring coverage and accuracy, a major bottleneck hindering the further application of InSAR in landslide monitoring in mountainous areas. Therefore, efficiently extracting stable pixels from the interferogram is crucial for determining the effectiveness of deformation inversion. Relying solely on single-channel coherence or amplitude filtering, particularly in high-noise environments, can easily lead to loss of deformation information and misinterpretation. With the introduction of the SqueeSAR method, "homogeneous pixel selection" has become a core step in InSAR time series processing. Its goal is to use a group of pixels with similar scattering characteristics to effectively suppress coherent speckle noise and local inconsistencies through statistical averaging of spatial neighborhoods, thereby improving the accuracy of information extraction and the stability of deformation estimation. It can also improve the accuracy of polarization parameter estimation and classification stability in PolSAR applications. In recent years, a large number of studies have proposed statistical discrimination methods based on multidimensional features such as amplitude, phase, and scattering mechanism, which significantly enhance the ability to discriminate pixel quality. However, these methods are mostly focused on regular scenes such as cities. For natural mountain landslide areas (i.e., densely vegetated mountainous areas) with diverse landform types and drastic temporal changes, there is still a problem of low discrimination accuracy. Summary of the Invention

[0003] The embodiments of the present application provide a SAR homogeneous pixel selection method and device based on ViT and polarization decomposition, which can solve the problem of poor discrimination accuracy of homogeneous pixels in mountainous areas covered with dense vegetation.

[0004] In a first aspect, an embodiment of the present application provides a SAR homogeneous pixel selection method based on ViT and polarization decomposition, comprising:

[0005] Acquire multi-scene fully polarimetric SAR images of the target area, and register the multi-scene fully polarimetric SAR images to obtain a multi-scene registered fully polarimetric SAR image;

[0006] For each scene of the fully polarimetric SAR image after registration, Freeman-Durden decomposition is performed on the registered fully polarimetric SAR image to obtain the surface scattering power, secondary scattering power and volume scattering power.

[0007] Based on the surface scattered power, secondary scattered power and volume scattered power obtained by Freeman-Durden decomposition, the time-series average Freeman-Durden three-component scattering characteristics are obtained;

[0008] The time-series averaged Freeman-Durden three-component scattering features are input into the ViT network for processing to obtain the object classification image of the target area; the object classification image is used to describe the object type corresponding to each pixel point in the main image during the registration process;

[0009] Calculate the slope angle, aspect angle and geometric distortion factor of each pixel in the main image;

[0010] According to the slope angle, aspect angle and geometric distortion factor, homogeneous pixels in the ground object classification image are determined.

[0011] Optionally, the time-series averaged Freeman-Durden three-component scattering characteristics include average surface scattering power, average secondary scattering power, and average volume scattering power.

[0012] Optionally, based on the surface scattered power, secondary scattered power, and volume scattered power obtained by Freeman-Durden decomposition, a time-series averaged Freeman-Durden three-component scattering characteristic is obtained, including:

[0013] The average surface scattering power of the fully polarized SAR image after multi-scene registration is taken as the average surface scattering power;

[0014] The average secondary scattering power of the fully polarized SAR image after multi-scene registration is taken as the average secondary scattering power;

[0015] The average value of the volume scattering power of the fully polarized SAR image after multi-scene registration is taken as the average volume scattering power.

[0016] Optionally, calculate the slope angle and aspect angle for each pixel in the main image, including:

[0017] The slope angle β of the pixel point (i, j) in the main image is calculated by the following formula: i,j and the slope angle α i,j :

[0018]

[0019]

[0020] DI=(h i-1,j-1 +2h i,j-1 +h i+1,j-1 )-(h i-1,j+1 +2h i,j+1 +h i+1,j+1 )

[0021] DJ=(h i+1,j-1 +2h i+1,j +h i+1,j+1 )-(h i-1,j-1 +2h i-1,j +h i-1,j+1 )

[0022] Among them, h i-1,j-1 Indicates the elevation of the pixel point (i-1, j-1) to the upper left of the pixel point (i, j), h i,j-1 Indicates the elevation of the adjacent pixel point (i, j-1) above the pixel point (i, j), h i+1,j-1 Indicates the elevation of the pixel point (i+1, j-1) to the upper right of the pixel point (i, j), h i-1,j+1 Indicates the elevation of the pixel point (i-1, j+1) to the lower left of the pixel point (i, j), h i,j+1 Indicates the elevation of the adjacent pixel point (i, j+1) below the pixel point (i, j), h i+1,j+1 Indicates the elevation of the pixel point (i+1, j+1) to the lower right of the pixel point (i, j), h i-1,j Indicates the elevation of the adjacent pixel point (i-1, j) to the left of the pixel point (i, j), h i+1,j It represents the elevation of the adjacent pixel point (i+1, j) to the right of the pixel point (i, j), and S represents the area of the pixel point (i, j).

[0023] Optionally, calculate the geometric distortion factor for each pixel in the main image, including:

[0024] The geometric distortion factor R of the pixel point (i, j) in the main image is calculated by the following formula: i,j :

[0025] R i,j = sin[θ-β i,j sin(A i,j )]

[0026]

[0027] Wherein, θ represents the radar incident angle when the main image is captured, γ represents the radar azimuth when the main image is captured, Descending represents ascending orbit shooting, and Ascending represents descending orbit shooting.

[0028] Optionally, determining homogeneous pixels in the ground feature classification image based on the slope angle, aspect angle, and geometric distortion factor includes:

[0029] Use sliding windows to divide the object classification image into multiple rectangular windows;

[0030] For each rectangular window, perform the following steps:

[0031] For each pixel point in the rectangular window, the target pixel points whose slope angle and aspect angle meet the preset regional threshold conditions and whose geometric distortion factor is the same as the geometric distortion factor of the pixel point are screened out from the rectangular window, and the screened target pixel points are regarded as the homogeneous pixel points of the pixel points.

[0032] Optionally, the sliding window size is 15×15.

[0033] In a second aspect, an embodiment of the present application provides a SAR homogeneous pixel selection device based on ViT and polarization decomposition, comprising:

[0034] An acquisition module is used to acquire multi-scene full-polarimetric SAR images of the target area and register the multi-scene full-polarimetric SAR images to obtain a multi-scene registered full-polarimetric SAR image;

[0035] A decomposition module is used to perform Freeman-Durden decomposition on each scene of the registered fully polarimetric SAR image to obtain surface scattering power, secondary scattering power and volume scattering power;

[0036] An averaging module is used to obtain the time-series averaged Freeman-Durden three-component scattering characteristics based on the surface scattered power, secondary scattered power, and volume scattered power obtained by Freeman-Durden decomposition;

[0037] The processing module is used to input the time-series average Freeman-Durden three-component scattering characteristics into the ViT network for processing to obtain the object classification image of the target area; the object classification image is used to describe the object type corresponding to each pixel point in the main image during the registration process;

[0038] A calculation module is used to calculate the slope angle, aspect angle and geometric distortion factor of each pixel in the main image;

[0039] The determination module is used to determine homogeneous pixel points in the ground object classification image according to the slope angle, aspect angle and geometric distortion factor.

[0040] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the above-mentioned SAR homogeneous pixel selection method based on ViT and polarization decomposition is implemented.

[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned SAR homogeneous pixel selection method based on ViT and polarization decomposition.

[0042] The above solution of the present application has the following beneficial effects:

[0043] In the embodiment of the present application, the full polarization SAR image is polarized by using the Freeman-Durden decomposition method, and the parameter information obtained by polarization decomposition is processed by the pixel-level ViT network, and finally the homogeneous pixel points of the ground feature classification image output by the ViT network are selected based on the slope angle, aspect angle and geometric distortion factor of each pixel point in the main image. Among them, due to the polarization decomposition characteristics compared to the single polarization intensity image, it can provide different polarization channel information, and Freeman-Durden is better than the traditional one based on The decomposed classification method is more suitable for describing the scattering mechanism of ground objects in complex mountainous scenes. The ViT network relies on the self-attention mechanism to capture the global dependencies of the image, thereby effectively modeling the complex distribution characteristics, texture structure and spatial relationships of ground objects in remote sensing images. Therefore, the homogeneous pixel selection method of this application can effectively improve the accuracy of distinguishing homogeneous pixels in mountainous areas with dense vegetation coverage.

[0044] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0046] Figure 1 A flowchart of a SAR homogeneous pixel selection method based on ViT and polarization decomposition provided in one embodiment of the present application;

[0047] Figure 2 A schematic diagram of the classification results output by the ViT network in an example;

[0048] Figure 3is a schematic diagram of fully polarimetric SAR data after registration in an example;

[0049] Figure 4 A schematic diagram of InSAR phase optimization results obtained using different methods in an example;

[0050] Figure 5 A schematic diagram of the results of homogeneous pixel selection under different numbers of data sets in an example;

[0051] Figure 6 A schematic structural diagram of a SAR homogeneous pixel selection device based on ViT and polarization decomposition provided in one embodiment of the present application;

[0052] Figure 7 A schematic diagram of the structure of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0053] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0054] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0055] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0056] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0057] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0058] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0059] In view of the problem of low accuracy in identifying homogeneous pixels in natural mountain landslide areas (generally referring to mountainous areas covered with dense vegetation) under the background of diverse landform types and drastic temporal changes, the embodiment of the present application provides a SAR homogeneous pixel selection method based on ViT and polarization decomposition. The method uses the Freeman-Durden decomposition method to polarize the fully polarized SAR image, and uses the pixel-level ViT network to process the parameter information obtained by polarization decomposition. Finally, based on the slope angle, aspect angle and geometric distortion factor of each pixel in the main image, homogeneous pixels are selected for the land feature classification image output by the ViT network. Among them, due to the polarization decomposition characteristics compared to single polarization intensity images, it can provide different polarization channel information, and Freeman-Durden is better than traditional images based on polarization decomposition. The decomposed classification method is more suitable for describing the scattering mechanism of ground objects in complex mountainous scenes. The ViT network relies on the self-attention mechanism to capture the global dependencies of the image, thereby effectively modeling the complex distribution characteristics, texture structure and spatial relationships of ground objects in remote sensing images. Therefore, the homogeneous pixel selection method of this application can effectively improve the accuracy of distinguishing homogeneous pixels in mountainous areas with dense vegetation coverage.

[0060] The SAR homogeneous pixel selection method based on ViT and polarization decomposition provided by the present application is exemplarily described below with reference to specific embodiments.

[0061] like Figure 1 As shown, the SAR homogeneous pixel selection method based on ViT and polarization decomposition provided in the embodiment of the present application includes the following steps:

[0062] Step 11: Acquire multi-scene fully polarimetric SAR images of the target area, and register the multi-scene fully polarimetric SAR images to obtain a multi-scene registered fully polarimetric SAR image.

[0063] The target areas mentioned above are those requiring surface deformation monitoring, such as mountainous areas covered by dense vegetation. Specifically, fully polarimetric SAR imagery can be acquired using satellites (such as the ALOS-2 satellite). For surface deformation monitoring, registering multi-view fully polarimetric SAR images is a key step in ensuring pixel alignment between images. Common SAR image registration methods can be used to register multi-view fully polarimetric SAR images, resulting in a fully polarimetric SAR image after multi-view registration.

[0064] In some optional examples, the above-mentioned full-polarization SAR images can be 9 scenes.

[0065] Step 12: For each scene of the registered fully polarimetric SAR image, perform Freeman-Durden decomposition on the registered fully polarimetric SAR image to obtain surface scattering power, secondary scattering power, and volume scattering power.

[0066] The Freeman-Durden decomposition is a SAR polarimetric decomposition method based on physical scattering mechanisms. It decomposes the polarimetric covariance matrix into a combination of three typical scattering mechanisms: surface scattering, double scattering (i.e., secondary scattering), and volume scattering. The contributions of these three scattering mechanisms to the total scattered power P can be expressed as:

[0067] P=P s +P d +P v ≡(|S hh | 2 +2|S hv | 2 +|S vv | 2 )

[0068] P s =f s (1+|β| 2 )

[0069] P d =f d (1+|α| 2 )

[0070]

[0071] Among them, P s 、P d and P v They represent surface scattering power, secondary scattering power and volume scattering power respectively, S hh The complex value representing the backscattering characteristics of the radar in the horizontal transmission-horizontal reception (HH) polarization mode, S hv The complex value representing the backscattering characteristics of the radar in the horizontal transmission-vertical reception (HV) polarization mode, Svv The complex value representing the backscattering characteristics of the radar in vertical transmit-vertical receive (VV) polarization mode, f s 、f d 、f v Represents the basic scattering intensity of the three scattering mechanisms (f s Represents the basic scattering intensity of surface scattering, f d Represents the basic scattering intensity of secondary scattering, f v represents the basic scattering intensity of volume scattering, while α and β are related to the reflection direction and target structure and are parameters estimated from the polarized echo signal. The advantage of this decomposition method is that it can be used to preliminarily determine the dominant backscattering mechanism in polarimetric SAR data, thereby helping to identify the dominant scattering type. This characteristic makes it commonly used for identifying ground object types and determining surface cover conditions.

[0072] In some embodiments of the present application, the Freeman-Durden decomposition method can be used to perform a plan decomposition on the registered full-polarization SAR image to obtain the surface scattering power P of the registered full-polarization SAR image. s , secondary scattering power P d and volume scattered power P v .

[0073] Step 13: Based on the surface scattered power, secondary scattered power, and volume scattered power obtained by Freeman-Durden decomposition, a time-series averaged Freeman-Durden three-component scattering characteristic is obtained.

[0074] The above-mentioned time-series averaged Freeman-Durden three-component scattering characteristics include average surface scattering power, average secondary scattering power and average volume scattering power.

[0075] In some embodiments of the present application, the average value of the surface scattering power of the fully polarimetric SAR image after multi-scene registration can be used as the average surface scattering power; the average value of the secondary scattering power of the fully polarimetric SAR image after multi-scene registration can be used as the average secondary scattering power; and the average value of the volume scattering power of the fully polarimetric SAR image after multi-scene registration can be used as the average volume scattering power.

[0076] Step 14: Input the time series average Freeman-Durden three-component scattering features into the ViT network for processing to obtain a ground object classification image of the target area.

[0077] The above-mentioned object classification image is used to describe the object type corresponding to each pixel point in the main image during the registration process. In some embodiments of the present application, the object types that may be included in the object classification image include: cultivated land, bare land, woodland, water bodies, buildings, etc.

[0078] In some embodiments of the present application, the above-mentioned ViT network is a computer vision model based on the Transformer architecture, which is mainly used in the present application to process the time-series averaged Freeman-Durden three-component scattering characteristics to obtain the classification of homogeneous objects in mountainous areas. In the present application, the homogeneous object samples are mainly established based on three conditions: ① For vegetation-covered areas, the sample library is annotated based on the scattering characteristics of the objects and the vegetation coverage rate; ② Based on the SAR imaging characteristics of mountainous areas, in addition to the secondary scattering characteristics presented by buildings, a homogeneous sample set of strong scattering objects such as ridge lines is established; ③ For obviously different object categories, sample data sets can be produced using Pauli pseudo-color images and optical images. Through ①②③, a homogeneous pixel selection sample library suitable for vegetation-covered mountainous areas is constructed, the time-series averaged Freeman-Durden three-component scattering feature dataset is integrated with the deep representation capability of the ViT network, and a spatial consistency constraint mechanism is introduced to achieve accurate selection and classification pre-modeling of homogeneous objects.

[0079] Specifically, the overall architecture of the ViT-based polarimetric synthetic aperture radar (PolSAR) image classification system consists of a data preprocessing layer, a self-attention mechanism layer, and an MLP layer. First, image patches are extracted from an image represented by the pixel-centered temporal average Freeman-Durden three-component scattering feature and fed into the ViT backbone network to construct a feature representation. Finally, supervised classification is performed using a softmax classifier combined with a cross-entropy loss function to determine the feature type. The MLP structure, together with the self-attention module, forms the basic unit of the Transformer encoder. This structure further enhances the model's discriminative capabilities after constructing a global contextual representation, effectively improving the classification accuracy of polarimetric SAR images.

[0080] For ease of understanding, the processing of each image patch is taken as an example to introduce in detail the specific process of the present invention in combining the ViT network with the Freeman-Durden decomposition polarization characteristics to achieve homogeneous object classification.

[0081] Assume the input image is Where H and W are the height and width of the image (the image represented by the temporal average Freeman-Durden three-component scattering characteristics), respectively, and C is the number of channels, that is, the average scattering matrix obtained by temporal averaging of Freeman-Durden decomposition (the temporal average Freeman-Durden three-component scattering characteristics). First, the average scattering matrix is divided into N fixed-size image blocks of size P×P, each of which is represented by x i , where i = 1, 2, ... N, where N represents the number of patches obtained by segmenting the image. Each patch is flattened into a vector form:

[0082]

[0083] In the above formula, Flatten represents the flattening operation, P i Represents the i-th patch.

[0084] Then, the flattened x i By linearly projecting it into the embedding space of fixed dimension D, we get z i , W is the weight matrix, and b is the bias vector.

[0085]

[0086] In order to introduce the global semantic information of the image, it is also necessary to predefine a CLS token for classification, which is recorded as a learnable parameter After concatenating the CLS token with the embedding vectors of all patches, the complete input sequence X is formed. i .

[0087]

[0088] In addition, to preserve spatial location information, it is necessary to add a corresponding learnable position code to each token. The i-th row E i It represents the position information of the i-th token, and finally the embedding and position encoding are added element by element to obtain the input sequence of the Transformer module.

[0089] The dependencies between image blocks are captured through a multi-head self-attention mechanism. The specific implementation process is as follows:

[0090] First, input X i Mapped to query matrix Q, key matrix K and value matrix V, respectively, and trainable linear projection parameters d h Represents the dimension of each attention head.

[0091] Q=X i W Q ,K=X i W K ,V=X i W V

[0092] For each attention head, in order to better capture the different relationships and features in the input sequence, the softmax function is used to normalize and calculate the attention weight matrix A, head i Represents a single attention head in the multi-head attention mechanism (Mult i-Head Attention).

[0093]

[0094] head i =Attention(Q i ,K i ,V i )

[0095] In order to improve the model's ability to capture diverse features, the above process can be performed h times in parallel to implement a multi-head attention mechanism, where each head corresponds to an independent set of Q i , K i 、V i The final output of the attention layer is obtained by concatenating the outputs of each head and applying a linear transformation. This self-attention mechanism can effectively capture long-range dependencies between image patches, compensating for the lack of local perception capabilities of traditional convolutional networks, thereby improving the overall performance of PolSAR image classification and assisting in the extraction of homogeneous feature pixels.

[0096] To further enhance the nonlinear modeling and feature expression capabilities of the model, a multi-layer perception mechanism (MLP) module is introduced after each self-attention module. The image blocks of the CLS token embedded with position encoding and classification are sequentially input into the MLP module for calculation. The MLP module consists of two layers of fully connected neural networks. The weight matrix of each layer of the fully connected network is D ff represents the dimension of the module's intermediate layer, b1 and b2 are the corresponding bias vectors, and a nonlinear activation function, GeLU, is embedded to learn complex scattering data distributions. Furthermore, to ensure training stability and improve gradient propagation efficiency, layer normalization is introduced before and after the MLP module, and the input and output are summed via residual connections, effectively preventing gradient vanishing and accelerating model convergence.

[0097] MLP(x i )=GeLU(x i W1+b1)W2+b2

[0098] After completing the self-attention mechanism and multi-layer perception module processing, the feature vector corresponding to the first CLS token is used as the image block x i The global representation of the output feature sequence is denoted as The long-distance dependencies and semantic features between all patches are integrated, which is the input basis for subsequent classification operations. In order to obtain the predicted probability of each category, the logits vector (i.e. the original output of the fully connected layer, the D-dimensional vector without any nonlinear activation function) is input into the softmax function for normalization, where p(c i |x) indicates that the input feature x belongs to category ci The predicted probability of .

[0099]

[0100] z i and z j The elements in the Logits vector representing the output of the model correspond to the original scores of the i-th and j-th categories, respectively. i The known label of the i-th category is the target category to be predicted in the classification task, and D represents the total number of label categories.

[0101] To guide the model training process, the cross entropy loss function is used as the optimization target. Assume that the true label is y∈{1,2,…,D}, then the loss function of the input feature x is Defined as:

[0102]

[0103] p(y|x) represents the probability of the true label y given the input feature x.

[0104] By minimizing the loss function formula Combined with the gradient backpropagation mechanism, the model can continuously update its parameters, thereby improving the prediction accuracy in polarimetric SAR image classification tasks. This classification module, as the final output layer of the ViT network, can accurately classify and identify target areas and establish a homogeneous terrain constraint model for mountainous areas in the InSAR time series phase.

[0105] Step 15: Calculate the slope angle, aspect angle, and geometric distortion factor of each pixel in the main image.

[0106] It should be noted that the geometric relationship of the undulating terrain in mountainous areas greatly affects the quality of SAR imaging. Common geometric distortions will lead to missing or aliasing of ground object information. The polarization scattering responses of the same type of ground objects in different geometric distortion attribute areas are significantly different, and the requirements of "scattering consistency" for homogeneous pixels cannot be met. Therefore, this application introduces slope, aspect and SAR radar imaging characteristics as constraints to achieve homogeneous terrain constraint modeling in mountainous areas. The establishment of conditions mainly considers two aspects: ① For terrain characteristics, the slope and aspect information of the pixel can be calculated according to the digital elevation model and formula, and the terrain homogeneity constraint condition is set according to the threshold. ② For radar imaging principles, geometric distortion factor modeling is introduced. Through ①②, a homogeneous terrain constraint model suitable for vegetation-covered mountainous areas is adaptively constructed.

[0107] Specifically, the slope angle β of the pixel point (i, j) in the main image can be calculated by the following formula: i,j and the slope angle α i,j :

[0108]

[0109] DI=(h i-1,j-1 +2h i,j-1 +h i+1,j-1 )-(h i-1,j+1 +2h i,j+1 +h i+1,j+1 )

[0110] DJ=(h i+1,j-1 +2h i+1,j +h i+1,j+1 )-(h i-1,j-1 +2h i-1,j +h i-1,j+1 )

[0111] Among them, h i-1,j-1 Indicates the elevation of the pixel point (i-1, j-1) to the upper left of the pixel point (i, j), h i,j-1 Indicates the elevation of the adjacent pixel point (i, j-1) above the pixel point (i, j), h i+1,j-1 Indicates the elevation of the pixel point (i+1, j-1) to the upper right of the pixel point (i, j), h i-1,j+1 Indicates the elevation of the pixel point (i-1, j+1) to the lower left of the pixel point (i, j), h i,j+1 Indicates the elevation of the adjacent pixel point (i, j+1) below the pixel point (i, j), h i+1,j+1 Indicates the elevation of the pixel point (i+1, j+1) to the lower right of the pixel point (i, j), h i-1,j Indicates the elevation of the adjacent pixel point (i-1, j) to the left of the pixel point (i, j), h i+1,jIndicates the elevation of the adjacent pixel (i+1, j) to the right of pixel (i, j), and S indicates the area of pixel (i, j). Pixel (i, j) indicates the pixel with coordinates (i, j), which can be any pixel in the main image. The adjacent pixel (i-1, j-1) indicates the pixel with coordinates (i-1, j-1), the adjacent pixel (i, j-1) indicates the pixel with coordinates (i, j-1), the adjacent pixel (i+1, j-1) indicates the pixel with coordinates (i+1, j-1), and the adjacent pixel (i-1, j+ 1) represents the pixel with coordinate position (i-1, j+1), the adjacent pixel (i, j+1) represents the pixel with coordinate position (i, j+1), the adjacent pixel (i+1, j+1) represents the pixel with coordinate position (i+1, j+1), the adjacent pixel (i-1, j) represents the pixel with coordinate position (i-1, j), and the adjacent pixel (i+1, j) represents the pixel with coordinate position (i+1, j). It can be understood that when calculating the slope angle and aspect angle of the pixel at the edge position, the elevation of the non-existent adjacent pixel is 0. For example, when calculating the slope angle and aspect angle of the leftmost pixel, there are no adjacent pixels in the upper left, left, and lower left directions, so the elevations corresponding to these positions are 0.

[0112] The above formula shows the method of calculating terrain slope and aspect information based on the digital elevation model (DEM). First, the elevation difference Δhi is obtained in the local neighborhood of the pixel (i, j) ,j , get the gradient components DI and DJ of the matrix rows and columns, and then use the inverse tangent function to solve the slope angle β corresponding to each pixel i,j and the slope angle α i,j , where the slope calculation takes into account the normalization factor of the pixel area S, thereby achieving accurate expression of the terrain inclination and orientation of each pixel point in SAR, which is suitable for landform recognition and surface geometric feature modeling.

[0113] In some embodiments of the present application, the geometric distortion factor R of the pixel point (i, j) in the main image can be calculated by the following formula: i,j :

[0114] R i,j = sin[θ-β i,j sin(A i,j )]

[0115]

[0116] Where θ represents the radar incident angle when the main image is captured, γ represents the radar azimuth when the main image is captured, Descending indicates that the image is captured when the satellite is ascending (i.e., the main image is captured when the satellite is ascending), and Ascending indicates that the image is captured when the satellite is descending (i.e., the main image is captured when the satellite is descending).

[0117] The above formula shows the calculation method of the geometric distortion factor based on the terrain and radar imaging information. By using the slope angle β at the pixel (i, j) i,j and the slope angle α i,j As well as the radar incident angle θ and azimuth angle γ, the geometric distortion factor R index size of each pixel is obtained, so that a homogeneous terrain constraint model can be constructed by setting adaptive regional thresholds for the slope angle and aspect angle and the R index satisfies the same geometric distortion conditions.

[0118] Step 16: Determine homogeneous pixel points in the ground feature classification image based on the slope angle, aspect angle, and geometric distortion factor.

[0119] In some embodiments of the present application, the selection of homogeneous pixels in the ground object classification image can be completed by dividing the windows.

[0120] Specifically, a sliding window can be used to divide the object classification image into multiple rectangular windows, and then the following steps are performed for each rectangular window:

[0121] For each pixel point in the rectangular window, the target pixel points whose slope angle and aspect angle meet the preset regional threshold conditions and whose geometric distortion factor is the same as the geometric distortion factor of the pixel point are screened out from the rectangular window, and the screened target pixel points are regarded as the homogeneous pixel points of the pixel points.

[0122] The above-mentioned preset area threshold condition can be: the slope angle and / or the aspect angle is less than 20 degrees. Therefore, for a certain pixel point within a rectangular window, the pixel point with a slope angle and / or the aspect angle within the rectangular window less than 20 degrees and a geometric distortion factor the same as the geometric distortion factor of the pixel point will be regarded as a homogeneous pixel point of the pixel point (a homogeneous pixel point can be understood as a pixel point belonging to the same ground feature type as the pixel point), and a homogeneous pixel set is obtained, thereby completing the selection of the notification pixel point within the rectangular window. In order to improve the accuracy of the selection, the size of the sliding window can be set to 15×15.

[0123] After identifying homogeneous pixels, adaptive weighted filtering calculations can be performed based on commonly used filtering methods to obtain the filtered optimized phase, thereby facilitating the improvement of the accuracy of surface deformation monitoring in surface deformation estimation.

[0124] The SAR homogeneous pixel selection method based on ViT and polarization decomposition provided by the present application is exemplarily described below with reference to specific examples.

[0125] In this example, the experimental dataset used was derived from nine scenes of L-band fully polarimetric data acquired by the ALOS-2 satellite, with a spatial resolution of approximately 3 meters. The data was located in a mountainous area. The region experiences simultaneous rain and heat, high humidity, excellent vegetation growth conditions, and a high forest coverage rate, making it suitable for verifying the adaptability of the present invention in complex natural environments. After radiometric correction, geometric registration, and despeckling, the image data was processed to extract polarimetric scattering features, and interferometry processing was performed on the HH polarization channel to provide standardized input data for model training and classification evaluation of the proposed method.

[0126] The time series average Freeman-Durden decomposition data is used as training data, and the homogeneous samples are used as training samples to input the ViT network for 100 iterations. The classification results (i.e., the ground feature classification image) are as follows: Figure 2 shown.

[0127] In order to compare the advantages of the present invention, the present invention will Figure 3 The fully polarized SAR data after registration shown is used as the data source, and the method of this application and the traditional method are used to perform SAR intensity image filtering and InSAR time series phase estimation respectively. Among them, the traditional method includes: ①DespecKS method: In the time series intensity image, the double-sample KS test is used to calculate the cumulative distribution function of the intensity value of the reference pixel and the remaining pixels in the window in the time dimension. If the difference is not significant at the significance level, it is judged to be a homogeneous pixel, thereby achieving a robust homogeneous pixel set, and adaptive window filtering reduces imaging noise. ②BM3D method: Utilizing the non-local similarity in a single SAR image, image blocks similar to the reference block are searched within the search window according to the Euclidean distance threshold, and these blocks with similar structures are grouped to form block groups. The block groups are then stacked to form a three-dimensional data array, and joint transformation, denoising and inverse transformation reconstruction are performed in this three-dimensional domain, thereby achieving efficient speckle noise suppression and structural information retention. The results of the above experiments are shown in Figure 2. Figure 4 、 Figure 5 As shown in Table 1, Figure 4 The InSAR phase optimization results obtained by different methods are shown. Figure 5 The results of homogeneous pixel selection under different numbers of datasets are shown. Figure 4 (d) ViTPol is the phase optimization result completed based on the homogeneous selection method of this application. Figure 4 The second row of pictures in (a) are enlarged pictures of the rectangular boxes in the first row of pictures. Figure 4 The second row of pictures in (b) are enlarged pictures of the rectangular boxes in the first row of pictures. Figure 4The second row of pictures in (c) are enlarged pictures of the rectangular boxes in the first row of pictures. Figure 4 The second row of pictures in (d) are enlarged pictures of the rectangular boxes in the first row of pictures. Figure 5 Where KSSHPS represents the homogeneous point selection based on the DespecKS method, ViTPolSHP represents the homogeneous point selection completed by the method of this application, and N represents the number of SAR images involved in the calculation. Table 1 records the values of the mean, ENL, and SSI corresponding to the original image, the homogeneous point selection based on the DespecKS method, and the homogeneous point selection completed by the method of this application.

[0128] Table 1

[0129] index Original image DespecKS ViTPolSHP <![CDATA[Mean (*10 7 )]]> 8.210 8.517 8.282 ENL 1.340 1.888 16.302 SSI / 0.551 0.180

[0130] The mean is one of the most commonly used metrics in statistics, used to describe the central tendency of data. The Equivalent Number of Looks (ENL), calculated by calculating the mean and variance of homogeneous regions, measures the degree of speckle suppression. A larger ENL value indicates a stronger speckle suppression effect. The Speckle Suppression Index (SSI) measures the speckle suppression effect of a filtering algorithm relative to the original image. Generally speaking, the SSI is less than 1, and lower SSI values indicate a stronger speckle suppression filter.

[0131] It can be seen that the InSAR homogeneous pixel set extracted by the method of this application is more consistent in intensity and texture features, and also performs well in small data sets. It can effectively suppress noise interference while maintaining phase edge information, showing superior performance in interferometric phase filtering tasks.

[0132] It is worth noting that, assuming that all pixels within a homogeneous region have the same radar reflectivity and geophysical parameters, the signal-to-noise ratio of the central pixel can be improved by averaging its homogeneous neighborhood, thereby enhancing the phase signal of interest. Polarimetric decomposition is a key process for extracting ground scatterers and their corresponding scattering characteristics from SAR images. It aims to reveal the scattering mechanisms of different ground object types, thereby enabling deep mining and effective utilization of polarimetric information. This process is crucial for understanding the physical properties of targets, classifying ground object types, and improving interpretation accuracy. Polarimetric target decomposition decomposes complex radar echo signals into several physically meaningful scattering components, enabling more intuitive analysis of ground object scattering characteristics. Freeman-Durden decomposition effectively distinguishes different scattering mechanisms while preserving physical meaning, providing a solid foundation for subsequent ground object classification. Furthermore, the powerful self-attention mechanism and excellent learning capabilities of the ViT network are leveraged to learn complex features and relationships from multi-temporal polarimetric decomposition feature data to achieve pixel classification. Unlike traditional data-driven homogeneous pixel selection methods, deep learning-based methods significantly reduce the reliance on large datasets, enabling homogeneous filtering on small datasets.

[0133] In summary, using the multi-temporal averaged Freeman-Durden polarimetric decomposition features as the prior classification basis of the ViT network and integrating terrain features with radar geometric distortion information for adaptive modeling can improve the accuracy of distinguishing homogeneous pixels in mountainous areas with dense vegetation coverage, construct a homogeneous pixel set that is more suitable for mountainous scenes, and thus provide key support for the suppression of speckle noise of distributed scatterers in multi-temporal InSAR.

[0134] The following is an illustrative description of the relevant equipment provided in this application in conjunction with specific embodiments.

[0135] like Figure 6 As shown, an embodiment of the present application provides a SAR homogeneous pixel selection device based on ViT and polarization decomposition, and the SAR homogeneous pixel selection device 600 includes:

[0136] An acquisition module 601 is configured to acquire multi-view full-polarimetric SAR images of a target area and register the multi-view full-polarimetric SAR images to obtain a multi-view registered full-polarimetric SAR image.

[0137] A decomposition module 602 is configured to perform Freeman-Durden decomposition on each registered fully polarimetric SAR image to obtain surface scattered power, secondary scattered power, and volume scattered power.

[0138] An averaging module 603 is configured to obtain a time-series average Freeman-Durden three-component scattering signature based on the surface scattering power, secondary scattering power, and volume scattering power obtained by Freeman-Durden decomposition;

[0139] Processing module 604 is used to input the time-series average Freeman-Durden three-component scattering characteristics into the ViT network for processing to obtain a ground object classification image of the target area; the ground object classification image is used to describe the ground object type corresponding to each pixel point in the main image during the registration process;

[0140] The calculation module 605 is used to calculate the slope angle, aspect angle and geometric distortion factor of each pixel in the main image;

[0141] The determination module 606 is configured to determine homogeneous pixels in the ground feature classification image according to the slope angle, the aspect angle, and the geometric distortion factor.

[0142] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0143] like Figure 7 As shown, an embodiment of the present application provides a terminal device, such as Figure 7 As shown, the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 7 Only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the steps of any of the above method embodiments when executing the computer program D102.

[0144] Specifically, when the processor D100 executes the computer program D102, it performs polarization decomposition on the fully polarized SAR image by using the Freeman-Durden decomposition method, and processes the parameter information obtained by the polarization decomposition using the pixel-level ViT network, and finally selects homogeneous pixels from the ground feature classification image output by the ViT network based on the slope angle, aspect angle and geometric distortion factor of each pixel in the main image. Among them, due to the polarization decomposition characteristics compared to single polarization intensity images, it can provide different polarization channel information, and Freeman-Durden is better than traditional polarization-based images. The decomposed classification method is more suitable for describing the scattering mechanism of ground objects in complex mountainous scenes. The ViT network relies on the self-attention mechanism to capture the global dependencies of the image, thereby effectively modeling the complex distribution characteristics, texture structure and spatial relationships of ground objects in remote sensing images. Therefore, the homogeneous pixel selection method of this application can effectively improve the accuracy of distinguishing homogeneous pixels in mountainous areas with dense vegetation coverage.

[0145] The processor D100 may be a central processing unit (CPU), or may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0146] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart memory card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card, etc. equipped on the terminal device D10. Furthermore, the memory D101 may also include both an internal storage unit of the terminal device D10 and an external storage device. The memory D101 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory D101 may also be used to temporarily store data that has been output or is to be output.

[0147] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0148] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0149] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0150] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0151] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the SAR homogeneous pixel selection device / terminal device based on ViT and polarization decomposition, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, mobile hard disk, magnetic disk or optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0152] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0153] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0154] In the embodiments provided in this application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0155] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0156] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A SAR homogeneous pixel selection method based on ViT and polarization decomposition, characterized in that: include: Acquire multi-view full-polarimetric SAR images of the target area, and register the multi-view full-polarimetric SAR images to obtain a multi-view registered full-polarimetric SAR image; For each scene of the registered fully polarimetric SAR image, Freeman-Durden decomposition is performed on the registered fully polarimetric SAR image to obtain surface scattering power, secondary scattering power, and volume scattering power; Based on the surface scattered power, secondary scattered power and volume scattered power obtained by Freeman-Durden decomposition, the time-series average Freeman-Durden three-component scattering characteristics are obtained; Inputting the time-series average Freeman-Durden three-component scattering features into the ViT network for processing, obtaining a ground object classification image of the target area; the ground object classification image is used to describe the ground object type corresponding to each pixel point in the main image during the registration process; Calculating the slope angle, aspect angle, and geometric distortion factor of each pixel in the main image; Homogeneous pixel points in the ground object classification image are determined according to the slope angle, the aspect angle and the geometric distortion factor.

2. The SAR homogeneous pixel selection method according to claim 1, characterized in that: The time-series average Freeman-Durden three-component scattering characteristics include average surface scattering power, average secondary scattering power and average volume scattering power.

3. The SAR homogeneous pixel selection method according to claim 2, characterized in that: The surface scattered power, secondary scattered power and volume scattered power obtained based on the Freeman-Durden decomposition are used to obtain the time-series average Freeman-Durden three-component scattering characteristics, including: The average surface scattering power of the fully polarized SAR image after multi-scene registration is taken as the average surface scattering power; The average secondary scattering power of the fully polarized SAR image after multi-scene registration is taken as the average secondary scattering power; The average value of the volume scattering power of the fully polarized SAR image after multi-scene registration is taken as the average volume scattering power.

4. The SAR homogeneous pixel selection method according to claim 1, characterized in that: Calculating the slope angle and the aspect angle of each pixel in the main image, including: The slope angle β of the pixel point (i, j) in the main image is calculated by the following formula: i,j and the slope angle α i,j : <h2 style=";text-align:left;direction:ltr">DI=(h<h2 style=";text-align:left;direction:ltr"> i-1,j-1 <h2 style=";text-align:left;direction:ltr"> +2h<h2 style=";text-align:left;direction:ltr"> i,j-1 <h2 style=";text-align:left;direction:ltr"> +h<h2 style=";text-align:left;direction:ltr"> i+1,j-1 <h2 style=";text-align:left;direction:ltr"> )-(h<h2 style=";text-align:left;direction:ltr"> i-1,j+1 <h2 style=";text-align:left;direction:ltr"> +2h<h2 style=";text-align:left;direction:ltr"> i,j+1 <h2 style=";text-align:left;direction:ltr"> +h<h2 style=";text-align:left;direction:ltr"> i+1,j+1 <h2 style=";text-align:left;direction:ltr"> ) DJ=(h i+1,j-1 +2h i+1,j +h i+1,j+1 )-(h i-1,j-1 +2h i-1,j +h i-1,j+1 ) Among them, h i-1,j-1 Indicates the elevation of the pixel point (i-1, j-1) to the upper left of the pixel point (i, j), h i,j-1 Indicates the elevation of the adjacent pixel point (i, j-1) above the pixel point (i, j), h i+1,j-1 Indicates the elevation of the pixel point (i+1, j-1) to the upper right of the pixel point (i, j), h i-1,j+1 Indicates the elevation of the pixel point (i-1, j+1) to the lower left of the pixel point (i, h), h i,j+1 Indicates the elevation of the adjacent pixel point (i, j+1) below the pixel point (i, h), h i+1,j+1 Indicates the elevation of the pixel point (i+1, j+1) to the lower right of the pixel point (i, h), h i-1,j Indicates the elevation of the adjacent pixel point (i-1, j) to the left of the pixel point (i, j), h i+1,j It represents the elevation of the adjacent pixel point (i+1, j) to the right of the pixel point (i, j), and S represents the area of the pixel point (i, j).

5. The SAR homogeneous pixel selection method according to claim 4, characterized in that: Calculating the geometric distortion factor of each pixel in the main image, including: The geometric distortion factor R of the pixel point (i, j) in the main image is calculated by the following formula: i,j : R i,j =sin[θ-β i,j sin(A i,j )] Wherein, θ represents the radar incident angle when the main image is captured, γ represents the radar azimuth when the main image is captured, Descending represents ascending trajectory shooting, and Ascending represents descending trajectory shooting.

6. The SAR homogeneous pixel selection method according to claim 5, characterized in that: Determining homogeneous pixel points in the ground object classification image according to the slope angle, the aspect angle, and the geometric distortion factor includes: Dividing the ground object classification image into a plurality of rectangular windows using a sliding window; For each rectangular window, perform the following steps: For each pixel point in the rectangular window, target pixel points whose slope angle and aspect angle meet the preset area threshold conditions and whose geometric distortion factor is the same as the geometric distortion factor of the pixel point are filtered out from the rectangular window, and the filtered target pixel points are used as homogeneous pixel points of the pixel point.

7. The SAR homogeneous pixel selection method according to claim 6, characterized in that: The size of the sliding window is 15×15.

8. A SAR homogeneous pixel selection device based on ViT and polarization decomposition, characterized in that: include: An acquisition module is used to acquire multi-scene full-polarimetric SAR images of the target area and register the multi-scene full-polarimetric SAR images to obtain a multi-scene registered full-polarimetric SAR image; a decomposition module for performing Freeman-Durden decomposition on each registered fully polarimetric SAR image to obtain surface scattering power, secondary scattering power, and volume scattering power; An averaging module is used to obtain the time-series averaged Freeman-Durden three-component scattering characteristics based on the surface scattered power, secondary scattered power, and volume scattered power obtained by Freeman-Durden decomposition; a processing module, configured to input the time-series averaged Freeman-Durden three-component scattering feature into a ViT network for processing to obtain a ground object classification image of the target area; the ground object classification image is used to describe the ground object type corresponding to each pixel point in the main image during the registration process; A calculation module, configured to calculate the slope angle, aspect angle, and geometric distortion factor of each pixel point in the main image; The determination module is used to determine homogeneous pixel points in the ground feature classification image according to the slope angle, the aspect angle and the geometric distortion factor.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the SAR homogeneous pixel selection method based on ViT and polarization decomposition is implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the SAR homogeneous pixel selection method based on ViT and polarization decomposition is implemented as claimed in any one of claims 1 to 7.