Field extraction method based on PolSAR multi-class segmentation multi-task network

Through the PolSAR multi-category segmentation multi-task network, combined with polarization decomposition and hybrid framework, the problem of insufficient accuracy and adaptability of optical remote sensing data in cloudy and rainy areas is solved, and efficient farmland plot extraction and all-weather monitoring are achieved.

CN120356045APending Publication Date: 2025-07-22CENT SOUTH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510382319.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art has poor accuracy and adaptability in farmland plots based on optical remote sensing data in cloudy and rainy areas, making it difficult to achieve efficient arable land monitoring.

Method used

The method of multi-category segmentation multi-task network is adopted based on PolSAR, and the feature matrix is constructed through polarization decomposition theory, combined with convolutional neural network and Transformer to build a hybrid framework, integrating region segmentation, boundary detection and pixel-to-boundary distance field regression tasks, and using multi-task loss function for end-to-end optimization to realize field extraction.

Benefits of technology

It improves the accuracy and adaptability of farmland plot extraction, breaks through the limitation of the lack of optical data in cloudy and rainy areas, realizes stable all-weather arable land monitoring, and improves the generalization and extraction accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356045A_ABST
    Figure CN120356045A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a field extraction method based on a PolSAR multi-class segmentation multi-task network, and belongs to the technical field of data processing, and the method specifically comprises the steps: 1, extracting a polarization parameter set corresponding to polarization SAR data based on a polarization decomposition theory, and constructing a feature matrix according to the polarization parameter set; step 2, constructing a hybrid framework based on a convolutional neural network and Transform, and performing fusion weighting on local features and global features corresponding to the feature matrix according to the hybrid framework; step 3, establishing a multi-task joint learning model, integrating three complementary tasks of region segmentation, boundary detection and pixel-to-boundary distance field regression, and performing end-to-end optimization by adopting a multi-task loss function and a feature matrix after fusion weighting to obtain an extraction model; and 4, inputting the polarimetric SAR data corresponding to the target area into the extraction model to obtain a field extraction result. Through the scheme disclosed by the invention, the extraction accuracy and adaptability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of data processing, and in particular, to a method for extracting fields based on a PolSAR multi-class segmentation multi-task network. Background Art

[0002] Currently, for the extraction requirements of complex scenarios, the multi-task collaborative learning framework is becoming a new trend in technological evolution. This paradigm realizes the complementary enhancement of feature representation and the joint constraint of error propagation by jointly optimizing semantically related sub-tasks (such as plot range segmentation, boundary localization, distance field calculation, etc.). Representative achievements include: The ResUNet-a model innovatively constructs a four-task joint optimization network, and simultaneously realizes the generation of arable land masks, boundary detection, pixel-level distance field regression, and image reconstruction in an improved U-Net architecture. Its multi-scale feature interaction mechanism effectively improves the recognition robustness of small-scale plots; BSiNet uses a dual-branch decoder to realize the decoupled learning of regional features and boundary features; SEANet introduces a boundary attention module to enhance the contour perception ability; CTMENet fuses the CNN and Transformer architectures and uses the self-attention mechanism to capture the global spatial correlation of farmland arrangements; The CLPs model designs a boundary weighted fusion strategy to solve the fracture problem under complex terrain; The PLR-Net constructs a point-line-plane joint learning framework and realizes sub-pixel level accuracy improvement through triple constraints of vertex detection-boundary tracking-region segmentation.

[0003] It should be particularly emphasized that current research mostly focuses on optical remote sensing data, and there has been no exploration of multi-task deep learning for farmland plot extraction based on synthetic aperture radar (SAR) data. However, in areas with high cloud and rain coverage, the availability of optical images is severely restricted.

[0004] Obviously, there is an urgent need for a method for extracting fields based on a PolSAR multi-class segmentation multi-task network with high extraction accuracy and adaptability. Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide a method for extracting fields based on a PolSAR multi-class segmentation multi-task network, which at least partially solves the problem of poor extraction accuracy and adaptability in the prior art.

[0006] Embodiments of the present disclosure provide a method for extracting fields based on a PolSAR multi-class segmentation multi-task network, including:

[0007] Step 1, extracting a set of polarization parameters corresponding to the polarimetric SAR data based on the polarization decomposition theory and constructing a feature matrix accordingly;

[0008] Step 2: Construct a hybrid framework based on a convolutional neural network and Transformer, and accordingly perform fusion weighting on the local features and global features corresponding to the feature matrix;

[0009] Step 3: Establish a multi-task joint learning model, integrate three complementary tasks of region segmentation, boundary detection, and pixel-to-boundary distance field regression, and perform end-to-end optimization using a multi-task loss function and the fusion-weighted feature matrix to obtain an extraction model;

[0010] Step 4: Input the polarimetric SAR data corresponding to the target region into the extraction model to obtain the result of field block extraction.

[0011] According to a specific implementation manner of an embodiment of the present disclosure, the specific steps of Step 1 include:

[0012] Step 1.1: Decompose the polarimetric SAR data through the Cloude-Pottier decomposition theory to generate the polarimetric entropy H, average scattering angle α, and anisotropy parameter A, and decompose the polarimetric SAR data into surface scattering P s , dihedral angle scattering P d , and volume scattering power ratio P v ;

[0013] Step 1.2: Based on the polarimetric entropy H, average scattering angle α, anisotropy parameter A, surface scattering P s , dihedral angle scattering P d , and volume scattering power ratio P v , construct a three-dimensional feature system through a hierarchical feature fusion strategy, and accordingly construct a feature matrix.

[0014] According to a specific implementation manner of an embodiment of the present disclosure, the three-dimensional feature system includes a basic layer, an energy layer, and a cross-derived layer.

[0015] According to a specific implementation manner of an embodiment of the present disclosure, the step of constructing a three-dimensional feature system based on the polarimetric entropy H, average scattering angle α, anisotropy parameter A, surface scattering P s , dihedral angle scattering P d , and volume scattering power ratio P v , through a hierarchical feature fusion strategy, includes:

[0016] Step 1.2.1: Use the polarimetric entropy H, average scattering angle α, and anisotropy parameter A as the basic layer to analyze the scattering mechanism characteristics;

[0017] Step 1.2.2: Use the surface scattering P s , dihedral angle scattering P d , and volume scattering power ratio P vAs an energy layer, it is used to quantify the scattered power distribution;

[0018] Step 1.2.3, based on polarization entropy H, average scattering angle α, anisotropy parameter A, surface scattering P s , dihedral angle scattering P d and the proportion of volume scattering power P v as a cross-derived layer, and generate high-order features through non-linear combination.

[0019] According to a specific implementation manner of the embodiment of the present disclosure, the specific steps of step 2 include:

[0020] Step 2.1, use a depthwise separable convolutional neural network to construct a five-level feature pyramid, extract multi-scale local texture features in the feature matrix, and connect a CBAM attention module after each convolutional layer to strengthen spatial-channel feature screening;

[0021] Step 2.2, when the feature matrix passes through each level of the pyramid layer, it is sequentially divided into image blocks of a preset size, linearly mapped to generate a serialized input Transformer encoder, and axial attention is designed in the Transformer encoder to reduce the computational complexity, and at the same time, a deformable convolutional kernel is introduced to replace the standard position encoding;

[0022] Step 2.3, use gated cross-attention to dynamically fuse the detailed features of the depthwise separable convolutional neural network and the global semantics of the Transformer, where the features output by the depthwise separable convolutional neural network are used as key vectors and value vectors, and the output of the Transformer encoder is used as a query vector, and the contribution weight is adaptively adjusted through a gating coefficient.

[0023] According to a specific implementation manner of the embodiment of the present disclosure, the specific steps of step 3 include:

[0024] Step 3.1, define the region segmentation sub-task loss function as

[0025]

[0026] where m is the number of multi-classification categories, n is the number of pixels, P ij is the predicted probability of the j-th pixel for the i-th class, and G ij is the actual probability of the j-th pixel for the i-th class;

[0027] Step 3.2, define the pixel-to-boundary distance regression sub-task loss function as

[0028]

[0029] where is the shortest distance from the pixel to the predicted farmland area, is the ground truth distance;

[0030] Step 3.3, define the boundary detection subtask loss function as

[0031]

[0032] where k is the level of boundary prediction, is the k-level boundary loss;

[0033] Step 3.4, based on the maximization of the mean squared error uncertainty likelihood estimation, establish a multi-task loss function according to the region segmentation subtask loss function, the pixel-to-boundary distance regression subtask loss function, and the boundary detection subtask loss function

[0034]

[0035] where the parameters σ1, σ2, and σ3 are the noise parameters of the region segmentation subtask, the boundary detection subtask, and the pixel-to-boundary distance regression subtask, respectively;

[0036] Step 3.5, train a multi-task joint learning model according to the multi-task loss function and the feature matrix after fusion weighting to obtain an extraction model.

[0037] The field extraction scheme based on the PolSAR multi-class segmentation multi-task network in the embodiments of the present disclosure includes: Step 1, extract the polarization parameter set corresponding to the polarimetric SAR data based on the polarization decomposition theory and construct a feature matrix accordingly; Step 2, construct a hybrid framework based on the convolutional neural network and the Transformer, and fuse and weight the local features and global features corresponding to the feature matrix accordingly; Step 3, establish a multi-task joint learning model, integrate three complementary tasks of region segmentation, boundary detection, and pixel-to-boundary distance field regression, and perform end-to-end optimization using the multi-task loss function and the feature matrix after fusion weighting to obtain an extraction model; Step 4, input the polarimetric SAR data corresponding to the target region into the extraction model to obtain the field extraction result.

[0038] The beneficial effects of the embodiments of the present disclosure are as follows: Through the solution of the present disclosure, the characteristics of time-series PolSAR data are integrated with the CNN-Transformer hybrid model to form a multi-level feature learning framework. The shallow network strengthens the extraction of local scattering features through convolution operations, and the deep layer uses the self-attention mechanism of Transformer to capture global context associations, realizing the collaborative representation from micro-texture to macro-structure. Compared with the traditional optical data-driven single-task model, it has three advantages: ① It verifies the ability of PolSAR data to identify the cultivated land range within a specific accuracy threshold, breaking through the limitation of the lack of optical data in cloudy and rainy areas; ② A heterogeneous feature fusion module is designed. In the shallow layer, deformable convolution is used to adapt to the geometric changes of farmland edges, and in the deep layer, a cross-attention mechanism is introduced to achieve cross-modal alignment of polarization scattering features and spatial semantics; ③ A multi-class and multi-task joint optimization mechanism is constructed to synchronously execute multi-class plot segmentation, boundary detection, and pixel-to-boundary distance regression tasks, improving the generalization of the model through parameter sharing and gradient interaction, and enhancing the extraction accuracy and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0040] Figure 1 It is a schematic flowchart of a method for extracting field plots based on a PolSAR multi-class segmentation and multi-task network provided by the embodiments of the present disclosure;

[0041] Figure 2 It is a schematic diagram of the specific implementation process of a method for extracting field plots based on a PolSAR multi-class segmentation and multi-task network provided by the embodiments of the present disclosure;

[0042] Figure 3 It is a schematic diagram of PolSAR data and corresponding ground truth samples provided by the embodiments of the present disclosure;

[0043] Figure 4 It is a schematic diagram of the results of farmland plots extracted by different methods provided by the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0045] The following describes the embodiments of the present disclosure through specific specific examples. Those skilled in the art can easily understand the other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0046] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using other structures and / or functionality in addition to one or more of the aspects described herein.

[0047] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure in a schematic manner. The diagrams only show the components related to the present disclosure and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0048] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0049] In recent years, the revolutionary progress of remote sensing technology has made it a key supporting technology for the cultivated land dynamic monitoring system, strongly driving the intelligent transformation of the agricultural resource management mode. The technological development context has evolved from the initial land cover classification relying on single satellite images to the innovative stage of multi-modal remote sensing data fusion and artificial intelligence algorithm collaboration. Its multi-dimensional perception ability and spatio-temporal modeling advantages have continuously broken through the accuracy ceiling and efficiency threshold of cultivated land monitoring. Notably, the leapfrog improvement in the spatial resolution, temporal continuity, and open sharing of multi-source remote sensing data has enabled the high-precision positioning of farmland plot boundaries and the refined analysis of internal heterogeneous features to become a reality. The large-scale cultivated land intelligent interpretation system based on remote sensing is gradually replacing traditional manual surveys, highlighting significant engineering practice value, and the semantic segmentation and object recognition technologies empowered by deep learning have led the automation technology innovation of cultivated land information extraction.

[0050] In the field of intelligent extraction of farmland plots, the technical paradigm has shifted from manually designed feature engineering to being fully driven by deep learning. Compared with traditional algorithms, deep learning demonstrates significant competitiveness in maintaining the geometric accuracy and morphological integrity of farmland boundaries due to its autonomous feature learning ability and end-to-end optimization characteristics. However, this task still faces multiple technical bottlenecks: the problem of blurred ridge boundaries caused by the canopy shielding effect during the lush vegetation period; the generation of pseudo-edges due to the internal texture heterogeneity and noise interference of plots; the challenge to the model generalization performance in scenarios with coexisting multi-scale farmlands (such as terraced micro-plots and plain contiguous farms); the boundary adhesion effect between densely distributed plots increasing the difficulty of semantic segmentation; the over-segmentation error caused by crop growth differences (such as pest and disease areas); and the stringent requirements of the model's dynamic adaptability for the evolution of phenological period temporal features. Notably, the breakthrough of instance segmentation technology based on deep convolutional neural networks provides new ideas for the above problems. Related research has significantly enhanced the geometric fidelity of farmland boundaries by optimizing frameworks such as Mask R-CNN and YOLO.

[0051] For the extraction requirements in complex scenarios, the multi-task collaborative learning framework is becoming a new trend in technological evolution. This paradigm realizes the complementary enhancement of feature representation and the joint constraint of error propagation by jointly optimizing semantically related sub-tasks (such as plot range segmentation, boundary localization, distance field calculation, etc.). Representative achievements include: The ResUNet-a model innovatively constructs a four-task joint optimization network, which synchronously realizes cultivated land mask generation, boundary detection, pixel-level distance field regression, and image reconstruction in an improved U-Net architecture. Its multi-scale feature interaction mechanism effectively improves the recognition robustness of small-scale plots; BSiNet uses a dual-branch decoder to realize the decoupled learning of regional features and boundary features; SEANet introduces a boundary attention module to enhance the contour perception ability; CTMENet fuses the CNN and Transformer architectures and uses the self-attention mechanism to capture the global spatial associations of farmland arrangements; The CLPs model designs a boundary weighted fusion strategy to solve the fracture problem under complex terrains; The PLR-Net constructs a point-line-plane joint learning framework and realizes sub-pixel level accuracy improvement through triple constraints of vertex detection-boundary tracking-region segmentation.

[0052] It should be particularly emphasized that current research mainly focuses on optical remote sensing data, and there is no exploration of multi-task deep learning for farmland plot extraction based on synthetic aperture radar (SAR) data. From April to October in southern China, the cloud and rain coverage rate is as high as 70%, severely restricting the availability of optical images. Against this background, developing SAR cultivated land intelligent interpretation technology with all-weather monitoring capabilities has great application value. Especially the unique response mechanism of multi-polarized SAR data to the surface dielectric characteristics and geometric structures provides an irreplaceable technical path for monitoring the cultivated land status in cloudy and rainy regions.

[0053] The embodiments of the present disclosure provide a method for extracting farmland plots based on a PolSAR multi-class segmentation multi-task network, and the method can be applied to the detection process of farmland plots in synthetic aperture radar scenarios.

[0054] See Figure 1 , which is a schematic flowchart of a method for extracting farmland plots based on a PolSAR multi-class segmentation multi-task network provided by the embodiments of the present disclosure. As Figure 1 and Figure 2 shown, the method mainly includes the following steps:

[0055] Step 1, extract the polarization parameter set corresponding to the polarimetric SAR data based on the polarization decomposition theory and construct a feature matrix accordingly;

[0056] Furthermore, the specific steps of Step 1 include:

[0057] Step 1.1, decompose the polarimetric SAR data through the Cloude-Pottier decomposition theory to generate the polarimetric entropy H, the average scattering angle α, and the anisotropy parameter A, and decompose the polarimetric SAR data into surface scattering P s , dihedral scattering P d , and the volume scattering power ratio P v ;

[0058] Step 1.2, based on the polarimetric entropy H, the average scattering angle α, the anisotropy parameter A, surface scattering P s , dihedral scattering P d , and the volume scattering power ratio P v , construct a three-dimensional feature system through a hierarchical feature fusion strategy, and construct a feature matrix accordingly.

[0059] Furthermore, the three-dimensional feature system includes a basic layer, an energy layer, and a cross-derived layer.

[0060] Furthermore, the step of constructing a three-dimensional feature system based on the polarimetric entropy H, the average scattering angle α, the anisotropy parameter A, surface scattering P s , dihedral scattering P d , and the volume scattering power ratio P v , through a hierarchical feature fusion strategy, includes:

[0061] Step 1.2.1, use the polarimetric entropy H, the average scattering angle α, and the anisotropy parameter A as the basic layer to analyze the scattering mechanism characteristics;

[0062] Step 1.2.2, use surface scattering P s , dihedral scattering P d , and the volume scattering power ratio P v as the energy layer to quantify the scattering power distribution;

[0063] Step 1.2.3, use the polarimetric entropy H, the average scattering angle α, the anisotropy parameter A, surface scattering P s , dihedral scattering P d , and the volume scattering power ratio P v as the cross-derived layer, and generate high-order features through non-linear combination.

[0064] This study focuses on the problem of dynamic monitoring of cultivated land in cloudy and rainy terrain areas. Aiming at the characteristics of fragmented farmland distribution and diverse crop types in Guangdong, Hunan and other places, it endeavors to break through the cognitive bottleneck of the spatio-temporal scattering mechanism of polarimetric SAR. Although traditional optical vegetation indices (NDVI / RVI) can reflect the growth trend of crops, they are limited by the canopy spectral saturation effect and insufficient characterization of structural features, making it difficult to accurately analyze the spatial configuration and biodiversity characteristics of crops. In contrast, the temporal PolSAR system has unique advantages: its microwave penetration characteristics can overcome cloud interference and enable short-cycle revisit monitoring, and the temporal polarization feature matrix can invert the morphological structure (stem and leaf orientation), physiological state (water content) of crops at different growth stages, and surface parameters (soil roughness).

[0065] For the extraction of farmland plots, the most advanced method is to use multi-task deep learning technology to simultaneously obtain the scope and boundaries of farmland plots. However, existing studies are all based on optical data, and optical remote sensing is vulnerable to cloudy, rainy and snowy weather, resulting in many provinces in southern China being unable to obtain available optical images with large-scale coverage during the rainy season every year, and only complete and available images can be obtained in the rainless winter. This situation is unacceptable for the extraction of farmland plots in the rainy season in the cloudy and rainy areas of southern China. Polarimetric SAR has the advantage of stable earth observation and can extract continuous temporal characteristics of cultivated land; replacing the CNN feature extraction architecture with a more advanced hybrid architecture of CNN and Transformer can model the long-range context information of the image and consider the shape and arrangement rules of the fields; further detailed classification of the crops in the cultivated land helps to distinguish adjacent different crop fields and reduce field adhesion. Based on this, this project plans to use polarimetric SAR data, add an advanced hybrid architecture of CNN and Transformer, and a crop classification sub-task to extract farmland plots. The main research contents include: 1) Based on polarimetric target decomposition methods such as Cloude-Pottier decomposition and Freeman-Durden decomposition, extract polarimetric parameter sets (H / A / α, volume scattering power ratio, etc.), and construct different feature matrices including polarimetric entropy, scattering angle, anisotropy, etc.; 2) Introduce an advanced hybrid architecture of CNN and Transformer; 3) Add a crop classification sub-task.

[0066] In specific implementation, polarimetric parameter sets (H / A / α, volume scattering power ratio, etc.) can be extracted based on polarimetric target decomposition methods such as Cloude-Pottier decomposition and Freeman-Durden decomposition, and different feature matrices including polarimetric entropy, scattering angle, anisotropy, etc. can be constructed.

[0067] Based on polarization decomposition theories such as Cloude-Pottier and Freeman-Durden, a multi-dimensional scattering feature characterization framework can be established. The Cloude-Pottier decomposition generates polarization entropy H, mean scattering angle α, and anisotropy A parameters. Among them, the H parameter characterizes the randomness of scattering (value range 0-1), the α parameter discriminates the dominant scattering type (0° corresponds to surface scattering, 45° is dipole scattering, and 90° indicates double-bounce scattering), and the A parameter characterizes the difference of the sub-optimal scattering components. Combining with the Freeman-Durden decomposition, the power ratios of surface scattering (Ps), dihedral angle scattering (Pd), and volume scattering (Pv) are obtained, which can quantify the random scattering intensity of the vegetation canopy. A three-dimensional feature system is constructed through a hierarchical feature fusion strategy: the basic layer (H / α / A) analyzes the characteristics of the scattering mechanism, the energy layer (Ps / Pd / Pv) quantifies the scattering power distribution, and the cross-derived layer generates high-order features through non-linear combinations (such as Pv / (Ps+Pd)) to significantly improve the distinguishability between the interior of the farmland and the surrounding ground objects, providing a feature construction paradigm that combines physical mechanisms and statistical characteristics for the interpretation of polarimetric SAR images. The Cloude-Pottier decomposition of the polarization coherence matrix is as follows:

[0068]

[0069] where λ i , U i are the eigenvalues and eigenvectors of the coherence matrix. Since the coherence matrix is a semi-Hermitian matrix, all three eigenvalues are greater than zero. It can be found that the decomposition based on eigenvalues has no fixed decomposition basis. This makes the decomposition method not interfered by the artificially preset decomposition basis, but it also makes the interpretation of eigenvalues and eigenvectors no longer simple. To better interpret ground objects, there are three widely used parameters. The first parameter is the scattering angle α:

[0070]

[0071] The scattering angle contains the structural information of the ground object, and the structure of the ground object can be analyzed through the scattering angle. When α→0°, the scatterer is an isotropic surface; when α→45°, the scatterer is a dipole; when α→90°, the scatterer is an isotropic dihedral angle. The second parameter is the scattering entropy H:

[0072]

[0073] The degree of disorder of ground objects can be analyzed according to the magnitude of the scattering entropy. When H < 0.3, the ground scatterers can be regarded as deterministic targets. When H increases, the ground scatterers are composed of a mixture of various scattering targets. In the extreme case when H = 1, the ground scatterers are completely random. Note that the scattering entropy and the eigenvalues do not have a mapping relationship. As a supplement to the scattering entropy, anisotropy can be used to further analyze the scattering characteristics of ground objects:

[0074]

[0075] It can be found that anisotropy is mainly used to analyze the relationship between the second scattering mechanism and the third scattering mechanism. Anisotropy is only effective when the scattering entropy is relatively large because when the scattering entropy is too small, the second scattering mechanism and the third scattering mechanism are vulnerable to noise. Freeman divided the scattering mechanisms of the earth's surface into surface scattering, dihedral angle scattering, and volume scattering. Among them, surface scattering is generated by a slightly rough surface. The scattering matrix corresponding to surface scattering is:

[0076]

[0077] It can be found that the surface scattering model is closely related to the ground soil moisture. Therefore, the surface scattering model plays an important role in soil moisture inversion. The second scattering mechanism is dihedral angle scattering, which is generated by a pair of mutually perpendicular planes. The dielectric constants of the two perpendicular planes and the extinction coefficients of different polarization channels are considered simultaneously. Assuming that the dihedral angle scatterers satisfy a definite distribution, the corresponding coherence matrix is:

[0078]

[0079] The third scattering mechanism is volume scattering, which is assumed to be composed of completely random dipoles. Taking the horizontal dipole as an example, the corresponding scattering matrix is:

[0080]

[0081] Performing incoherent superposition on the above three scattering mechanisms, the form of the final three-component decomposition method is:

[0082]

[0083] Step 2: Construct a hybrid framework based on a convolutional neural network and a Transformer, and accordingly perform fusion weighting on the local features and global features corresponding to the feature matrix;

[0084] Based on the above embodiments, the specific content of the said Step 2 includes:

[0085] Step 2.1: Build a five-level feature pyramid using a depthwise separable convolutional neural network to extract multi-scale local texture features in the feature matrix. After each convolutional layer, connect a CBAM attention module to enhance spatial-channel feature screening;

[0086] Step 2.2: When the feature matrix passes through each level of the pyramid layer, it is sequentially divided into image patches of a preset size, linearly mapped to generate a serialized input to the Transformer encoder. Design axial attention in the Transformer encoder to reduce the computational complexity, and at the same time introduce deformable convolutional kernels to replace the standard position encoding;

[0087] Step 2.3: Use gated cross-attention to dynamically fuse the detailed features of the depthwise separable convolutional neural network with the global semantics of the Transformer. Among them, the features output by the depthwise separable convolutional neural network are used as key vectors and value vectors, and the output of the Transformer encoder is used as the query vector, and the contribution weights are adaptively adjusted through the gating coefficient.

[0088] In specific implementation, use a hybrid architecture of CNN and Transformer to effectively fuse the global features and local detailed information of the image.

[0089] Convolutional neural networks (CNNs) have established a core position in the field of computer vision due to their unique local perception and parameter sharing mechanisms. However, these characteristics also bring significant advantages and limitations. In terms of advantages, CNNs extract local features (such as edges and textures) layer by layer through small convolutional kernels like 3×3, and achieve translational invariance with the help of max pooling, reaching a Top-5 accuracy of over 95% in the ImageNet image classification task; the parameter sharing mechanism significantly reduces the computational load, and ResNet-50 only requires 25.6 million parameters to process 224×224 inputs, with the number of parameters reduced by 99% compared to fully connected networks; their hardware friendliness enables real-time object detection on edge devices such as Jetson Nano (YOLOv5s reaches 30 FPS). In terms of disadvantages, the local receptive field of CNNs limits their global modeling ability. In remote sensing images, the association of cross-regional plots requires the use of dilated convolutions to expand the receptive field, resulting in a 42% increase in computational costs; the vanishing gradient problem in deep networks needs to be alleviated by introducing residual connections (such as ResNet), but the model interpretability is still poor, making it difficult to trace the physical meaning of specific convolutional kernels; in terms of data dependence, at least thousands of labeled samples are required for medical image segmentation tasks to avoid overfitting. The Transformer model has shown revolutionary breakthroughs in the field of artificial intelligence with its self-attention mechanism and parallel architecture, but it also faces significant technical challenges, and its application scenarios are rapidly expanding from traditional natural language processing to multi-modal fields. In terms of advantages, Transformer dynamically models long-range dependencies through global attention weights, achieving breakthroughs in tasks such as machine translation and protein structure prediction; the parallel computing feature enables the training efficiency of models with hundreds of billions of parameters such as GPT-4 to be tripled, supporting multi-modal unified architectures (such as the cross-modal retrieval accuracy of CLIP for images and texts exceeding supervised models). In terms of defects, the complexity of self-attention leads to the memory occupancy being 16 times that of CNNs when processing 4K-length sequences, and it is necessary to rely on sparse attention (such as Longformer) or chunking strategies to compress the computational load; in terms of data dependence, BERT requires pre-training on over 100 GB of corpus, and it is highly sensitive to hyperparameters. Self-supervised techniques need to be combined in small sample scenarios to alleviate overfitting. In computer vision, ViT has surpassed ResNet with an ImageNet classification accuracy of 88.55%. Although CNNs are good at capturing local features such as ridge textures and crop patches, their limited receptive field makes it difficult to model cross-regional global associations (such as the spatial topological relationship of irrigation canal networks), resulting in hierarchical misalignment of mountain terraces or breaks in contiguous farmland in plains; while the Transformer can model long-range dependencies with its self-attention mechanism, it is less sensitive to high-frequency details and is prone to missing the boundaries of fragmented plots. Hybrid architectures can achieve complementarity through hierarchical feature interactions (such as the convolutional-attention alternating module of MobileViT).

[0090] Use a hybrid architecture of CNN and Transformer to collaboratively optimize local feature extraction and global context modeling capabilities, and achieve efficient and accurate remote sensing image interpretation. The core implementation steps are as follows: In the architecture design stage, depthwise separable convolutions are used at the front end to construct a five-level feature pyramid to extract multi-scale local texture features (such as farmland ridges and vegetation patches). After each convolutional layer, a CBAM attention module is connected to strengthen spatial-channel feature screening; in the feature interaction stage, the third-level feature map is divided into 16×16 image patches, linearly mapped to generate a serialized input for the Transformer encoder. Axial Attention is designed in the encoder to reduce computational complexity, and at the same time, deformable convolutional kernels are introduced to replace the standard position encoding to enhance geometric adaptability to irregular farmland boundaries; Gated Cross-Attention is used to dynamically fuse the detailed features of CNN and the global semantics of Transformer. Among them, the CNN features are used as Key / Value, and the Transformer output is used as Query, and the contribution weights are adaptively adjusted through the gating coefficient. The local sensitivity of CNN makes up for the missed detection problem of Transformer for fragmented plots, while the long-range dependence of Transformer solves the modeling bottleneck of CNN in cross-region association and effectively considers the shape and arrangement rules of fields, providing an optimal balance solution for real-time interpretation of complex agricultural scenarios.

[0091] Step 3, establish a multi-task joint learning model, integrate three complementary tasks of region segmentation, boundary detection, and pixel-to-boundary distance field regression, and use a multi-task loss function and a fused and weighted feature matrix for end-to-end optimization to obtain an extraction model;

[0092] Based on the above embodiments, step 3 specifically includes:

[0093] Step 3.1, define the loss function of the region segmentation sub-task as

[0094]

[0095] where m is the number of multi-classification categories, n is the number of pixels, P ij is the predicted probability of the j-th pixel for the i-th class, and G ij is the actual probability of the j-th pixel for the i-th class;

[0096] Step 3.2, define the loss function of the pixel-to-boundary distance regression sub-task as

[0097]

[0098] where, is the shortest distance from the pixel to the predicted farmland area, is the ground truth distance;

[0099] Step 3.3, define the boundary detection sub-task loss function as

[0100]

[0101] where k is the level of boundary prediction, is the k-th level boundary loss;

[0102] Step 3.4, based on the maximization of the mean square error uncertainty likelihood estimation, establish a multi-task loss function according to the region segmentation sub-task loss function, the pixel-to-boundary distance regression sub-task loss function, and the boundary detection sub-task loss function

[0103]

[0104] where the parameters σ1, σ2, and σ3 are the noise parameters of the region segmentation sub-task, the boundary detection sub-task, and the pixel-to-boundary distance regression sub-task, respectively;

[0105] Step 3.5, train a multi-task joint learning model according to the multi-task loss function and the fused weighted feature matrix to obtain an extraction model.

[0106] In specific implementation, a multi-task joint learning model can be established to integrate three complementary tasks of region segmentation, boundary detection, and pixel-to-boundary distance field regression, and an adaptive weight loss function is used for end-to-end optimization.

[0107] By jointly optimizing the three tasks of farmland plot range segmentation, boundary detection, and pixel-to-boundary distance regression, an intelligent interpretation framework with deep fusion of feature complementarity and geometric constraints is constructed. The range task predicts the coverage area of the farmland plot and the pixel-to-boundary distance to provide auxiliary constraints. On the other hand, the boundary task provides a more accurate prediction for the division of the field boundary. The farmland plot range task is designed as a multi-classification branch to distinguish different crop categories within the cultivated land, weaken the influence of boundary adhesion between adjacent different crop fields caused by the dense distribution of farmland plots, and improve the accuracy of the plot boundary. A different decoder is designed for each task to adapt to their different features.

[0108] The farmland range segmentation task is a multi-classification segmentation architecture for different crop categories and non-cultivated land. The crop categories include meadow, corn, wheat, barley, potato, rye, rapeseed, fallow land, and vegetables, etc. The decoder of this task is designed based on the decoding component of the ResU-Net network. This decoding method promotes the adaptation to farmland plots with different shapes and sizes by fusing multi-level features. The pixel-to-boundary distance regression sub-task and the range segmentation task share the same decoder, which is used to assist and constrain the farmland range segmentation task. The farmland range segmentation sub-task loss function is as follows:

[0109]

[0110] Among them, m is the number of categories in multi-classification, n is the number of pixels, and P ij is the predicted probability of the j-th pixel for the i-th class, and G ij is the actual probability of the j-th pixel for the i-th class. The loss function of the pixel-to-boundary distance regression sub-task is as follows:

[0111]

[0112] Among them, is the shortest distance from the pixel to the predicted farmland area, is the ground truth distance.

[0113] The side structure is used as the decoder for the field boundary task to generate edge maps at each stage and integrate them with weights into the field boundary of interest. The side structure can effectively capture multi-level boundary features, resulting in a more comprehensive hierarchical edge representation. In addition, the multi-level prediction of the side structure can also promote deeper supervision of the boundary detection results. To optimize the feature map at each stage, a compact dilated convolution-based module (CDCM) is added to enrich the edge information before decoding the features. This module first reduces the dimension of the output multi-channel features to 21 using a 1×1 convolutional kernel. This parameter value is aligned with the feature size reduction parameter setting of RCF. Then, the dimension-reduced features pass through a series of expanded convolutional kernels to expand the receptive field of the boundary features and enrich their representation. Since the field boundary has a large span, the rich long-range auxiliary information provided by the stacked dilated convolutional layers of CDCM better preserves the overall connectivity of the boundary and more accurately describes the classification features. In addition, this method maintains fine-grained detailed information without introducing downsampling, which is crucial for narrow-width boundary objects. The boundary sub-task loss function is as follows:

[0114]

[0115] Among them, k is the level of boundary prediction, is the boundary loss at the k-th level.

[0116] Based on the maximization of the mean square error uncertainty likelihood estimation, an adaptive multi-task loss function is formulated. The loss function is as follows:

[0117]

[0118] Among them, the parameters σ1, σ2, and σ3 correspond to the noise parameters of the three tasks respectively, reflecting their uncertainties. A higher noise value indicates an increase in uncertainty and a decrease in the task weight.

[0119] Step 4: Input the polarimetric SAR data corresponding to the target area into the extraction model to obtain the result of field block extraction.

[0120] Specifically, when implementing, after training the extraction model, when it is necessary to identify the field blocks in a certain place, the polarimetric SAR data corresponding to the target area can be collected and input into the extraction model to obtain the result of field block extraction.

[0121] The method for field block extraction based on the PolSAR multi-class segmentation multi-task network provided in this embodiment forms a multi-level feature learning framework by fusing the features of temporal PolSAR data and the CNN-Transformer hybrid model: the shallow network strengthens the extraction of local scattering features through convolutional operations, and the deep layer uses the self-attention mechanism of Transformer to capture global context associations, realizing the collaborative representation from micro-texture to macro-structure. Compared with the traditional optical data-driven single-task model, it has three advantages: ① It verifies the ability of PolSAR data to identify the cultivated land range within a specific accuracy threshold, breaking through the limitation of the lack of optical data in cloudy and rainy areas; ② Design a heterogeneous feature fusion module, adapt to the geometric changes of farmland edges through deformable convolutions in the shallow layer, and introduce a cross-attention mechanism in the deep layer to achieve cross-modal alignment of polarimetric scattering features and spatial semantics; ③ Construct a multi-class multi-task joint optimization mechanism, synchronously execute multi-class plot segmentation, boundary detection, and pixel-to-boundary distance regression tasks, and improve the generalization of the model through parameter sharing and gradient interaction, improving the extraction accuracy and adaptability.

[0122] The method of the present disclosure will be further described below in combination with a specific embodiment. Data description: ① The experimental area is located in Area A, and the vector information of farmland plots in this area in 2019 has been obtained, including the range and crop categories of the plots, which can meet the necessary requirements of the experiment. ② The polarimetric SAR data used in the experiment is Sentinel-1 GRD data, with a resolution of 10 meters, and the image acquisition time is January 2019 (to minimize the impact of changes in crop cultivation categories on the experimental results). It is subjected to preprocessing operations such as orbit correction, multi-look, and filtering, and then the parameter VV / VH is calculated, so as to jointly synthesize a pseudo-color image with VV and VH, as shown in Figure 3 (a) in. ③ By rasterizing the cultivated land vector according to the requirements of the range and resolution of the corresponding polarimetric SAR data, the sample image used in the experiment is obtained, as shown in Figure 3 (b) in.

[0123] To compare the advantages of the method of the present invention, the obtained data and sample slices are used to perform farmland plot extraction using the method of the present invention and the methods of previous scholars respectively. Among them, the previous method refers to CLPs, and the method of the present invention is an improvement based on CLPs, abbreviated as m-CLPs. As shown in Figure 4As shown, the results obtained by different methods are presented. Among them, (a) represents the extraction result of the CLPs method, and (b) represents the extraction result of the m-CLPs method. It can be clearly seen that the extraction effect of the method of the present disclosure is superior to the existing solutions. Further, the precision indicators of different methods are calculated, and it can be seen that the farmland plot results obtained by the method of the present invention are more accurate than those of the CLPs method.

[0124] It should be understood that each part of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof.

[0125] As described above, the above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present disclosure should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for extracting field plots based on a multi-class segmentation multi-task network of PolSAR, characterized in that, Including: Step 1: Extract the polarization parameter set corresponding to the polarimetric SAR data based on the polarization decomposition theory and construct a feature matrix accordingly; Step 2: Construct a hybrid framework based on the convolutional neural network and Transformer, and accordingly perform fusion weighting on the local features and global features corresponding to the feature matrix; Step 3: Establish a multi-task joint learning model, integrate three complementary tasks of region segmentation, boundary detection, and pixel-to-boundary distance field regression, and perform end-to-end optimization using the multi-task loss function and the fusion-weighted feature matrix to obtain the extraction model; Step 4: Input the polarimetric SAR data corresponding to the target area into the extraction model to obtain the result of field block extraction.

2. The method according to claim 1, wherein The specific content of the above Step 1 includes: Step 1.1, decompose the polarimetric SAR data through the Cloude-Pottier decomposition theory to generate the polarimetric entropy H, the average scattering angle α, and the anisotropy parameter A, and decompose the polarimetric SAR data into the surface scattering P s , the dihedral scattering P d , and the volume scattering power ratio P v ; Step 1.2, based on the polarization entropy H, average scattering angle α, anisotropy parameter A, surface scattering P s , dihedral angle scattering P d , and the proportion of volume scattering power P v , a three-dimensional feature system is constructed through a hierarchical feature fusion strategy, and a feature matrix is constructed accordingly.

3. The method according to claim 2, wherein The three-dimensional feature system includes a basic layer, an energy layer, and a cross-derived layer.

4. The method according to claim 3, characterized in that Based on the polarization entropy H, average scattering angle α, anisotropy parameter A, surface scattering P s , dihedral angle scattering Pd, and the proportion of volume scattering power P v , the steps of constructing a three-dimensional feature system through a hierarchical feature fusion strategy include: Step 1.2.1: Use the polarization entropy H, the average scattering angle α, and the anisotropy parameter A as the basic layer to analyze the scattering mechanism characteristics; Step 1.2.2, taking the surface scattering P s , the dihedral angle scattering P d and the volume scattering power ratio P v as the energy layer for quantifying the scattering power distribution; Step 1.2.3, using the polarization entropy H, average scattering angle α, anisotropy parameter A, surface scattering P s , dihedral angle scattering Pd, and the proportion of volume scattering power P v as the cross-derived layer, and generating high-order features through non-linear combination.

5. The method according to claim 4, characterized in that The specific content of the above Step 2 includes: Step 2.1: Use a depthwise separable convolutional neural network to construct a five-level feature pyramid, extract the multi-scale local texture features in the feature matrix, and connect a CBAM attention module after each convolutional layer to strengthen the spatial-channel feature screening; Step 2.2: When the feature matrix passes through each level of the pyramid layer, it is sequentially segmented into image blocks of a preset size, linearly mapped to generate a serialized input to the Transformer encoder, and axial attention is designed in the Transformer encoder to reduce the computational complexity, and at the same time, a deformable convolutional kernel is introduced to replace the standard position encoding; Step 2.3: Use gated cross-attention to dynamically fuse the detailed features of the depthwise separable convolutional neural network and the global semantics of the Transformer. Among them, the features output by the depthwise separable convolutional neural network are used as the key vector and the value vector, and the output of the Transformer encoder is used as the query vector, and the contribution weight is adaptively adjusted through the gating coefficient.

6. The method according to claim 5, characterized in that, The specific content of the above Step 3 includes: Step 3.1: Define the loss function of the region segmentation sub-task as where m is the number of categories for multi-classification, n is the number of pixels, and P ij is the predicted probability of the j-th pixel belonging to the i-th class, and G ij is the actual probability of the j-th pixel belonging to the i-th class; Step 3.2: Define the loss function of the pixel-to-boundary distance regression sub-task as Among them, is the shortest distance from the pixel to the predicted farmland area, is the ground truth distance; Step 3.3: Define the loss function of the boundary detection sub-task as where k is the level of boundary prediction, is the boundary loss at the k-th level; Step 3.4: Based on the maximization of the mean square error uncertainty likelihood estimation, establish a multi-task loss function according to the loss function of the region segmentation sub-task, the loss function of the pixel-to-boundary distance regression sub-task, and the loss function of the boundary detection sub-task where the parameters σ1, σ2, and σ3 are the noise parameters of the region segmentation sub-task, the boundary detection sub-task, and the pixel-to-boundary distance regression sub-task respectively; Step 3.5: Train the multi-task joint learning model according to the multi-task loss function and the fusion-weighted feature matrix to obtain the extraction model.

Citation Information

Cited By

  • Farmland parcel identification method based on optical-Ka frequency band SAR (Synthetic Aperture Radar) feature fusion

    CN120544048A