Ultrahigh-resolution SAR image farmland extraction method and system based on hybrid deep learning model, storage medium and electronic equipment
By combining a hybrid deep learning model with the local residual module and VMamba's VSS module, the problem that farmland extraction in optical remote sensing images is easily affected by weather and consumes a lot of computing resources is solved, achieving efficient and accurate farmland extraction results.
Patent Information
- Application Number
- CN202510837559.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-03
AI Technical Summary
Existing farmland extraction methods are easily affected by weather in optical remote sensing images, traditional convolutional neural networks have difficulty capturing the spatial relationships of large-scale farmlands, and Transformer-based models consume large amounts of computing resources, resulting in low farmland extraction accuracy and efficiency.
A method based on a hybrid deep learning model is adopted. By combining the local residual module and the VSS module of VMamba, a local feature encoder and global feature modeling are constructed. Local shallow features are fused with deep semantic features, and transposed convolution upsampling is used to generate farmland extraction results.
The accuracy and efficiency of farmland extraction have been improved, and it can accurately extract farmland morphology and edges in complex backgrounds, supporting precision agriculture and land resource management.
Smart Images

Figure CN120747739A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of agricultural monitoring and remote sensing, and in particular to a method, system, storage medium and electronic device for extracting farmland from ultra-high-resolution SAR images based on a hybrid deep learning model. Background Art
[0002] Farmland extraction has important applications in remote sensing image analysis, particularly in precision agriculture, land resource management, and environmental monitoring. Accurate farmland extraction provides a scientific basis for agricultural planning and decision-making, helping to optimize land use, improve crop production efficiency, and sustainably utilize resources. In precision agriculture, farmland extraction provides fundamental data for crop growth monitoring, irrigation management, and pest and disease early warning. In land resource management and environmental protection, farmland extraction helps track changes in cultivated land, monitor land use patterns, and provide data support for ecological and environmental protection policies. Therefore, farmland extraction plays a crucial role in improving agricultural production efficiency and ensuring ecological and environmental sustainability.
[0003] However, existing farmland extraction methods face a series of challenges that affect their application in large-scale agricultural monitoring. First, optical remote sensing images are easily affected by weather factors. Weather phenomena such as clouds and haze can cause image quality to deteriorate, which in turn affects the accuracy and stability of farmland extraction. Secondly, among deep learning methods for extracting farmland, although traditional convolutional neural network methods can extract local features, they have certain limitations in modeling global features and cannot effectively capture the spatial relationships and interactions of large-scale farmland. In addition, the Transformer-based model can better capture global information, but it consumes a lot of computing resources and is less efficient when processing large-scale remote sensing data, resulting in high computational overhead in practical applications, which limits its promotion and application.
[0004] In summary, in order to overcome these challenges, improve the accuracy and efficiency of farmland extraction, explore a variety of new deep learning methods, and balance the accuracy, efficiency and computational overhead of farmland extraction by integrating the advantages of different models, which remains an important research direction in this field. Summary of the Invention
[0005] The purpose of this application is to provide a method and system for farmland extraction from ultra-high-resolution SAR images based on a hybrid deep learning model to solve or alleviate the problems existing in the above-mentioned prior art.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] The present application provides a method for extracting farmland from ultra-high-resolution SAR images based on a hybrid deep learning model, comprising: step S101, preprocessing ultra-high-resolution SAR image data and generating edge strength data thereof; step S102, using a local residual module to form a local feature encoder to extract local features, and outputting feature maps layer by layer; step S103, modeling the global relationship of features with linear complexity by introducing the VSS module of VMamba; step S104, reshaping the output feature map size of the global feature modeling, and performing conventional convolution to complete feature extraction of the feature encoder; step S105, fusing local shallow features with deep semantic features through a shallow-deep feature fusion module, completing jump connections, and providing semantic information for upsampling; step S106, using transposed convolution to upsample to the input image size, obtaining extraction results, and performing accuracy evaluation.
[0008] Preferably, in step S101, the ultra-high resolution SAR image data is preprocessed and edge strength data is generated, specifically:
[0009] Multi-view processing, filtering and noise reduction, terrain correction and decibel conversion are performed on ultra-high resolution SAR data to optimize image quality and improve the accuracy of subsequent feature extraction. Edge intensity data is then calculated based on the processed ultra-high resolution SAR data:
[0010] Let I(x,y) be the SAR scattering intensity. In a window of size N×N centered at a certain pixel, the scattering intensity is approximated as a quadratic polynomial model:
[0011] I(x,y)≈P·v(x,y) T
[0012] Where (x, y) is the coordinate of the origin of the N×N window center, P = [A, B, C, D, E, F], v (x, y) = [x 2 ,y 2 ,xy,x,y,1]; the coefficient vector P is solved by the least squares criterion.
[0013] Construct the local design matrix v and observation vector I, then the corresponding minimization P objective function is:
[0014]
[0015] Then the analytical solution of vector P is:
[0016] P=(v t v) -1 v T I
[0017] Based on the fitting model, the first-order and second-order partial derivatives of x and y are obtained:
[0018]
[0019] Since x=0, y=0 at the center of the local window, the gradient amplitude at the center point can be expressed as:
[0020]
[0021] The Hessian matrix is a square matrix that describes the local curvature of a multivariate scalar function. The Hessian matrix is constructed from the second-order partial derivatives at the center, and we can get:
[0022]
[0023] Let K = max(|λ1|,|λ2|) be the local curvature measure to reflect the sensitivity of the image to structural changes at that point, where λ1 and λ2 are the eigenvalues of the matrix. The magnitude of image changes and the sudden change in local curvature can better highlight farmland in a complex background. Combining gradient information and local curvature, the specific formula for edge strength data can be expressed as:
[0024]
[0025] Preferably, in step S102, a local feature encoder is constructed using a local residual module to extract local features and output feature maps layer by layer, specifically:
[0026] The pre-processed ultra-high resolution SAR data and edge strength data are used as the input of the deep learning model, and the local feature encoder is constructed using the local residual module to extract local features. This step is divided into four sub-stages, each stage contains several local residual modules, and the feature map output by each sub-stage is named and
[0027] The local residual module is based on the residual block, but the second traditional convolutional layer in the residual block is replaced by a dynamic snake convolution. Dynamic snake convolution is more sensitive to tubular structures and can better adapt to ridges and roads. Its formula is as follows:
[0028]
[0029] Where i represents the number of sub-stages, i∈{1,2,3,4}; and Represents the intermediate output and output of the i-th stage respectively; BN represents the batch normalization operation, Conv 3×3 It is a convolution operation with a convolution kernel size of 3×3, DSConv is a dynamic snake convolution operation, and Relu is the activation function.
[0030] Preferably, in step S103, the global relationship of features is modeled with linear complexity by introducing the VSS module of VMamba, specifically:
[0031] Local features are mapped to a higher-dimensional semantic space through patch embedding, where long-range dependencies are established. This phase is also divided into four sub-stages, each of which uses several VSS modules to maintain the resolution of the feature map.
[0032] Two-dimensional selective scanning is the core mechanism of the VSS module to achieve global relationship modeling with O(n) complexity. Its workflow can be summarized into three main steps: expansion, selection, and reorganization. First, the two-dimensional feature map is expanded into independent one-dimensional sequences along four diagonal directions, covering the contextual association paths of all pixels. The sequence in each direction is input into the parameter-adaptive S6 module, which dynamically adjusts the hidden state weights through a gating mechanism to selectively enhance important features and suppress redundant information. Finally, the sequences processed in the four directions are folded back into two-dimensional space along the original path. The feature responses of different scanning directions are retained through weighted aggregation, and the enhanced feature map that integrates global long-range dependencies is finally output. Its formula is as follows:
[0033] x v =expand(x,v)
[0034]
[0035] Among them, x and The input and output of the two-dimensional selection scan; v represents the four scanning directions, v∈{1,2,3,4}; expand represents the expansion operation, S6 represents the S6 module processing, and merge is the sequence merge operation.
[0036] Preferably, in step S104, the output result feature map size of the global feature modeling is reshaped, and conventional convolution is performed to complete feature encoder feature extraction, specifically:
[0037] After the fourth sub-stage of global feature extraction, the deep semantic features are swapped in the channel dimension. Subsequently, a 3×3 convolution optimizes the feature representation, and the encoding stage ends. The formula is as follows:
[0038]
[0039] Among them, x is the output result of the fourth sub-stage of global feature extraction, Input for the decoding stage; Conv 3×3 It is a convolution operation with a convolution kernel size of 3×3, and reverse is a feature dimension adjustment operation.
[0040] Preferably, in step S105, the shallow-deep feature fusion module is used to fuse local shallow features and deep semantic features to complete the skip connection and provide semantic information for upsampling, specifically:
[0041] Define the decoded output as and The shallow-deep feature fusion module concatenates the feature representations from the encoder and decoder along the channel dimension, thereby introducing richer multi-scale contextual information. The concatenated features are then fed into the VSS module for semantic information fusion. Next, by introducing learnable coefficients α and β, the shallow features, deep features, and their fusion are adaptively weighted, improving the model's stability and adaptability under different spatial structures. The formula is as follows:
[0042]
[0043] Among them, F i The output result of the deep-shallow feature module, i∈{1,2,3,4}, concat indicates that the channel is a splicing operation, V indicates VSS module processing, and α and β are learnable coefficients.
[0044] Preferably, in step S106, transposed convolution is used to upsample to the input image size to obtain the extraction result and perform accuracy evaluation, specifically:
[0045] For the output of the shallow-deep feature fusion module, transposed convolution upsampling is used to gradually restore the feature map size, and finally restore it to the input image size. The calculation formula is as follows:
[0046]
[0047] i∈{1,2,3,4},F i Output result of deep-shallow feature module; BN stands for batch normalization operation, TransConv 3×3 is the transposition operation of the convolution kernel size of 3×3, and Relu is the activation function. 4 Obtained by upsampling That is, the image after semantic segmentation of farmland.
[0048] The accuracy of the hybrid deep learning model was evaluated using both unlabeled and labeled images from the test set. The unlabeled images were directly captured from preprocessed ultra-high-resolution SAR images with edge strength information, while the labeled images were obtained by labeling the unlabeled images according to preset classification labels. The accuracy evaluation metrics used were precision, recall, F1, and IoU, which evaluated the model's ability to extract farmland. The calculation formula is as follows:
[0049]
[0050] Among them, TP, FP, FN and TN represent the number of true positive, false positive, false negative and true negative pixels in the prediction results, respectively.
[0051] The present application also provides an ultra-high-resolution SAR image farmland extraction system based on a hybrid deep learning model, comprising:
[0052] a data processing unit configured to segment the pre-processed ultra-high resolution SAR image and calculate and derive an edge intensity image;
[0053] The feature encoding unit is configured to use a hybrid feature encoder based on CNN-VMamba to extract features from the input data and obtain features of data at different scales;
[0054] The feature fusion unit is configured to fuse shallow features with decoded deep features based on a two-dimensional selective scanning mechanism, organically combining the three through dynamic weighting to improve feature expression capabilities;
[0055] The feature decoding unit is configured to perform semantic segmentation on the input data based on the fused features, restore the image size through upsampling, and obtain the extraction result after semantic segmentation.
[0056] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the program is a method for extracting farmland from ultra-high-resolution SAR images based on a hybrid deep learning model as described above.
[0057] An embodiment of the present application also provides an electronic device, comprising: a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, an ultra-high-resolution SAR image farmland extraction system based on a hybrid deep learning model as described above is implemented.
[0058] Beneficial effects:
[0059] In the method for farmland extraction from ultra-high-resolution SAR images based on a hybrid deep learning model provided in this application, the ultra-high-resolution SAR image is first preprocessed and the corresponding edge intensity map is generated; then, a feature encoder is constructed based on the local residual module to extract and output the local feature map layer by layer; then the VSS module of VMamba is introduced for global feature modeling, and the feature size is reshaped by conventional convolution to complete the feature encoding on the encoder side; then, in the shallow-deep feature fusion module, the shallow features of the encoding stage are fused with the deep semantic features of the decoding stage to establish a jump connection; finally, the fused features are upsampled to the original image size using transposed convolution to generate the final farmland extraction result. The method of this application simultaneously utilizes the main features in the ultra-high-resolution data and the edge features in the edge intensity data. By fusing the main features and edge features, the feature expression ability of the model is improved, and the morphology and edges of the farmland can be accurately located and described, achieving efficient and accurate farmland extraction effects, which is beneficial to the implementation of precision agriculture and land resource management. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The drawings and descriptions that constitute part of this application are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. Among them:
[0061] Figure 1 A schematic diagram of a process according to the present invention;
[0062] Figure 2 A structural diagram of the hybrid deep learning model according to the present invention;
[0063] Figure 3 A display image for data preprocessing provided according to some embodiments of the present application;
[0064] Figure 4 A result diagram of semantic segmentation and farmland extraction of pre-processed data based on the present application according to some embodiments of the present application;
[0065] Figure 5 Schematic diagram of the structure according to the present invention. DETAILED DESCRIPTION
[0066] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments. Each example is provided by way of explanation of the present application and does not limit the present application. In fact, it will be clear to those skilled in the art that modifications and variations can be made in the present application without departing from the scope or spirit of the present application. For example, a feature shown or described as part of one embodiment can be used in another embodiment to produce yet another embodiment. Therefore, it is expected that the present application includes such modifications and variations within the scope of the appended claims and their equivalents.
[0067] Exemplary Methods
[0068] like Figure 1 As shown in FIG, the farmland extraction method of ultra-high-resolution SAR images based on a hybrid deep learning model includes:
[0069] Step S101: pre-processing the ultra-high resolution SAR image data and generating edge intensity data;
[0070] Specifically, multi-view processing, filtering and noise reduction, terrain correction and decibel conversion are performed on ultra-high-resolution SAR data to optimize image quality and improve the accuracy of subsequent feature extraction.
[0071] In the embodiments of the present application, an existing product generated using ultra-high-resolution data provided by TerraSAR-X is used as an example for illustration. The resolution of this product is 0.8m×0.8m / pixel, and terrain correction is performed using DEM data provided by ALOS PALSAR.
[0072] Calculate edge strength data based on processed TerraSAR-X data:
[0073] Let I(x,y) be the SAR scattering intensity. In a 5×5 window centered at a pixel, the scattering intensity is approximated as a quadratic polynomial model:
[0074] I(x,y)≈P·v(x,y) T
[0075] Where (x, y) is the origin coordinate of the center of the 5×5 window, P = [A, B, C, D, E, F], v (x, y) = [x 2 ,y 2 ,xy,x,y,1]; the coefficient vector P is solved by the least squares criterion.
[0076] Construct the local design matrix v and observation vector I, then the corresponding minimization P objective function is:
[0077]
[0078] Then the analytical solution of vector P is:
[0079] P=(v T v) -1 v T I
[0080] Based on the fitting model, the first-order and second-order partial derivatives of x and y are obtained:
[0081]
[0082] Since x=0, y=0 at the center of the local window, the gradient amplitude at the center point can be expressed as:
[0083]
[0084] The Hessian matrix is a square matrix that describes the local curvature of a multivariate scalar function. The Hessian matrix is constructed from the second-order partial derivatives at the center, and we can get:
[0085]
[0086] Let K = max(|λ1|,|λ2|) be the local curvature measure to reflect the sensitivity of the image to structural changes at that point, where λ1 and λ2 are the eigenvalues of the matrix.
[0087] The amplitude of image changes and the sudden change of local curvature can better highlight farmland in a complex background. Combining gradient information and local curvature, the specific formula of edge intensity data can be expressed as:
[0088]
[0089] Edge strength reflects the local changes between each pixel and its adjacent pixels. It maintains a low value in homogeneous farmland areas and a high value in areas with complex structures such as buildings, thereby highlighting edge features to a certain extent.
[0090] Figure 3 The result of data preprocessing is shown. The value of each pixel in the edge strength data can reflect the edge strength of the pixel. The input data of the embodiment of the present application is specified to be an image size of 512×512.
[0091] Step S102: Using a local residual module to construct a local feature encoder to extract local features and output feature maps layer by layer;
[0092] The pre-processed ultra-high resolution SAR data and edge strength data are used as the input of the deep learning model, and the local feature encoder is constructed using the local residual module to extract local features. This step is divided into four sub-stages, and the feature map output by each sub-stage is named and
[0093] In an embodiment of the present application, the number of local residual modules included in each stage is configured as {2, 2, 2, 2}.
[0094] The local residual module is based on the residual block, but the second traditional convolutional layer in the residual block is replaced by a dynamic snake convolution. Dynamic snake convolution is more sensitive to tubular structures and can better adapt to ridges and roads. Its formula is as follows:
[0095] The intermediate output and output of Conv 3×3 is a convolution operation with a convolution kernel size of 3×3, DSConv is a dynamic snake convolution operation, and Relu is the activation function.
[0096] In the embodiment of the present application, the local feature extraction stage of feature coding and The feature map sizes are divided into 64×256×256, 128×128×128, 256×64×64, and 512×32×32.
[0097] Step S103: Modeling the global relationship of features with linear complexity by introducing the VSS module of VMamba;
[0098] Specifically, local features are mapped to a higher-dimensional semantic space through Patch Embedding. In the embodiment of the present application, the feature dimension is mapped from 512 to 738. And on this basis, long-distance dependencies are established. This stage is also divided into four sub-stages, each of which uses several VSS modules for processing to keep the feature map resolution unchanged. In the example of the present application, the number of VSS modules contained in each stage is configured as {1,1,1,1}.
[0099] Two-dimensional selective scanning is the core mechanism of the VSS module to achieve global relationship modeling with O(n) complexity. Its workflow can be summarized into three main steps: expansion, selection, and reorganization. First, the two-dimensional feature map is expanded into independent one-dimensional sequences along four diagonal directions, covering the contextual association paths of all pixels. The sequence in each direction is input into the parameter-adaptive S6 module, which dynamically adjusts the hidden state weights through a gating mechanism to selectively enhance important features and suppress redundant information. Finally, the sequences processed in the four directions are folded back into two-dimensional space along the original path. The feature responses of different scanning directions are retained through weighted aggregation, and the enhanced feature map that integrates global long-range dependencies is finally output. Its formula is as follows:
[0100] x v =expand(x,v)
[0101]
[0102] Among them, x and The input and output of the two-dimensional selection scan; v represents the four scanning directions, v∈{1,2,3,4}; expand represents the expansion operation, S6 represents the S6 module processing, and merge is the sequence merge operation.
[0103] In the embodiment of the present application, the feature map size of the global feature extraction stage of feature encoding is 32×32×738.
[0104] Step S104: reshape the feature map size of the output result of global feature modeling, and perform conventional convolution to complete feature encoder feature extraction;
[0105] Specifically, after the fourth sub-stage of global feature extraction, the deep semantic features are swapped in the channel dimension. Subsequently, a 3×3 convolution optimizes the feature representation, and the encoding stage ends. The formula is as follows:
[0106]
[0107] Among them, x is the output result of the fourth sub-stage of global feature extraction, Input for the decoding stage; Conv 3×3 It is a convolution operation with a convolution kernel size of 3×3, and reverse is a feature dimension adjustment operation.
[0108] In the embodiment of this application, The feature map size is 512×32×32 and Consistent size.
[0109] like Figure 2 As shown, this is a structural diagram of a deep learning model for farmland extraction based on a serial hybrid structure provided in an embodiment of the present application.
[0110] Step S105: Using the shallow-deep feature fusion module, local shallow features are fused with deep semantic features to complete skip connections and provide semantic information for upsampling.
[0111] Specifically, define the decoding output as and Respectively
[0112] The feature map size corresponds to .
[0113] The shallow-deep feature fusion module concatenates the feature representations from the encoder and decoder along the channel dimension, thereby introducing richer multi-scale contextual information. The concatenated features are then fed into the VSS module for semantic information fusion. Next, by introducing learnable coefficients α and β, the shallow features, deep features, and their fusion are adaptively weighted, improving the model's stability and adaptability under different spatial structures. The formula is as follows:
[0114]
[0115] Among them, F iThe output result of the deep-shallow feature module, i∈{1,2,3,4}, concat indicates that the channel is a splicing operation, V indicates VSS module processing, and α and β are learnable coefficients.
[0116] Step S106: Use transposed convolution to upsample to the input image size to obtain the extraction result and perform accuracy evaluation;
[0117] Specifically, for the output of the shallow-deep feature fusion module, transposed convolution upsampling is used to gradually restore the feature map size, and finally restore it to the input image size. The calculation formula is as follows:
[0118]
[0119] i∈{1,2,3,4},F i Output result of deep-shallow feature module; BN stands for batch normalization operation, TransConv 3×3 It is a transposed convolution operation with a convolution kernel size of 3×3, and Relu is the activation function. 4 Obtained by upsampling That is, the image after semantic segmentation of farmland.
[0120] The accuracy of the hybrid deep learning model was evaluated using both unlabeled and labeled images from the test set. The unlabeled images were directly captured from preprocessed ultra-high-resolution SAR images with edge strength information, while the labeled images were obtained by labeling the unlabeled images according to preset classification labels. The accuracy evaluation metrics used were precision, recall, F1, and IoU, which evaluated the model's ability to extract farmland. The calculation formula is as follows:
[0121]
[0122] Among them, TP, FP, FN and TN represent the number of true positive, false positive, false negative and true negative pixels in the prediction results, respectively.
[0123] Figure 4 The extraction results of this application example based on a hybrid deep learning model are presented. The results are compared with human visual interpretation and are consistent with real-world conditions. The extraction accuracy is high, and the morphological description is accurate. Compared with existing methods, this method significantly improves extraction accuracy.
[0124] Exemplary Systems
[0125] Figure 5This is a schematic diagram of the structure of a system for extracting farmland from ultra-high-resolution SAR images based on a hybrid deep learning model, according to an embodiment of the present application. The system comprises: a data processing unit configured to segment preprocessed high-resolution SAR images and calculate and derive edge intensity images; a feature encoding unit configured to extract features from input data using a hybrid feature encoder based on CNN-VMamba to obtain features of data at different scales; a feature decoding unit configured to perform semantic segmentation on the input data based on fused features, restore the image size through upsampling, and obtain an extraction result after semantic segmentation; and a feature decoding unit configured to perform semantic segmentation on the input data based on fused features, restore the image size through upsampling, and obtain an extraction result after semantic segmentation.
[0126] The ultra-high-resolution SAR image farmland extraction system based on a hybrid deep learning model provided in the embodiments of the present application can implement any of the above-mentioned ultra-high-resolution SAR image farmland extraction steps and processes based on a hybrid deep learning model, and achieve the same technical effects, which will not be repeated here one by one.
[0127] Exemplary devices
[0128] The present application provides an electronic device, including a storage medium and a processor, the processor being suitable for executing various programs; and a memory being used to store multiple programs; characterized in that when the memory executes the program on the processor, the method for extracting farmland from ultra-high-resolution SAR images based on a hybrid deep learning model is implemented.
[0129] Since the steps of farmland extraction from ultra-high-resolution SAR images based on the hybrid deep learning model have been introduced in detail in the specific implementation method example, they will not be described in detail here.
[0130] The processor includes a central processing unit (CPU), a network processor (NP), etc., and may also be a digital signal processor, an application-specific integrated circuit, an off-the-shelf programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. A general-purpose processor may be a microprocessor or any conventional processor.
[0131] The processor can be specifically configured as follows: preprocessing ultra-high-resolution SAR image data and generating its edge strength data; using the local residual module to form a local feature encoder to extract local features, and output feature maps layer by layer; introducing VMamba's VSS module to model the global relationship of features with linear complexity; reshaping the output feature map size of the global feature modeling, and performing conventional convolution to complete feature encoder feature extraction; through the shallow-deep feature fusion module, fusing local shallow features with deep semantic features, completing jump connections, and providing semantic information for upsampling; using transposed convolution to upsample to the input image size, obtaining the extraction result, and performing accuracy evaluation.
[0132] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0133] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code originally stored in a remote recording medium or a non-temporary machine storage medium downloaded via a network and to be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method for extracting farmland from ultra-high-resolution SAR images based on a hybrid deep learning model described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0134] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.
[0135] It should be noted that the various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from the other embodiments. In particular, the device and system embodiments are described briefly because they are generally similar to the method embodiments. For relevant parts, refer to the description of the method embodiments.
[0136] The device and system embodiments described above are merely illustrative. Units not described as separate may or may not be physically separate, and units not described as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiments. Persons of ordinary skill in the art will be able to understand and implement the present embodiments without inventive effort.
[0137] The foregoing description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for extracting farmland from ultra-high-resolution SAR images based on a hybrid deep learning model, characterized by: The steps include: Step S101: preprocessing ultra-high resolution SAR image data and generating edge intensity data; Step S102: Using a local residual module to construct a local feature encoder to extract local features and output feature maps layer by layer; Step S103: Modeling the global relationship of features with linear complexity by introducing the VSS module of VMamba; Step S104: reshape the feature map size of the output result of global feature modeling, and perform conventional convolution to complete feature encoder feature extraction; Step S105: Using the shallow-deep feature fusion module, local shallow features are fused with deep semantic features to complete skip connections and provide semantic information for upsampling. Step S106: Use transposed convolution to upsample to the input image size, obtain the extraction result, and perform accuracy evaluation.
2. The method for extracting farmland from ultra-high resolution SAR images based on a hybrid deep learning model according to claim 1, characterized in that: The step S101 is specifically as follows: Multi-view processing, filtering and noise reduction, terrain correction and decibel conversion are performed on ultra-high resolution SAR data to optimize image quality and improve the accuracy of subsequent feature extraction. Edge intensity data is then calculated based on the processed ultra-high resolution SAR data: Let I(x,y) be the SAR scattering intensity. In a window of size N×N centered at a certain pixel, the scattering intensity is approximated as a quadratic polynomial model: I(x,y)≈P·v(x,y) T Where (x, y) is the coordinate of the origin of the N×N window center, P = [A, B, C, D, E, F], v (x, y) = [x 2 ,y 2 ,xy,x,y,1]; the coefficient vector P is solved by the least squares criterion; Construct the local design matrix v and observation vector I, then the corresponding minimization P objective function is: Then the analytical solution of vector P is: P=(v T v) -1 v T I Based on the fitting model, the first-order and second-order partial derivatives of x and y are obtained: Since x=0, y=0 at the center of the local window, the gradient amplitude at the center point can be expressed as: The Hessian matrix is a square matrix that describes the local curvature of a multivariate scalar function. The Hessian matrix is constructed from the second-order partial derivatives at the center, and we can get: Let K = max(|λ1|,|λ2|) be the local curvature measure to reflect the sensitivity of the image to structural changes at that point, where λ1 and λ2 are the eigenvalues of the matrix; Combining gradient information and local curvature, the specific formula for edge strength data can be expressed as:
3. The method for extracting farmland from ultra-high resolution SAR images based on a hybrid deep learning model according to claim 1, characterized in that: The step S102 is specifically as follows: The pre-processed ultra-high resolution SAR data and edge strength data are used as the input of the deep learning model, and the local feature encoder is constructed using the local residual module to extract local features. This step is divided into four sub-stages, each stage contains several local residual modules, and the feature map output by each sub-stage is named and The local residual module is based on the residual block, but the second traditional convolutional layer in the residual block is replaced by a dynamic snake convolution. The dynamic snake convolution is more sensitive to tubular structures and can better adapt to ridges and roads. Its formula is as follows: The intermediate output and output of Conv 3×3 It is a convolution operation with a convolution kernel size of 3×3, DSConv is a dynamic snake convolution operation, and Relu is the activation function.
4. The method for extracting farmland from ultra-high resolution SAR images based on a hybrid deep learning model according to claim 1, characterized in that: The step S103 is specifically as follows: After patch embedding, local features are mapped to a higher-dimensional semantic space, and long-distance dependencies are established on this basis. This stage is also divided into four sub-stages, each of which uses several VSS modules for processing to maintain the resolution of the feature map. Two-dimensional selective scanning is the core mechanism of the VSS module for achieving global relationship modeling with O(n) complexity. Its workflow can be summarized into three main steps: expansion, selection, and reorganization. First, the two-dimensional feature map is expanded into independent one-dimensional sequences along four diagonal directions, covering the contextual association paths of all pixels. The sequence in each direction is input into the parameter-adaptive S6 module, which dynamically adjusts the hidden state weights through a gating mechanism to selectively enhance important features and suppress redundant information. Finally, the sequences processed in the four directions are reversely folded back into two-dimensional space along the original path. The feature responses of different scanning directions are retained through weighted aggregation, and the final output is an enhanced feature map that integrates global long-range dependencies. The formula is as follows: x v =expand(x,v) Among them, x and The input and output of the two-dimensional selection scan; v represents the four scanning directions, v∈{1,2,3,4}; expand represents the expansion operation, S6 represents the S6 module processing, and merge is the sequence merge operation.
5. The method for extracting farmland from ultra-high resolution SAR images based on a hybrid deep learning model according to claim 1, wherein: The step S104 is specifically as follows: After the fourth sub-stage of global feature extraction, the deep semantic features are swapped in the channel dimension. Subsequently, a 3×3 convolution is performed to optimize the feature representation, and the encoding stage ends. The formula is as follows: Among them, x is the output result of the fourth sub-stage of global feature extraction, Input for the decoding stage; Conv 3×3 It is a convolution operation with a convolution kernel size of 3×3, and reverse is a feature dimension adjustment operation.
6. The method for extracting farmland from ultra-high resolution SAR images based on a hybrid deep learning model according to claim 1, characterized in that: The step S105 is specifically as follows: Define the decoded output as and The shallow-deep feature fusion module concatenates the feature representations from the encoder and decoder in the channel dimension, thereby introducing richer multi-scale contextual information. The concatenated features are then fed into the VSS module for semantic information fusion. Next, by introducing learnable coefficients α and β, the shallow features, deep features, and their fusion are adaptively weighted, improving the model's stability and adaptability under different spatial structures. The formula is as follows: Among them, F i The output result of the deep-shallow feature module, i∈{1,2,3,4}, concat indicates that the channel is a splicing operation, V indicates VSS module processing, and α and β are learnable coefficients.
7. The method for extracting farmland from ultra-high resolution SAR images based on a hybrid deep learning model according to claim 1, characterized in that: The step S106 is specifically as follows: For the output of the shallow-deep feature fusion module, transposed convolution upsampling is used to gradually restore the feature map size, and finally restore it to the input image size; the calculation formula is as follows: i∈{1,2,3,4},F i Output result of deep-shallow feature module; BN stands for batch normalization operation, TransConv 3×3 It is a transposed convolution operation with a convolution kernel size of 3×3, and Relu is the activation function; the final F 4 Obtained by upsampling That is, the image after semantic segmentation of farmland; The accuracy of the hybrid deep learning model was evaluated using unlabeled and labeled images from the test set. The unlabeled images were directly captured from preprocessed ultra-high-resolution SAR images and edge strength information, while the labeled images were obtained by labeling the unlabeled images according to preset classification labels. The accuracy evaluation indicators were precision, recall, F1, and IoU, which evaluated the model's farmland extraction ability. The calculation formula is as follows: Among them, TP, FP, FN and TN represent the number of true positive, false positive, false negative and true negative pixels in the prediction results, respectively.
8. A farmland extraction system from ultra-high resolution SAR images based on a hybrid deep learning model, characterized by: include: The data processing unit is configured to: segment the pre-processed ultra-high resolution SAR image and calculate and derive an edge intensity image; Obtaining target training set unlabeled images and target training set labeled images, wherein the target training set labeled images are obtained by labeling the target training set unlabeled images with categories according to preset classification labels; The feature encoding unit is configured to use a hybrid feature encoder based on CNN-VMamba to extract features from the input data and obtain features of data at different scales; The feature fusion unit is configured to: fuse shallow encoding features and deep decoding features based on a two-dimensional selective scanning mechanism, and organically combine the three through dynamic weighting to improve feature expression capabilities; The feature decoding unit is configured to: perform semantic segmentation on the input data based on the fused features, restore the image size through upsampling, and obtain the extraction result after semantic segmentation.
9. A storage medium storing a plurality of programs, characterized in that: The program application is loaded and executed by a processor to implement the ultra-high resolution SAR image farmland extraction method based on a hybrid deep learning model as described in any one of claims 1-7.
10. An electronic device comprising a storage medium and a processor; the processor being adapted to execute various programs; and a memory being adapted to store a plurality of programs; characterized in that: When the memory executes the program on the processor, the method for extracting farmland from ultra-high-resolution SAR images based on a hybrid deep learning model as described in any one of claims 1 to 7 is implemented.