Multi-scale feature fusion-based agricultural remote sensing farmland segmentation model construction method
Patent Information
- Application Number
- CN202611099171.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-04
AI Technical Summary
(1)类别严重不平衡:背景(裸土、道路、建筑、水体)面积占比与作物占比不平衡,模型易倾向于将耕地误分为背景,导致耕地召回率偏低
[0018]The beneficial effects of the method described in this invention are as follows: by using a two-stage serial structure and a priori constraint mechanism for cultivated land, the interference of background categories on crop classification tasks is effectively reduced, and the stability of crop classification in complex agricultural scenarios is improved; by using a multi-temporal remote sensing image fusion mechanism, the ability to express temporal spectral differences at different crop growth stages is enhanced, and the separability of crop categories is improved; by using boundary detection branches and auxiliary supervision branches, the recovery ability of slender structures such as farmland boundaries, field ridges, and ditches is enhanced, and the continuity of boundaries is improved; by using a separate jump fusion mechanism, the model's response ability to key crop areas and important spatial locations is improved.
Smart Images

Figure CN122695243A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of agricultural segmentation model construction, specifically involving a method for constructing an agricultural remote sensing farmland segmentation model through multi-scale feature fusion. Background Technology
[0002] In high-resolution agricultural remote sensing image processing, farmland extraction and crop classification models face the following key technical bottlenecks: (1) Severe category imbalance: The proportion of background (bare soil, roads, buildings, water bodies) area is unbalanced with the proportion of crops. The model tends to misclassify farmland as background, resulting in a low farmland recall rate.
[0003] (2) Blurred boundaries of small linear features: Slender structures such as field roads and ditches are prone to feature loss during multiple downsampling processes, resulting in problems such as broken segments and jagged boundaries in the segmentation results.
[0004] (3) Insufficient spectral separability of crops: In single-temporal remote sensing images, different crops (taking corn, rice and soybean as examples in this paper) have strong overlap in key bands such as near-infrared, making it difficult to effectively distinguish them by relying solely on single-temporal spectral information.
[0005] (4) The skip connection fusion mechanism is simple. Existing U-Net-like networks usually use simple splicing or element-wise addition to fuse encoder and decoder features, which lacks the ability to adaptively select shallow texture information and deep semantic information, and is difficult to effectively recover fine-grained boundary structures in complex agricultural scenes.
[0006] (5) Insufficient cross-regional generalization ability. Remote sensing images from different regions are affected by sensors, climate and surface conditions, and their spectral distributions vary significantly. Traditional normalization methods may destroy the physical consistency of the spectrum and reduce the model's transferability and generalization performance.
[0007] Therefore, there is an urgent need to propose an agricultural remote sensing segmentation model that can take into account the accuracy of farmland extraction, the ability to classify crops in a fine-grained manner, the ability to preserve boundaries, and the generalization performance across regions. Summary of the Invention
[0008] To address the aforementioned technical problems, this invention provides a method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion. This method can construct an agricultural remote sensing segmentation model that can balance farmland extraction accuracy, crop fine-grained classification ability, boundary preservation ability, and cross-regional generalization performance.
[0009] The method is specifically as follows: S1. Input layer construction: Select farmland remote sensing images from multiple time phases, select several bands for each time phase, and stitch all band images along the channel dimension to form a channel fusion image, which is used as the unified input of the model. S2. First stage network construction: An encoder-decoder network architecture is adopted. The data processing flow is as follows: the channel fusion image is passed through the encoder for 4 levels of downsampling, the bottleneck layer and the decoder for 4 levels of upsampling in sequence to output a farmland probability map. The farmland probability map is thresholded to obtain a farmland prior mask. A boundary detection branch is added to the encoder and an auxiliary supervision branch is added to the decoder. S3. Second-stage network construction: The encoder-decoder network architecture is shared with the first-stage network, but no boundary detection branch and auxiliary supervision branch are introduced; The data processing flow is as follows: The channel fusion image and the farmland prior mask are input together, and then pass through the encoder's 4-level downsampling, the bottleneck layer and the decoder's 4-level upsampling in sequence to output the farmland segmentation result; S4. Segmentation Model Training: Training is conducted in two stages. In the first stage, the network's loss function is composed of a weighted sum of the main loss, deep supervision loss, and boundary loss. In the second stage, the network's loss function adopts the main loss function.
[0010] Furthermore, the encoder consists of an initial convolutional layer and a 4-level downsampling structure. Each downsampling level sequentially performs 2×2 max pooling, two 3×3 convolutions, and a CBAM attention mechanism. The CBAM attention mechanism is composed of channel attention and spatial attention concatenated: channel attention performs global average pooling and global max pooling on the input features respectively, and after passing through a shared multilayer perceptron, the results are summed and then activated by a sigmoid operation to generate channel weights; spatial attention performs average pooling and max pooling along the channel axis respectively, and after concatenation, the results are convolved to generate spatial weights, which are then recalibrated element-wise on the channel-weighted features.
[0011] Furthermore, the output of the 4th level downsampled by the encoder constitutes the bottleneck layer. The bottleneck layer aggregates high-level semantic information of the entire image. Each feature point corresponds to the largest receptive field range in the original image, serving as a feature compression and transition hub from the encoder to the decoder.
[0012] Furthermore, the decoder consists of a 4-level upsampling structure, which restores spatial resolution step by step from the bottleneck layer. Each level first performs a 2x upsampling through bilinear interpolation, then performs a lateral connection with the corresponding level features of the encoder, and then refines the features through two 3×3 convolutions and the CBAM attention mechanism.
[0013] Furthermore, a self-designed split-skip fusion mechanism is introduced in the first to third stages of decoder upsampling, and deployed at the skip connection position of the decoder. The split-skip fusion mechanism includes encoder branch and decoder branch. The encoder branch performs 1×1 pointwise convolution on encoder skip features, and the decoder branch performs 3×3 convolution on decoder upsampled features. After the outputs of the two branches are concatenated along the channel dimension, cross-branch feature interaction and nonlinear fusion are achieved through two consecutive 3×3 convolutions, and finally input into the CBAM attention mechanism.
[0014] Furthermore, a separate boundary detection branch is derived from the initial convolutional layer features of the encoder. This boundary detection branch contains three modules, each of which specifically includes: The first module consists of a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation function layer. The second module consists of 3×3 convolutional layers, batch normalization layers, and ReLU activation function layers. The third module consists of 1×1 convolutional layers and Sigmoid activation function layers; The boundary detection branch outputs a boundary probability map with the same resolution as the channel fused image.
[0015] Furthermore, auxiliary supervision branches are introduced in the second and third upsampling stages of the decoder. Each auxiliary supervision branch specifically includes: a 3×3 convolutional layer, a batch normalization layer, a ReLU activation function layer, a 1×1 convolutional layer, and a bilinear interpolation upsampling layer. The output of each auxiliary supervision branch is a semantic segmentation prediction map with the same resolution and the same number of categories as the main output of the decoder.
[0016] Furthermore, the boundary loss is specifically calculated as follows: the weighted binary cross-entropy of the boundary probability map is used as the boundary loss function. .
[0017] Furthermore, the deep supervision loss is specifically calculated as follows: cross-entropy loss and Dice loss are calculated on the semantic segmentation prediction graphs output by the second-level auxiliary supervision branch upsampled by the decoder, and then weighted and fused to obtain the loss function. The cross-entropy loss and Dice loss are calculated on the semantic segmentation prediction graphs output by the upsampled third-level auxiliary supervision branch of the decoder, and then weighted and fused to obtain the loss function. loss function and loss function Together they constitute the loss of deep oversight.
[0018] The beneficial effects of the method described in this invention are as follows: by using a two-stage serial structure and a priori constraint mechanism for cultivated land, the interference of background categories on crop classification tasks is effectively reduced, and the stability of crop classification in complex agricultural scenarios is improved; by using a multi-temporal remote sensing image fusion mechanism, the ability to express temporal spectral differences at different crop growth stages is enhanced, and the separability of crop categories is improved; by using boundary detection branches and auxiliary supervision branches, the recovery ability of slender structures such as farmland boundaries, field ridges, and ditches is enhanced, and the continuity of boundaries is improved; by using a separate jump fusion mechanism, the model's response ability to key crop areas and important spatial locations is improved. Attached Figure Description
[0019] Figure 1 This is a flowchart of the network workflow for Stage 1 in this embodiment of the invention; Figure 2 This is a flowchart of the network workflow for Stage 2 in this embodiment of the invention. Detailed Implementation
[0020] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0021] Example 1 This embodiment provides a method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion. The method adopts an encoder-decoder construction form and constructs a collaborative segmentation network model with an overall framework of "two-stage concatenation + boundary enhancement".
[0022] The first stage (Stage 1) performs binary classification of cultivated land / non-cultivated land, outputting a priori mask for cultivated land. The second stage (Stage 2) performs fine-grained crop classification under the constraint of the cultivated land mask, outputting classification results for corn, rice, and soybeans. Both stages share the encoder-decoder architecture. Stage 1 introduces an additional boundary detection branch and a deep supervision mechanism, while Stage 2 is simplified to a pure semantic segmentation structure to focus on spectral differences between crop categories. The specific workflows of Stage 1 and Stage 2 are as follows: Figure 1 and Figure 2 As shown.
[0023] 1. Input layer design The input is multi-temporal, multi-channel fused remote sensing imagery. This embodiment acquires Sentinel-2 multispectral images of the same region in June and August. For each temporal phase, four bands (blue (B), green (G), red (R), and near-infrared (NIR)) are extracted and stitched together along the channel dimension to form an 8-channel fused image, which serves as the unified input to the network. The number of temporal phases and the band combination can be adjusted according to the agricultural planting structure and remote sensing data sources of different regions.
[0024] 2. Stage 1 Design: Cultivated Land / Non-Cultivated Land Classification The overall data flow of Stage 1 is as follows: 8-channel fused image → encoder 4-level downsampling → bottleneck layer → decoder 4-level upsampling → farmland probability map; at the same time, the boundary detection branch is derived from the shallow layer of the encoder to output the boundary probability map, and the deep supervision auxiliary branch is derived from the middle layer of the decoder to participate in the loss calculation, generating a farmland prior mask to constrain Stage 2.
[0025] 2.1 Encoder The encoder consists of an initial convolutional layer and a 4-level downsampling structure. Each downsampling level sequentially performs 2×2 max pooling, two 3×3 convolutions (including BatchNorm + ReLU), and the CBAM attention mechanism. The CBAM (Convolutional Block Attention Module) is composed of channel attention and spatial attention concatenated: channel attention performs global average pooling and global max pooling on the input features, sums them after passing through a shared multilayer perceptron, and generates channel weights after a sigmoid activation operation; spatial attention performs average pooling and max pooling along the channel axes, concatenates them, and generates spatial weights through a 7×7 convolution, recalibrating the channel-weighted features element-wise.
[0026] The structure at each level is as follows: Initial convolutional layer: two 3×3 convolutions, 64 output channels, feature map size remains unchanged at 512×512; Downsampling Level 1: 2×2 max pooling (size halved to 256×256) → two 3×3 convolutions → CBAM, output channel count 128; Downsampling Level 2: 2×2 max pooling (size halved to 128×128) → two 3×3 convolutions → CBAM, output channel count 256; Downsampling Level 3: 2×2 max pooling (size halved to 64×64) → two 3×3 convolutions → CBAM, output channel count 512; Downsampling level 4 (bottleneck layer): 2×2 max pooling (size halved to 32×32) → two 3×3 convolutions → CBAM, output channel count 1024.
[0027] The number of encoder channels changed as follows: 64 → 128 → 256 → 512 → 1024.
[0028] 2.2 Bottleneck Layer The output of the fourth downsampling stage of the encoder constitutes the bottleneck layer, with 1024 channels and a feature map size of 32×32. The bottleneck layer aggregates high-level semantic information from the entire image, with each feature point corresponding to the largest receptive field in the original image, serving as a feature compression and transition hub between the encoder and decoder.
[0029] 2.3 Decoder The decoder consists of a 4-level upsampling structure that recovers spatial resolution step by step from the bottleneck layer. Each level first performs 2× upsampling through bilinear interpolation, then performs lateral connection with the corresponding level features of the encoder, and finally refines the features through two 3×3 convolutions and the CBAM attention module.
[0030] Unlike the direct channel splicing of standard U-Net, this invention introduces a self-designed Separated Skip Fusion (SSF) mechanism in the first three sampling stages of the decoder. Deployed at the skip connection positions in the decoder, it achieves effective collaborative expression of shallow spatial details and deep semantic information through three stages: differentiated projection, collaborative fusion, and adaptive recalibration. This splits the single splicing operation of encoding and decoding features into two independent heterogeneous branches, which are preprocessed separately and then merged for refinement. The SSF module structure is as follows: Encoder branch (conv_skip): Performs 1×1 pointwise convolution on the encoder skip features to complete channel projection while keeping the spatial topology of the feature map unchanged, thus preserving the spatial location information of long and thin linear features such as field ridges and ditches to the maximum extent. Decoder branch (conv_feat): Performs 3×3 convolution on the upsampled features of the decoder to enhance the deep semantic discrimination ability by expanding the context receptive field; The outputs of the two branches are concatenated along the channel dimension, and then cross-branch feature interaction and nonlinear fusion are achieved through two consecutive 3×3 convolutions; Finally, the CBAM attention module is input to perform joint adaptive recalibration of the channel and spatial dimensions of the fusion result.
[0031] The decoder's structure at each stage is as follows: Upsampling Level 1: Bilinear interpolation 2×upsampling → SSF fusion (bridged with the 512-channel feature of the coding downsampling Level 3) → Two 3×3 convolutions → CBAM, output channel number 512, feature map size 64×64; Upsampling Level 2: Bilinear interpolation 2×upsampling → SSF fusion (bridged with the 256-channel feature of the encoding downsampling Level 2) → Two 3×3 convolutions → CBAM, output channel number 256, feature map size 128×128; Upsampling Level 3: Bilinear interpolation 2×upsampling → SSF fusion (bridged with the 128-channel feature of the coding downsampling Level 1) → Two 3×3 convolutions → CBAM, output channel number 128, feature map size 256×256; Upsampling Level 4: Bilinear interpolation 2×upsampling → concatenation with the feature channels of the initial convolutional layer → two 3×3 convolutions → CBAM, output channel number 64, feature map size 512×512.
[0032] The number of decoder channels changed as follows: 512 → 256 → 128 → 64.
[0033] SSF skip connections are only applied to the first three levels of decoding upsampling, while the shallowest layer (level 4 upsampling) uses standard channel splicing. During the decoder's progressive upsampling process, the SSF module effectively preserves the spatial details of slender linear features such as farmland boundaries, field roads, and irrigation ditches, alleviating the problem of shallow textures in the encoder being overwhelmed by strong semantic features in the decoder in traditional skip connections, and improving the continuity of farmland boundaries and the accuracy of fine-grained feature structure restoration.
[0034] 2.4 Boundary Detection Branch To address the challenges of easily broken farmland boundaries and the difficulty in preserving slender structures such as field ridges and ditches, a separate boundary detection branch is derived from the initial convolutional layer features (64 channels, original resolution 512×512) of the encoder. The boundary detection head sequentially includes: First round: 3×3 convolution (Conv) → Batch Normalization (BN) → ReLU activation function; Second round: 3×3 convolution (Conv) → Batch Normalization (BN) → ReLU activation function; 1×1 convolution (Conv) → Sigmoid activation function; The final output is a boundary probability map with the same resolution as the original image. During the training phase, the Sobel operator is used to automatically generate ground truth boundary maps from the labeled images, and binary cross-entropy loss is used to supervise the predicted boundaries, enhancing the model's ability to perceive boundary structures such as field ridges, roads, and ditches.
[0035] 2.5 In-depth supervision mechanism In deep segmentation networks, the loss signal, when backpropagating from the network's end to the shallow layers, experiences a significant amplitude decay after multiple gradient multiplications, leading to slow parameter updates and insufficient feature learning in the shallow encoder and intermediate decoder layers. To address this, auxiliary supervision branches are introduced at the second upsampling stage (256 channels) and the third upsampling stage (128 channels) of the decoder, injecting additional training signals in the middle of the decoding path. This effectively alleviates the vanishing gradient problem in deep network training and improves gradient propagation stability.
[0036] The auxiliary branch first maps the intermediate layer features to the number of class channels using a 1×1 convolution, then upsamples them to the original input resolution (512×512), forming the intermediate layer prediction results and participating in the loss function calculation. The intermediate layer prediction results are aux1 and aux2; aux1: Intermediate layer features of decoder Up2 level (256 channels, 128×128 resolution) → mapped to the number of class channels via aux_head1 (3×3Conv → BN → ReLU → 1×1Conv) → bilinear interpolation upsampling to 512×512 → output a semantic segmentation prediction map with the same resolution and number of classes as the main output; aux2: Intermediate layer features of decoder Up3 level (128 channels, 256×256 resolution) → processed in the same way by aux_head2 → upsampled to 512×512 → output semantic segmentation prediction map of the same resolution; Another key role of deep supervision is to force the intermediate layer features of the decoder to have complete semantic discrimination ability in the early stage of training, rather than waiting until the last layer of the network to complete the "semantic enlightenment" - the higher the quality of the intermediate layer features of the decoder, the more sufficient the information interaction between the boundary detection branch derived from the initial convolutional layer of the encoder and the deep semantic features, the boundary detection and semantic segmentation form a positive synergistic effect, thereby improving the overall segmentation accuracy in complex farmland scenes.
[0037] 2.6 Output: Farmland probability map and farmland prior mask At the decoder end, 64-channel features are mapped to 2-channel output via 1×1 convolution. The posterior probability of cultivated land / non-cultivated land is calculated pixel-by-pixel using the Softmax function, and the class with the highest probability is selected as the cultivated land prediction result. During training, the model's forward propagation returns a quadruple (main output, auxiliary output 1, auxiliary output 2, boundary output), where auxiliary outputs 1 and 2 correspond to the deep supervised predictions of the second and third levels of decoding upsampling, respectively, and the boundary output corresponds to the boundary detection branch. After thresholding, a binary cultivated land prior mask is generated and passed to Stage 2 as a spatial constraint—crop category discrimination is performed only within the predicted cultivated land area, while non-cultivated land areas are directly masked.
[0038] 3. Stage 2 Design: Crop Fine-Grain Classification Stage 2 shares the encoder-decoder infrastructure and CBAM and SSF modules with Stage 1, but does not introduce boundary detection branches or deep supervision mechanisms; parameters are learned independently. The data flow is as follows: 8-channel fused image + farmland prior mask spatial constraints → encoder level 4 downsampling → bottleneck layer → decoder level 4 upsampling → farmland segmentation results.
[0039] 3.1 Input and Spatial Prior Constraints Stage 2 receives two inputs: The main input is the same 8-channel multi-temporal fusion remote sensing image as Stage 1; The spatial prior constraint is the farmland prior mask generated by Stage1.
[0040] During the inference phase, a priori mask of cultivated land is used to shield non-cultivated land areas—only pixels within the cultivated land mask area are classified as corn, rice, or soybeans, while non-cultivated land areas are forcibly set as background. This spatial prior constraint mechanism alleviates the problem of an excessively large proportion of background categories, reduces the false detection rate of background, and improves the stability of crop classification.
[0041] 3.2 Encoder The stage 2 encoder and decoder have the same hierarchical structure, channel number changes, and size sequence as stage 1: The difference between Stage 2 and Stage 1 is that Stage 2 does not include a boundary detection branch or a deep supervision mechanism: the intermediate layers of the decoder no longer generate auxiliary supervision branches, and the initial convolutional layers of the encoder are not connected to the boundary detection head. During training, the forward propagation only returns a single logits tensor, without calculating auxiliary loss and boundary loss, allowing the model to focus on the spectral and texture differences between crop categories, while reducing computational overhead.
[0042] Stage 2 retains the SSF multi-scale fusion module because there are still thin dividing lines between different crop plots within the farmland (such as the transition zone between corn and rice). The 1×1 convolution of the encoder branch of SSF can effectively preserve these fine-grained spatial information and avoid being overwhelmed by deep semantic features.
[0043] 3.3 Output: Crop classification results At the decoder end, 64-channel features are mapped to 3-channel outputs (corresponding to corn, rice, and soybean) using 1×1 convolution. Softmax is then used to calculate the posterior probability of each category, and the category with the highest probability is taken as the final predicted label. During the inference stage, the classification results are overlaid with a prior mask for cultivated land—the area outside the mask is forcibly set as background (category 0), while the area inside the mask retains the crop classification results, generating a complete crop classification map.
[0044] 3.4 Loss Function and Training Strategy 3.4.1 Stage 1 Loss Function The total loss in Stage 1 is composed of a weighted average of the main loss, the deep supervision loss, and the boundary loss: Main loss: Cross-entropy loss + Dice loss with 2.0 times the weight. The cross-entropy class weights are set to [background=8.0, farmland=1.0], which alleviates the class imbalance problem between farmland and background by significantly increasing the misclassification cost of the background class; the Dice loss skips the background class (skip_bg=True) and only optimizes the Dice coefficient of the farmland region, ignoring the label value 255 (invalid region).
[0045] Deep supervision loss: Auxiliary outputs aux1 and aux2, each calculated with cross-entropy loss + Dice loss (class weights consistent with the main loss), added to the total loss with 0.3 times the weight.
[0046] Boundary loss: Binary cross-entropy (BCE), the Sobel operator automatically generates the boundary truth value, which is added to the total loss with a weight of 0.5.
[0047] 3.4.2 Stage 2 Loss Function Stage 2 only includes the main loss and does not include deep supervision or boundary loss: Cross-entropy loss is applied with equal weights for each category (corn, rice, soybean) [1.0, 1.0, 1.0]. Dice loss includes all categories (skip_bg=False), ignoring label value 255.
[0048] 3.5 Summary of the Two-Phase Transition Mechanism The complete data flow of the two-stage concatenated segmentation network is as follows: Input: June and August Sentinel-2 multispectral images (B, G, R, NIR) are stitched together along the channel dimension to generate an 8-channel fused image; Stage 1 Forward Propagation: 8-channel image → Encoder 4-level downsampling (64→128→256→512→1024) → Bottleneck layer (1024 channels, 32×32) → Decoder 4-level upsampling (512→256→128→64), Decoder upsampling levels 1-3 are skipped via SSF module, Decoder upsampling levels 2 and 3 lead to deep supervision auxiliary branches → Output farmland probability map (2 types); Boundary detection: Encoder initial convolutional layer features (64 channels) → Boundary detection head (two rounds of 3×3 Conv+BN+ReLU → 1×1 Conv → Sigmoid) → Boundary probability map → BCE boundary loss supervision; Farmland prior mask generation: Farmland probability map is segmented by thresholding → binary farmland prior mask, which is then passed to Stage2; Stage 2 Forward Propagation: 8-channel imagery + farmland prior mask spatial constraints → Encoder 4-level downsampling → Bottleneck layer → Decoder 4-level upsampling (including SSF skip connections and CBAM, no-boundary detection branch and deep supervision) → Output crop classification results (3 categories: maize / rice / soybean); Final output: The area outside the farmland mask is forcibly set as the background, while the crop classification results inside the mask are retained → Complete crop classification map.
[0049] Through the above two-stage differentiated design—Stage 1 enhances the accuracy of cultivated land outlines with boundary detection and deep supervision, while Stage 2 focuses on crop spectral differences with a simplified structure—the two work together to achieve high-precision segmentation of cultivated land and crop categories in complex agricultural remote sensing scenarios.
[0050] Example 2 This embodiment further defines Embodiment 1 and provides a further explanation of the construction method in Embodiment 1.
[0051] 1. Further introduction to the boundary detection branch: 1.1 Problem Background and Technical Motivation In the U-Net-based remote sensing semantic segmentation framework, the encoder extracts multi-scale semantic features through progressive downsampling. However, continuous pooling operations in standard encoders lead to irreversible degradation of high-frequency spatial information—especially the structural details of ground feature boundaries and linear targets (field ridges, roads, ditches). After the feature map undergoes four 2x downsampling operations, the linear structures with a width of 1-3 pixels at the original 512×512 resolution degrade to an indistinguishable level in the deep feature map. This problem is particularly prominent in agricultural remote sensing scenarios, as the segmentation accuracy of cultivated land plots directly depends on the accurate depiction of linear features such as field ridges and irrigation canals.
[0052] Traditional U-Net relies on skip connections to pass shallow high-resolution features to the decoder to compensate for the loss of spatial details. However, skip connections are essentially an implicit feature reuse mechanism—shallow features are directly concatenated into the decoder without any explicit structural constraints. Whether the network truly retains boundary information and to what extent it retains it depends entirely on the indirect guidance of gradient backpropagation, lacking measurable and supervised explicit modeling.
[0053] The proposed construction method transforms the implicit expectation of "preserving boundary information" into a supervised explicit learning objective by introducing an independent boundary detection branch in the shallow layer of the encoder, thereby imposing structure-aware constraints on the feature extraction process of the encoder.
[0054] 1.2 Design Principles (a) Shallow feature source selection Select the output of the encoder's first-level convolutional block. The technical basis for the source of boundary-aware features is as follows: Resolution advantage: This level is unsampled, so the feature map retains the original input resolution, resulting in the highest spatial localization accuracy. Linear targets with a width of 1-3 pixels, such as field ridges and ditches, still retain complete spatial resolution at this level. Texture sensitivity: Shallow convolution kernels mainly respond to low-level visual primitives such as local texture gradients, color contrast, and edge orientation. These primitives are precisely the physical basis for constituting the boundaries of ground features. Shortest gradient propagation distance: The boundary supervision signal acts directly on the front end of the encoder, and the gradient does not need to travel through the entire decoder to reach the shallow layer, avoiding the gradient decay problem commonly found in deep backpropagation.
[0055] (b) Boundary-aware branch structure The boundary detection branch consists of four levels: First layer: Local edge response enhancement layer (3×3 convolution + Batch Normalization + ReLU): 3×3 convolution kernels enhance local texture gradients in a directional manner, Batch Normalization stabilizes the feature distribution, and ReLU makes the branches focus only on edge regions with significant gradients by sparsifying the response.
[0056] The second layer: edge feature refinement layer (3×3 convolution + Batch Normalization + ReLU): Based on the enhanced edge response, the edge direction, intensity and continuity are further modeled. The equivalent receptive field of the two cascaded 3×3 convolution layers just covers the width range (1-5 pixels) of typical linear features, providing sufficient discriminative context while ensuring edge localization accuracy.
[0057] The third layer: Channel compression and boundary response mapping layer (1×1 convolution): compresses the multi-channel edge feature map into a single-channel boundary response map to achieve cross-channel information fusion and projection.
[0058] Fourth layer: Probability normalization layer (Sigmoid): Maps the boundary response to the [0, 1] interval, generating a pixel-level boundary probability map. The closer the value is to 1, the higher the confidence that the pixel is located at the boundary of a ground feature.
[0059] (c) Boundary truth generation and loss function During the training phase, the Sobel operator is used to extract the boundary truth map from the semantic label map Y: ; ; ; Semantic label map (annotated data known during training), pixel value 0 = background, 1 = farmland; Sobel horizontal (x-axis) gradient operator and The result of convolution is used to detect the vertical boundary. Sobel vertical (y-axis) gradient operator and The result of convolution is used to detect horizontal boundaries; : The first in the boundary truth graph Line 1 The pixel value in the column, 1 = the pixel is on the boundary of the feature, 0 = not on the boundary; The boundary loss employs weighted binary cross-entropy to mitigate class imbalance between boundary and non-boundary pixels.
[0060] The first branch of boundary detection, the Sigmoid, outputs the... Line 1 Boundary probability values of the column, range
[0061] Total number of pixels in the image; Loss weights for positive samples (boundary pixels); Loss weights for negative samples (non-boundary pixels); : Boundary-aware loss function value, the boundary component in the overall loss; Where positive and negative sample weights , The pixels are allocated inversely based on the actual number of boundary pixels and non-boundary pixels to ensure that the two types of pixels contribute equally to the total loss.
[0062] 1.3 Theoretical Advantages (1) Structural decoupling: The boundary detection branch and the main segmentation branch share the shallow features of the encoder. After training, they are completely removed during the inference stage without increasing the computational overhead of inference.
[0063] (2) Interpretability of explicit constraints: Boundary loss directly measures the network’s positioning accuracy of ground object boundaries. When the model performance degrades, attribution analysis can be performed by separating and observing the changing direction of boundary loss and main task loss.
[0064] (3) Gradient synergy effect: The boundary branch and the main segmentation branch form gradient superposition on the shallow features of the encoder:
[0065] The boundary loss weighting coefficient, For Stage 1, the total loss function is... The main segmentation loss is the sum of the cross-entropy loss of the main output path in Stage 1 and the Dice loss. This synergistic effect achieves dual optimization of semantic consistency and spatial accuracy without increasing inference parameters.
[0066] 2. Further introduction to the auxiliary supervision branch: 2.1 Problem Background and Technical Motivation In the standard U-Net decoder, deep semantic features recover spatial resolution through progressive upsampling. However, upsampling operations (bilinear interpolation or transposed convolution) are essentially smooth interpolation processes based on local neighborhoods, and their ability to reconstruct sharp boundaries is limited. In particular, linear features in deep features occupy very few pixels—when the feature map resolution is 1 / 4 of the original, a 3-pixel-wide linear target occupies only 1 / 16 of the original area—and are easily diluted or blurred during the smoothing process of progressive upsampling.
[0067] The method of this invention imposes explicit constraints on the semantic discrimination quality and spatial structure integrity of intermediate features by introducing an auxiliary prediction branch in the intermediate layer of the decoder.
[0068] 2.2 Design Principles (a) Selection and structure of auxiliary branch levels The module introduces auxiliary supervision branches at two intermediate levels of the decoder: The first auxiliary supervision branch (aux_head1) originates from the u3 level. Its feature map resolution is 1 / 4 of the original input, and it has 256 channels. This level integrates deep semantic information with mid-level structural information. The auxiliary branch is primarily used to constrain the coarse-grained semantic structural integrity. The second auxiliary supervision branch (aux_head2) originates from the u2 level. Its feature map resolution is half that of the original input, and it has 128 channels. This level offers richer spatial details, and the auxiliary branch primarily constrains finer-grained edge sharpness.
[0069] Each auxiliary branch structure is: 3×3 convolution (channel compression) → Batch Normalization → ReLU → 1×1 convolution (mapping to class space) → bilinear interpolation upsampling to the original resolution.
[0070] (b) Loss function design The total loss function is a weighted combination of multiple tasks:
[0071] in Both are semantic segmentation cross-entropy loss, weighted coefficients = =0.3、 =0.5, ensuring that the gradient of the main output path dominates, and the auxiliary branches only provide supplementary structural regularization.
[0072] 2.3 Theoretical Advantages (1) Gradient path shortening: The auxiliary branch directly injects the supervision signal into the middle layer of the decoder, which shortens the effective gradient path received by the shallow encoder from traversing all decoder layers in the standard U-Net to only traversing 1-2 upsampling layers, effectively alleviating the gradient decay problem in long-distance backpropagation.
[0073] (2) Hierarchical structural constraints: Auxiliary branches at different levels focus on structural features of different granularities—aux_head1 constrains coarse-grained semantic structure (overall shape of the plot), and aux_head2 constrains fine-grained edges (continuity and width of field ridges). This hierarchical constraint system from coarse to fine can more systematically guarantee the reconstruction quality of spatial structure than applying a single constraint only to the final output layer.
[0074] (3) Zero overhead during inference: All auxiliary branches are activated only during the training phase and removed directly during the inference phase, reverting to the inference path of the standard U-Net without introducing any additional computational overhead or parameter storage burden. This design paradigm of "constraints during training and simplification during inference" is one of the core technical features of this invention.
[0075] 3. The synergistic mechanism between boundary perception enhancement and assisted reconstruction: The two branches mentioned above form a synergistic enhancement effect of structure awareness during training: Forward information collaboration: The edge response extracted from shallow features by the encoder boundary awareness branch is transmitted to each layer of the decoder through the encoder forward path and skip connections, indirectly improving the attention response strength of the auxiliary branch to the prediction of the boundary region.
[0076] Backward gradient cooperative boundary loss shallow encoder layer and auxiliary loss In the encoder, multiple levels of gradient superposition are formed, collectively enhancing the encoder's ability to represent ground feature boundaries. Positive synergy is achieved when the cosine similarity of each gradient direction is positive; when directions conflict, the weighting coefficients are adjusted accordingly. Ensure that the gradient of the main task dominates to avoid training instability.
[0077] Redundancy and complementarity: The boundary-aware branch constrains texture features at the lowest level of the encoder, while the reconstruction-aided branch constrains semantic features at different levels in the decoder. The two form a complementary structural constraint network at different depths and semantic granularities, so that any local failure of a single constraint can be compensated for by constraints at other levels, thus enhancing the structural robustness of the overall framework.
[0078] Through the aforementioned collaborative mechanism, the construction method of this invention achieves a technical upgrade from "implicit expectation preservation" to "explicit multi-level constraint modeling" for fine-grained linear features, significantly improving the structural reconstruction accuracy of key agricultural features such as field ridges, field roads, and irrigation ditches in the semantic segmentation of remote sensing images.
[0079] 4. Separated Skip Fusion (SSF) Mechanism for Agricultural Linear Features: In deep convolutional networks, shallow features are mainly composed of high-resolution local texture information, forming a visual detail feature space; deep features are mainly composed of low-resolution global semantic information, forming a semantic abstract feature space. Traditional U-Net directly concatenates the shallow features of the encoder and the deep features of the decoder during skip connections, assuming both types of features reside in the same representation space. When there are differences in expression between shallow texture information and deep semantic information, direct fusion can easily lead to feature conflicts: high-frequency texture noise in the shallow layer interferes with semantic discrimination, while low-frequency semantic information in the deep layer may weaken the spatial continuity of slender structures such as field ridges, roads, and ditches. This problem is particularly prominent in agricultural remote sensing scenarios.
[0080] To address this, the present invention proposes a Separated Skip Fusion (SSF) mechanism, which is deployed at the skip connection position of the decoder. Through three stages—differential projection, collaborative fusion, and adaptive recalibration—it achieves effective collaborative expression of shallow spatial details and deep semantic information.
[0081] Phase 1: Projection of differentiated features of two branches.
[0082] Independent transformation paths are constructed for two types of features with different sources and semantic levels. For skip features on the encoder side, 1×1 convolutions are used to perform channel projection and feature recalibration, compressing redundant information while preserving the spatial topology and retaining fine-grained structural information such as farmland boundaries, field ridges, and irrigation ditches to the greatest extent. For upsampled features on the decoder side, 3×3 convolutions are used to perform local semantic enhancement, improving the ability to express deep semantics by expanding the local receptive field. Through independent mapping of two branches, the representation space alignment of the two types of heterogeneous features is achieved before fusion.
[0083] Phase Two: Synergistic Integration and Structural Reconstruction.
[0084] The encoder and decoder features, after projection, are fused along the channel dimension, and cross-level information interaction and collaborative representation are achieved through two consecutive 3×3 convolution layers. The first convolution layer is used to establish the correlation between spatial detail information and high-level semantic information, realizing the collaborative calibration of boundary structure and category discrimination information; the second convolution layer further optimizes the fused feature distribution, suppresses redundant responses and enhances effective structural expression, thereby obtaining a joint feature representation that combines spatial detail and semantic discrimination capabilities.
[0085] Phase 3: Attention-driven adaptive recalibration.
[0086] After feature fusion, a CBAM attention module is introduced to model the importance of different semantic responses through a channel attention mechanism, achieving adaptive enhancement of key features. Simultaneously, a spatial attention mechanism is used to strengthen the response intensity of key areas such as farmland boundaries, roads, and ditches, improving the continuity and integrity of slender linear features. Joint recalibration of the channel and spatial dimensions further enhances the discriminative and structural representation capabilities of the fused features.
[0087] Through the above three-stage processing, the SSF mechanism achieves effective decoupling, alignment and fusion of shallow spatial detail features and deep semantic features, avoids information conflict caused by direct splicing in traditional skip connections, and improves the restoration accuracy of farmland boundaries and fine-grained structures in complex agricultural remote sensing scenarios.
[0088] Regarding its integration with the basic vision model architecture, the "differentiated projection - collaborative fusion - adaptive recalibration" design concept adopted by the SSF mechanism has good compatibility with the cross-level feature aggregation and attention guidance mechanisms in the current basic vision model. It can be embedded into different encoder-decoder frameworks as a lightweight feature adaptation module in agricultural remote sensing scenarios, and has good transfer and expansion capabilities.
Claims
1. A method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion, characterized in that, The method includes the following steps: S1. Input layer construction: Select farmland remote sensing images from multiple time phases, select several bands for each time phase, and stitch all band images along the channel dimension to form a channel fusion image, which is used as the unified input of the model. S2. First stage network construction: An encoder-decoder network architecture is adopted. The data processing flow is as follows: the channel fusion image is passed through the encoder for 4 levels of downsampling, the bottleneck layer and the decoder for 4 levels of upsampling in sequence to output a farmland probability map. The farmland probability map is thresholded to obtain a farmland prior mask. A boundary detection branch is added to the encoder and an auxiliary supervision branch is added to the decoder. S3. Second-stage network construction: Shares the encoder-decoder network architecture with the first-stage network, but does not introduce boundary detection branches and auxiliary supervision branches; The data processing flow is as follows: the channel fusion image and the farmland prior mask are input together, and then pass through the encoder for 4 levels of downsampling, the bottleneck layer and the decoder for 4 levels of upsampling in sequence to output the farmland segmentation result; S4. Segmentation Model Training: Training is conducted in two stages. In the first stage, the network's loss function is composed of a weighted sum of the main loss, deep supervision loss, and boundary loss. In the second stage, the network's loss function adopts the main loss function.
2. The method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion according to claim 1, characterized in that, The encoder consists of an initial convolutional layer and a 4-level downsampling structure. Each downsampling level sequentially performs 2×2 max pooling, two 3×3 convolutions, and a CBAM attention mechanism. The CBAM attention mechanism is composed of channel attention and spatial attention concatenated: channel attention performs global average pooling and global max pooling on the input features respectively, and after passing through a shared multilayer perceptron, the results are summed and then activated by a sigmoid operation to generate channel weights; spatial attention performs average pooling and max pooling along the channel axis respectively, and after concatenation, the results are convolved to generate spatial weights, and the channel-weighted features are recalibrated element by element.
3. The method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion according to claim 2, characterized in that, The output of the encoder downsampling stage 4 constitutes the bottleneck layer. The bottleneck layer aggregates high-level semantic information from the entire image. Each feature point corresponds to the largest receptive field in the original image, serving as a feature compression and transition hub from the encoder to the decoder.
4. The method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion according to claim 3, characterized in that, The decoder consists of a 4-level upsampling structure that recovers spatial resolution step by step from the bottleneck layer. Each level first performs a 2x upsampling through bilinear interpolation, then connects laterally with the corresponding level features of the encoder, and finally refines the features through two 3×3 convolutions and the CBAM attention mechanism.
5. The method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion according to claim 4, characterized in that, A self-designed split-skip fusion mechanism is introduced in the first to third stages of decoder upsampling and deployed at the skip connection position of the decoder. The split-skip fusion mechanism includes encoder branch and decoder branch. The encoder branch performs 1×1 pointwise convolution on encoder skip features and the decoder branch performs 3×3 convolution on decoder upsampled features. The outputs of the two branches are concatenated along the channel dimension and then subjected to two consecutive 3×3 convolutions to achieve cross-branch feature interaction and nonlinear fusion. Finally, the input is fed into the CBAM attention mechanism.
6. The method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion according to claim 5, characterized in that, A separate boundary detection branch is derived from the initial convolutional layer features of the encoder. This boundary detection branch contains three modules, each of which specifically includes: The first module consists of a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation function layer. The second module consists of 3×3 convolutional layers, batch normalization layers, and ReLU activation function layers. The third module consists of 1×1 convolutional layers and Sigmoid activation function layers; The boundary detection branch outputs a boundary probability map with the same resolution as the channel fused image.
7. The method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion according to claim 6, characterized in that, Auxiliary supervision branches are introduced at the second and third upsampling stages of the decoder. Each auxiliary supervision branch specifically includes: a 3×3 convolutional layer, a batch normalization layer, a ReLU activation function layer, a 1×1 convolutional layer, and a bilinear interpolation upsampling layer. The output of each auxiliary supervision branch is a semantic segmentation prediction map with the same resolution and number of classes as the main output of the decoder.
8. The method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion according to claim 7, characterized in that, The boundary loss is specifically calculated as follows: the weighted binary cross-entropy of the boundary probability map is used as the boundary loss function. .
9. The method for constructing an agricultural remote sensing farmland segmentation model based on multi-scale feature fusion according to claim 7, characterized in that, The deep supervision loss is specifically calculated as follows: Cross-entropy loss and Dice loss are calculated on the semantic segmentation prediction graphs output by the second-level auxiliary supervision branch upsampled from the decoder, and then weighted and fused to obtain the loss function. The cross-entropy loss and Dice loss are calculated on the semantic segmentation prediction graphs output by the upsampled third-level auxiliary supervision branch of the decoder, and then weighted and fused to obtain the loss function. loss function and loss function Together they constitute the loss of deep oversight.