Intelligent mapping method for rice distribution based on multi-scale asymmetric fusion network
By integrating multi-source satellite data through a multi-scale asymmetric fusion network (MSAFNet), the accuracy and robustness issues of automated rice distribution mapping in complex environments were solved, achieving efficient and high-precision rice distribution identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2025-07-25
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to achieve efficient and accurate automated mapping of rice distribution in complex environments. Multimodal remote sensing data fusion neglects modal independence and class imbalance, resulting in limited recognition effectiveness.
A multi-scale asymmetric fusion network (MSAFNet) is adopted to construct an RGB-SAR multimodal remote sensing image dataset by integrating multi-source satellite remote sensing data. Combined with phenological index and human-machine collaborative verification, a multi-scale semantic enhancement module and a category-balanced feature fusion module are designed to achieve multi-scale feature complementarity and category-adaptive weight allocation.
It significantly improves the accuracy and adaptability of rice distribution identification, enabling accurate identification of fragmented fields and small-target rice in complex environments, and enhancing the model's generalization ability and robustness.
Smart Images

Figure CN120564072B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural remote sensing and intelligent information processing technology, specifically to an intelligent mapping method for rice distribution based on a multi-scale asymmetric fusion network. Background Technology
[0002] With the development of remote sensing technology, multi-source satellite imagery has been widely used in fields such as agricultural monitoring, especially in rice distribution monitoring, providing a data foundation for large-scale dynamic analysis. Traditional rice mapping often relies on single optical or radar imagery, which is limited by insufficient samples, meteorological interference, and scarcity of training samples, making it difficult to achieve efficient and high-precision automated mapping.
[0003] While optical imagery has spatial and spectral advantages, it is easily affected by clouds and fog, resulting in discontinuous data acquisition. SAR imagery has all-weather imaging capabilities and can supplement the spatiotemporal gaps in optical imagery. How to efficiently integrate multimodal remote sensing data and give full play to their complementary advantages is the key to improving recognition accuracy.
[0004] In recent years, deep learning methods have driven the development of crop classification, and multimodal fusion networks have further improved interpretation capabilities. However, existing methods often focus on the correlation of features between modalities, ignoring modal independence and class imbalance, resulting in limited segmentation performance in complex environments. At the same time, problems such as insufficient sample representativeness and image contamination severely restrict the application of multi-source satellite imagery in early rice identification.
[0005] Therefore, how to balance complementarity, independence, and category balance in multimodal remote sensing data fusion, overcome the bottlenecks of sample scarcity and data contamination, and achieve efficient, accurate, and automated mapping of rice distribution in complex environments is a technical challenge that urgently needs to be solved. Summary of the Invention
[0006] The purpose of this invention is to provide an intelligent mapping method for rice distribution based on a multi-scale asymmetric fusion network, addressing the following technical problems: How to achieve high-precision automated identification of rice distribution using multimodal remote sensing mapping, and realize efficient and accurate automated mapping of rice distribution in complex environments?
[0007] The objective of this invention can be achieved through the following technical solutions: A smart mapping method for rice distribution based on multi-scale asymmetric fusion networks includes the following steps: S1. Integrate multi-source satellite remote sensing data to determine a multimodal remote sensing image dataset, and preprocess the multimodal remote sensing image dataset; through data acquisition and preprocessing: acquire and integrate multi-source satellite remote sensing data, including high-resolution optical remote sensing images and synthetic aperture radar (SAR) images, to form an RGB-SAR multimodal remote sensing image dataset. Perform multi-level preprocessing on the dataset, including denoising, radiometric and terrain correction, image registration, spatial normalization, and data augmentation using methods such as horizontal flipping, vertical flipping, and diagonal mirroring, to improve data quality and sample diversity; S2. Based on the preprocessed multimodal remote sensing image dataset, an RGB-SAR multimodal remote sensing sample set integrating geometric, phenological, and ecological features is constructed. Sample areas with rice phenological characteristics are automatically identified and labeled. Candidate areas are purified using spatial filtering and centroid vector method to establish a highly representative sample set. Uncertain samples are eliminated through a human-machine collaborative verification mechanism to ensure labeling accuracy.
[0008] S3. A multi-scale asymmetric fusion network is used to extract features from the RGB-SAR multimodal remote sensing sample set, extracting features from both RGB and SAR images, and achieving multi-scale feature complementarity and class-adaptive weight allocation in the feature fusion stage. Specifically, a multi-scale asymmetric fusion network (MSAFNet) is constructed, using a dual-branch parallel coding architecture to extract features from RGB and SAR images respectively. The network includes a multi-scale semantic enhancement module (MSEM) and a class-balanced feature fusion module (CBFM), achieving multi-scale feature complementarity and class-adaptive weight allocation in the feature fusion stage. S4. Input the feature map output by the multi-scale asymmetric fusion network into the improved UPerNet structure for decoding and reconstruction; generate high-resolution rice distribution mapping results, and verify the rice distribution mapping by combining agricultural statistics and existing distribution products.
[0009] Preferably, the multi-source satellite remote sensing data includes optical remote sensing images and SAR radar images; the multimodal remote sensing image data includes Sentinel-1 SAR data and Sentinel-2 multispectral data.
[0010] Preferably, the sample area is analyzed in a synergistic manner based on the Land Surface Water Index, the comprehensive index of vegetation phenology and water body change, and the Sentinel-1 Dual-Polarized Water Index, and is further validated by human-machine collaboration using high-resolution Google Earth imagery.
[0011] Preferably, the detection rule of the Land Surface Water Index method is defined by the following formula:
[0012]
[0013] in, , , These represent the surface reflectance in the near-infrared band, the surface reflectance in the red band, and the surface reflectance in the short-wave infrared band, respectively. Normalized Difference Vegetation Index; Surface moisture index; Use the following equations to calculate the relevant variables:
[0014] in, This represents the paddy field identification result based on the Land Surface Water Index method, where "1" indicates that it is identified as a paddy field and "0" indicates that it is not identified as a paddy field. This is an empirical threshold; Based on the comprehensive index of vegetation phenology and water body change, rice pixels are detected using the following rule, as shown in the formula:
[0015] CCVS represents the combined index of vegetation phenology and water body change. This indicates that the pixel was classified as rice using a CCVS-based method. This indicates the minimum surface moisture index value during the transplanting stage; for The threshold, for The threshold; , They are 0.1 and 0.6 respectively; For, its expression is:
[0016] in, and These represent the heading stage of rice. Value and value; and These represent the rice transplanting period. Value and value; For enhanced vegetation index; Using based and The method can detect common rice pixels, and the specific formula is as follows:
[0017] in, This indicates the results of remote sensing identification of rice based on optical phenological characteristics; "1" represents rice paddy, and "0" represents non-rice paddy. Calculation using Sentinel-1 SAR data An index is used to extract the area of rice-growing regions. The formula for calculating the index is:
[0018] in, The Sentinel-1 Dual-Polarized Water Index is a water extraction index; when its value is... A positive value indicates a body of water, while a negative value indicates a body of non-water. and This refers to the dual polarization method of radar images, where... This indicates vertical transmission and vertical reception. This indicates vertical transmission and horizontal reception, and 8 is the initial threshold selected in this study.
[0019] Preferably, the multi-scale asymmetric fusion network adopts a dual-branch parallel coding architecture, with encoder branches designed for RGB and SAR images that do not share weights; the multi-scale asymmetric fusion network includes a multi-scale semantic enhancement module and a category-balanced feature fusion module.
[0020] Preferably, the multi-scale semantic enhancement module is used to extract multimodal features in the channel and spatial dimensions through an adaptive multi-scale pooling mechanism, and its steps include: The input feature map is downsampled using max pooling at different scales to extract multi-scale spatial features, as shown in the formula:
[0021]
[0022] in, Refers to the input RGB image feature map. This is the feature map of the input SAR image. It is a max pooling operation; For pooled scale sets, For the scale RGB feature pooling results Scale is SAR feature pooling results; Multi-scale pooling units are used to obtain feature representations under different receptive fields, and the pooling results are sampled with consistent feature map sizes. The formula is as follows:
[0023]
[0024] in, It is an upsampling operation, which is bilinear interpolation, to restore the pooled feature map to its original spatial resolution; All are upsampled features of the response characteristics; Attention weights are generated through channel concatenation and a multilayer perceptron, and dynamic weighting of multi-scale features is achieved using the following formula:
[0025]
[0026] Upsampled features at different scales are concatenated along the channel direction to form a rich multi-scale feature representation; among them, This is a splicing operation along the channel dimension. For multi-scale numbers; and The spliced RGB and SAR multi-scale features are shown. By dynamically weighting the concatenated features using MLP, important scale features are highlighted, achieving adaptive feature fusion. The formula is as follows:
[0027]
[0028] in: , This refers to a multilayer perceptron, used to generate attention weights. For activation function, , The weighted RGB and SAR fused features; Combining attention mechanisms with residual connections enhances the representation of small targets and fine-grained features. The formula is as follows:
[0029]
[0030] in, For the feature map of the RGB branch, The feature map of the SAR branch, and It is the fused attention weight feature map.
[0031] Preferably, the category-balanced feature fusion module is used to dynamically adjust the contribution weights of each modality feature through a multi-level feature evaluation and fusion architecture, using a three-way attention mechanism and an adjustable category balancing mechanism; its steps include: At the channel dimension, RGB and SAR features are concatenated and convolved to fuse them, and global and local contextual information is extracted through multi-scale pooling. The formula is as follows:
[0032]
[0033] in, Input as RGB features As SAR feature input, It is a multilayer perceptron used to generate channel or spatial attention weights. Use the Sigmoid activation function; and These are the generated attention weights; A three-way parallel attention mechanism is introduced to evaluate feature importance from three levels: global statistics, channel correlation, and spatial distribution, generating adaptive fusion weights, as shown in the formula:
[0034]
[0035]
[0036] in, Input feature map; For local convolution operations, extract local spatial attention. Use the Sigmoid activation function; For the fused spatial attention weights, For global channel attention weights; Local channel attention weights; In the spatial dimension, a dual-path spatial attention gating system is employed, combining global spatial awareness with local context enhancement, and the fused features are calibrated. The formula is as follows:
[0037]
[0038] in: The fused feature map and These are convolutional operations for the RGB and SAR branches respectively, used for feature dimensionality reduction. , These are spatial attention weight maps for the RGB and SAR branches, respectively, with shapes consistent with the input features. A composite output feature map is generated by normalized weighted summation.
[0039] Preferably, it further includes: The fused feature maps are input into the improved UPerNet structure for decoding and reconstruction, generating high-resolution rice distribution mapping results.
[0040] The beneficial effects of this invention are: (1) This invention constructs an RGB-SAR multimodal spatiotemporal remote sensing image dataset by fusing Sentinel-1 SAR radar images and Sentinel-2 multispectral optical images, and combines phenological, geometric and ecological characteristics to significantly improve the representativeness of the samples and the quality of the data, providing a solid data foundation for remote sensing mapping of rice distribution. Especially under complex plot structures and variable weather conditions, it can effectively ensure the spatiotemporal continuity and integrity of remote sensing information.
[0041] (2) This invention proposes an automatic sample area extraction method based on the collaborative analysis of LSWI, CCVS and SDWI triple phenological indices, and combines it with Google Earth high-resolution imagery for spatial filtering and human-computer collaborative verification, which greatly improves the fineness of the rice sample area boundary and the accuracy of the labeling, and ensures the high quality and high representativeness of the training data from the source.
[0042] (3) This invention innovatively introduces a multi-scale asymmetric fusion mechanism in the network structure, designs a multi-scale semantic enhancement module and a category-balanced feature fusion module, and adopts a dual-branch parallel coding architecture to realize fine-grained complementary extraction of RGB and SAR image features and category-adaptive weighted fusion, which effectively improves the ability to identify fragmented fields and small target rice, and significantly enhances the model's adaptability and generalization ability to complex scenes.
[0043] Of course, any product implementing this invention does not necessarily need to achieve all the advantages described above at the same time. Attached Figure Description
[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is a flowchart illustrating the steps of an intelligent rice distribution mapping method based on a multi-scale asymmetric fusion network according to the present invention. Figure 2This is a schematic diagram of the framework for high-precision intelligent mapping of rice distribution based on a multi-scale asymmetric fusion network as described in this invention. Figure 3 This is a schematic diagram of the structure for refining sample region acquisition based on phenological prior knowledge as described in this invention; Figure 4 This is a schematic diagram of the structure of the multi-scale asymmetric fusion network described in this invention. Figure 5 This is a schematic diagram of the structure of the multi-scale semantic enhancement module described in this invention; Figure 6 This is a schematic diagram of the structure of the category-balanced feature fusion module described in this invention; Figure 7 This is a comparison chart of the accuracy of the rice mapping model described in this invention; Figure 8 This invention relates to rice mapping under cloud and fog interference conditions. Detailed Implementation
[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] Please see Figures 1-8 As shown, in an embodiment of the present invention, the computer is configured as follows: dual Intel Xeon Gold 5318Y core processors, NVIDIA H800 graphics processors with a clock speed of 3.40 GHz, 80 GB of memory, and Linux operating system.
[0048] This invention is an intelligent mapping method for rice distribution based on a multi-scale asymmetric fusion network, achieving high-precision mapping results; it is implemented using the PyTorch 2.10 deep learning framework toolkit.
[0049] Please see Figures 1-2 As shown, a smart mapping method for rice distribution based on a multi-scale asymmetric fusion network specifically includes the following steps: Step 1: Integrate the multimodal remote sensing image dataset. The specific steps are as follows: This study integrates multimodal data from multiple satellite sensors, including high-resolution optical remote sensing images and synthetic aperture radar (SAR) images, to form an RGB-SAR multimodal remote sensing image dataset. The acquired image data undergoes preprocessing, primarily including denoising, image alignment, and data annotation, to ensure accurate correspondence between different modalities. Furthermore, data augmentation techniques such as horizontal flipping, vertical flipping, and diagonal mirroring are employed to enhance data quality and sample diversity. Step 2: Automatic extraction and labeling of multimodal sample areas. Based on the preprocessed multimodal remote sensing image dataset, an RGB-SAR multimodal remote sensing sample set integrating geometric, phenological and ecological features is constructed. Sample areas with rice phenological characteristics are automatically identified and labeled. Candidate areas are purified using spatial filtering and centroid vector method to establish a highly representative sample set. Uncertain samples are eliminated through human-machine collaborative verification mechanism to ensure labeling accuracy.
[0050] Specifically, in one embodiment, such as Figure 3 As shown, the specific steps are as follows: based on the collaborative analysis of the Land Surface WaterIndex (LSWI), the Comprehensive Index of Vegetation Phenology and Water Change (CCVS), and the Sentinel-1 Dual-Polarized WaterIndex (SDWI), and combined with Google Earth high-resolution imagery, human-machine collaborative verification is carried out.
[0051] In one embodiment of the present invention, the detection rule for the Land Surface Water Index method is defined by the following formula:
[0052]
[0053] in, , , These represent the surface reflectance in the near-infrared band, the surface reflectance in the red band, and the surface reflectance in the short-wave infrared band, respectively. Normalized Difference Vegetation Index; Surface moisture index; It aims to detect flooding signals in rice paddies during sowing and transplanting. Studies have found that during transplanting, the LSWI value of rice paddies is typically higher than the corresponding NDVI value. Therefore, detection rules based on the LSWI method were determined, and relevant variables were calculated using the following equation:
[0054] in, This represents the paddy field identification result based on the Land Surface Water Index method, where "1" indicates that it is identified as a paddy field and "0" indicates that it is not identified as a paddy field. This is an empirical threshold; For the CCVS-based method, rice mapping is performed using two images from the tillering stage to the heading stage. Because rice paddies are submerged after transplanting, the LSWI value remains high from transplanting to heading, typically showing only a slight increase before heading. On the other hand, the EVI (Enhancing Vegetation Index) increases significantly, similar to other upland crops; therefore, the ratio of LSWI to EVI change (RCLE) in rice paddies from transplanting to heading is much lower than in upland crops. Simultaneously, the lowest LSWI value in rice paddies at the transplanting stage is also higher than in other upland crops. Based on these unique phenomena, rice pixels can be detected using the following rule, as shown in the formula:
[0055] CCVS represents the combined index of vegetation phenology and water body change. This indicates that the pixel was classified as rice using a CCVS-based method. This indicates the minimum surface moisture index value during the transplanting stage; for The threshold, for The threshold; , The values are 0.1 and 0.6 respectively; the expression is:
[0056] in, and These represent the heading stage of rice. Value and value; and These represent the rice transplanting period. Value and value; Using based and The method can detect common rice pixels, and the specific formula is as follows:
[0057] in, This represents the results of remote sensing identification of rice combined with optical phenological characteristics; Based on the above, when specifically extracting rice planting areas during the transplanting period, the SDWI (Sentinel-1 Dual-Polarized Water Index) was calculated using Sentinel-1 SAR data to extract the area of the rice planting area. The formula for calculating the index is:
[0058] in, The Sentinel-1 Dual-Polarized Water Index is a water extraction index. When its value is greater than 0 (positive value), it is identified as a water body; when its value is less than 0 (negative value), it is identified as a non-water body. and This refers to the dual polarization method of radar images, where... (Vertical transmit – Vertical receive) means vertical transmission and vertical reception. (Vertical transmit–Horizontal receive) represents vertical transmission and horizontal reception, and 8 is the initial threshold selected in this study.
[0059] Complementing the aforementioned methods, threshold segmentation is further performed to extract sample planting areas for the remaining growth period of rice. Combined with high-resolution remote sensing imagery, sample areas with rice phenological characteristics are automatically identified. Spatial filtering and centroid vector method are used to purify candidate regions, establishing a highly representative sample set. Uncertain samples are then eliminated through a human-machine collaborative verification mechanism, improving the accuracy of sample area boundaries and labeling.
[0060] Step 3: Multi-scale asymmetric fusion network structure design and feature extraction, the specific steps are as follows: This step involves constructing a multi-scale asymmetric fusion network (MSAFNet), such as... Figure 4 As shown, this network achieves efficient feature extraction and fusion of multimodal remote sensing images. It employs a dual-branch parallel coding architecture, designing separate encoder branches with non-shared weights for RGB and SAR images. Each branch can independently mine feature representations of different modalities. Combined with the residual structure, it effectively captures and utilizes the substantial differences in spatial, spectral, and structural information between RGB and SAR images, laying a solid foundation for subsequent feature interaction and fusion. During feature extraction, MSAFNet introduces two key modules to address common problems in rice segmentation of remote sensing images, namely class imbalance and multi-scale target issues, significantly improving segmentation performance and robustness. (1) Multi-scale semantic enhancement module (MSEM) such as Figure 5 As shown: The MSEM module innovatively employs an adaptive multi-scale pooling mechanism to address the scale differences between RGB and SAR image features. It downsamples the input feature map using max pooling at different scales to extract multi-scale spatial features, thereby enhancing the model's ability to perceive different target sizes. The formula is as follows:
[0061]
[0062] in, Refers to the input RGB image feature map. This is the feature map of the input SAR image. It is a max pooling operation; For pooled scale sets, For the scale RGB feature pooling results Scale is SAR feature pooling results; Next, in both channel and spatial dimensions, the module first uses a differential feature extraction unit to explicitly model the complementarity of the two modalities based on the principle of differential amplification. Then, a multi-scale pooling unit is used to obtain feature representations under different receptive fields, and the pooling results are upsampled to ensure consistent feature map size. The formula is as follows:
[0063]
[0064] in, It is an upsampling operation, which is bilinear interpolation, to restore the pooled feature map to its original spatial resolution; All are upsampled features of the response characteristics; Attention weights are generated through channel concatenation and a multilayer perceptron, and dynamic weighting of multi-scale features is achieved using the following formula:
[0065]
[0066] Upsampled features at different scales are concatenated along the channel direction to form a rich multi-scale feature representation; among them, This is a splicing operation along the channel dimension. For multi-scale numbers; and The spliced RGB and SAR multi-scale features are shown. Then, the spliced features are dynamically weighted using MLP to highlight important scale features, achieving adaptive feature fusion. The formula is as follows:
[0067]
[0068] in: , This refers to a multilayer perceptron, used to generate attention weights. For activation function, , The weighted RGB and SAR fused features; Finally, by combining attention mechanisms with residual connections, the representation of small targets and fine-grained features is enhanced, as shown in the formula:
[0069]
[0070] in, For the feature map of the RGB branch, The feature map of the SAR branch, and It is the fused attention weight feature map; this module not only effectively integrates local details and global semantic information, but also adaptively highlights modal complementarity, improving the segmentation sensitivity and accuracy of small target regions such as rice.
[0071] (2) Category Balanced Feature Fusion Module (CBFM) Figure 6 As shown, the CBFM module addresses the challenges of class imbalance and feature fusion in multimodal remote sensing segmentation by designing a multi-level feature evaluation and fusion architecture mechanism. First, the module concatenates and convolves RGB and SAR features along the channel dimension, and then extracts global and local contextual information through multi-scale pooling, as shown in the formula:
[0072]
[0073] in, Input as RGB features As SAR feature input, It is a multilayer perceptron used to generate channel or spatial attention weights. Use the Sigmoid activation function; and These are the generated attention weights; To further enhance the expressive power of small target features, CBFM introduces a three-way parallel attention mechanism, which evaluates feature importance from three levels: global statistics, channel correlation, and spatial distribution, and generates adaptive fusion weights, as shown in the formula:
[0074]
[0075]
[0076] in, Input feature map; For local convolution operations, extract local spatial attention. Use the Sigmoid activation function; For the fused spatial attention weights, For global channel attention weights; Local channel attention weights; Furthermore, the module incorporates an adjustable class balancing mechanism that uses adaptive thresholds to distinguish between foreground (small targets) and background (large targets), assigning higher weights to smaller class targets and significantly alleviating the detection challenges caused by class imbalance. In the spatial dimension, CBFM employs dual-path spatial attention gating, combining global spatial awareness with local context enhancement, and then further calibrates the fused features, using the following formula:
[0077]
[0078] in: The fused feature map and These are convolutional operations for the RGB and SAR branches respectively, used for feature dimensionality reduction. , These are spatial attention weight maps for the RGB and SAR branches, respectively, with shapes consistent with the input features. Finally, a composite output feature map is generated by normalized weighted summation, achieving deep complementarity of multimodal information and noise suppression.
[0079] Through the aforementioned mechanism, MSAFNet can fully leverage the complementary advantages of RGB and SAR data, dynamically adjust the contribution weights of each modality feature, and significantly enhance the ability to identify small targets and rice distributions with imbalanced categories. Experimental results show that the fusion strategy based on multi-scale pooling, attention mechanism, and class balancing mechanism achieves superior segmentation performance compared to traditional methods in complex remote sensing scenarios, especially in application scenarios involving small targets such as rice and extremely imbalanced categories, demonstrating stronger robustness and generalization ability.
[0080] Step 4: Feature Decoding and High-Resolution Mapping, the specific steps are as follows: The fused feature maps are input into the improved UperNet structure, as follows: Figure 4Decoding and reconstruction are performed as shown. UPerNet, as a multi-level feature decoding network, can fully utilize multi-scale feature information extracted during the encoding stage. Its core mechanism lies in effectively integrating spatial and semantic information from different scales through Feature Pyramid Pooling (FPN) and multi-level feature fusion, achieving accurate restoration of paddy field boundaries and details. During the decoding process, the network fuses shallow spatial detail features and deep global semantic features layer by layer, improving segmentation performance for complex scenes such as fragmented fields and small-target paddy areas. The final output paddy distribution map is characterized by high resolution, clear boundaries, and rich details.
[0081] This invention also includes joint optimization of the loss function, with the following specific steps: To address the problem of extreme class imbalance between foreground (e.g., rice) and background categories, a weighted combination of weighted binary cross-entropy loss and Dice loss is employed during model training. The weighted binary cross-entropy loss explicitly assigns weights and boosting factors to rice category pixels, while the Dice loss measures segmentation overlap, effectively mitigating gradient vanishing and improving minority class segmentation accuracy. The total loss function is a weighted sum of both losses, with adjustable weights, significantly improving the model's segmentation accuracy and robustness. The joint optimization strategy of weighted binary cross-entropy loss and Dice loss during model training, setting class weights and boosting factors for minor categories like rice, significantly alleviates the extreme class imbalance problem, effectively improving the accuracy and robustness of rice distribution segmentation, and enabling the model to maintain high-level segmentation performance in a wide range of diverse agricultural environments.
[0082] Furthermore, the reasoning and application of rice distribution mapping were verified. To verify the effectiveness of this scheme, the reasoning and application of rice distribution were carried out based on the trained multimodal remote sensing mapping model, including quantitative and qualitative comparative analysis. In the quantitative analysis, Table 1 shows the performance of mainstream multimodal segmentation methods (Convnext_v2, ASANet, CMX, SA-Gate, vfusenet) and the MSAFNet of this invention on the rice remote sensing dataset of Sichuan Tianfu New Area. The results show that MSAFNet achieves the best or second-best results in core indicators such as Kappa, aAcc, mFscore, and mIoU, with mIoU reaching 91.27%, which is significantly improved compared with the second-best method, highlighting its advantages in class imbalance and small target recognition. Meanwhile, MSAFNet outperforms existing methods in terms of average F1 score and recall, effectively improving the accuracy and segmentation completeness of rice region detection. The system of this invention supports comparative evaluation of various mainstream multimodal segmentation models and uses multi-dimensional indicators such as overall accuracy, average F1 score, and average intersection-union ratio for comprehensive performance evaluation. It can generate high-resolution rice distribution maps with clear boundaries and rich details in large-scale and complex environments, providing efficient and reliable technical support for applications such as intelligent agricultural management, rice planting monitoring, and area estimation.
[0083] Table 1
[0084] For ease of reading, the Chinese meanings of the English abbreviations in the table are as follows: Kappa—Kappa coefficient; aAcc—average accuracy; mAcc—average precision; mFscore—average F1 score; mPrecision—average precision; mRecall—average recall; mIoU—average intersection-union ratio; iou_Background—IoU of the background category; iou_rice—IoU of the rice category. In the qualitative analysis, representative test areas are selected to visually compare the segmentation results of different methods. Figure 7 The results show that MSAFNet not only performs excellently in large, contiguous rice paddy areas, but also accurately identifies fragmented, strip-shaped, and complex boundaries. The segmentation results are highly consistent with manually labeled data, and its detail restoration capability and robustness are significantly superior to other methods. Especially in extremely fragmented and high-noise scenarios, MSAFNet, relying on its multi-scale semantic enhancement and class-balanced feature fusion modules, significantly improves the recognition ability of small targets and complex boundaries, effectively suppressing missed and misclassified segments, demonstrating excellent segmentation performance and application value.
[0085] To verify the innovativeness and application effect of the technical solution of this invention, a specific implementation case is included, comprising an application scenario, as described below: Examples of specific implementation and application scenarios: This invention addresses the challenges of rice identification in complex terrains, particularly in Sichuan Province, characterized by hilly terrain, scattered plots, and frequent cloud cover. Traditional single-source optical remote sensing data suffers from insufficient accuracy due to cloud cover. This invention proposes a rice distribution mapping framework based on a multi-scale asymmetric fusion network (MSAFNet) to fully leverage the information potential of multi-source remote sensing data. First, it combines Sentinel-1 / 2 multi-source satellite data, fusing geometric, phenological, and ecological features to construct an RGB-SAR multi-modal rice remote sensing image dataset. Then, a multi-scale asymmetric mechanism is introduced into the network's feature extraction layer to deeply mine the complementary details between multi-modal data. Through a multi-scale semantic enhancement module (MSEM) and a class-balanced feature fusion module (CBFM), the selective fusion of different modal features is effectively achieved, significantly improving the rice paddy identification capability. The model's mapping performance under cloud and fog interference is also demonstrated. Figure 8 As shown.
[0086] According to detailed technical materials, this invention provides a high-precision intelligent mapping system for rice distribution based on a multi-scale asymmetric fusion network. This method integrates multi-source remote sensing data, such as Sentinel-1 SAR and Sentinel-2 multispectral data, to construct an RGB-SAR multimodal spatiotemporal dataset, fusing phenological, geometric, and ecological features to significantly improve sample representativeness and data quality. Within a deep learning framework, a multi-scale asymmetric fusion mechanism is innovatively introduced, designing a multi-scale semantic enhancement module (MSEM) and a class-balanced feature fusion module (CBFM) to achieve fine-grained complementary extraction and adaptive fusion of multimodal features, effectively addressing complex scenarios such as fragmented fields and extreme class imbalance. Simultaneously, multi-scale pooling and adaptive attention mechanisms are employed to enhance the recognition ability of small targets and imbalanced classes, and weighted binary cross-entropy and Dice loss are jointly optimized to improve mapping accuracy and robustness; meeting the needs of rice monitoring and area estimation in large-scale complex environments; this technology provides efficient and reliable remote sensing support for modern agricultural intelligent management and scientific decision-making.
[0087] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0088] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in this application, they should all fall within the protection scope of the present invention.
Claims
1. A smart mapping method for rice distribution based on multi-scale asymmetric fusion networks, characterized in that, Includes the following steps: S1. Integrate multi-source satellite remote sensing data to determine a multimodal remote sensing image dataset, and preprocess the multimodal remote sensing image dataset; S2. Based on the preprocessed multimodal remote sensing image dataset, construct an RGB-SAR multimodal remote sensing sample set that integrates geometric, phenological, and ecological features, and automatically identify and label sample areas with rice phenological features. S3. A multi-scale asymmetric fusion network is used to extract features from the RGB-SAR multimodal remote sensing sample set, extracting RGB and SAR image features, and realizing multi-scale feature complementarity and class adaptive weight allocation in the feature fusion stage. S4. Decode and reconstruct the feature maps output by the multi-scale asymmetric fusion network to generate high-resolution rice distribution mapping results, and verify the rice distribution mapping by combining agricultural statistics and existing distribution products. The multi-scale asymmetric fusion network adopts a dual-branch parallel coding architecture, with encoder branches designed for RGB and SAR images that do not share weights; the multi-scale asymmetric fusion network includes a multi-scale semantic enhancement module and a category-balanced feature fusion module. The multi-scale semantic enhancement module is used to extract multimodal features in the channel and spatial dimensions through an adaptive multi-scale pooling mechanism, and its steps include: The input feature map is downsampled using max pooling at different scales to extract multi-scale spatial features, as expressed by the formula: in, Refers to the input RGB image feature map. This is the feature map of the input SAR image. It is a max pooling operation; For pooled scale sets, For the scale RGB feature pooling results Scale is SAR feature pooling results; Multi-scale pooling units are used to obtain feature representations under different receptive fields, and the pooling results are sampled with consistent feature map sizes. The formula is as follows: in, It is an upsampling operation, which is bilinear interpolation, to restore the pooled feature map to its original spatial resolution; All are upsampled features of the response characteristics; Attention weights are generated through channel concatenation and a multilayer perceptron, and dynamic weights are applied to multi-scale features, as shown in the following formula: Upsampled features at different scales are concatenated along the channel direction to form a rich multi-scale feature representation; among them, This is a splicing operation along the channel dimension. For multi-scale numbers; and The spliced RGB and SAR multi-scale features are shown. By dynamically weighting the concatenated features using MLP, important scale features are highlighted, achieving adaptive feature fusion. The formula is as follows: in: , This refers to a multilayer perceptron, used to generate attention weights. For activation function, , The weighted RGB and SAR fused features; Combining attention mechanisms with residual connections enhances the representation of small targets and fine-grained features: in, For the feature map of the RGB branch, The feature map of the SAR branch, and It is the fused attention weight feature map; The category-balanced feature fusion module is used to adopt a multi-level feature evaluation and fusion architecture, and dynamically adjust the contribution weight of each modality feature through a three-way attention mechanism and an adjustable category balance mechanism.
2. The intelligent mapping method for rice distribution based on a multi-scale asymmetric fusion network according to claim 1, characterized in that, The multi-source satellite remote sensing data includes optical remote sensing images and SAR radar images; the multimodal remote sensing image data includes Sentinel-1 SAR data and Sentinel-2 multispectral data.
3. The intelligent mapping method for rice distribution based on a multi-scale asymmetric fusion network according to claim 1, characterized in that, The sample area was analyzed based on the Land Surface Water Index, the comprehensive index of vegetation phenology and water body change, and the Sentinel-1 Dual-Polarized Water Index, and was verified by human-machine collaboration using high-resolution images.
4. The intelligent mapping method for rice distribution based on a multi-scale asymmetric fusion network according to claim 3, characterized in that, The detection rules of the Land Surface Water Index method are as follows: in, , , These represent the surface reflectance in the near-infrared band, the surface reflectance in the red band, and the surface reflectance in the short-wave infrared band, respectively. Normalized Difference Vegetation Index; Surface moisture index; Calculate using the following equation: in, This represents the paddy field identification result based on the Land Surface Water Index method, where "1" indicates that it is identified as a paddy field and "0" indicates that it is not identified as a paddy field. This is an empirical threshold; Based on a comprehensive index of vegetation phenology and water body change, rice pixels are detected using the following rules: CCVS represents the combined index of vegetation phenology and water body change. This indicates that the pixel was classified as rice using a CCVS-based method. This indicates the minimum surface moisture index value during the transplanting stage; for The threshold, for The threshold; , They are 0.1 and 0.6 respectively; The expression is: in, and These represent the heading stage of rice. Value and value; and These represent the rice transplanting period. Value and value, For enhanced vegetation index; Using based and The method can detect common rice pixels, as follows: in, This represents the results of remote sensing identification of rice combined with optical phenological characteristics; Calculation using Sentinel-1 SAR data An index is used to extract the area of rice-growing regions. The formula for calculating the index is as follows: in, The Sentinel-1 Dual-Polarized Water Index is a water extraction index. When its value is greater than 0, it is considered a water body, and when it is less than 0, it is considered a non-water body. and This refers to the dual polarization method of radar images, where... This indicates vertical transmission and vertical reception. This indicates vertical transmission and horizontal reception, and 8 is the initial threshold selected in this study.
5. The intelligent mapping method for rice distribution based on a multi-scale asymmetric fusion network according to claim 1, characterized in that, Dynamically adjust the contribution weights of each modal feature; The steps include: RGB and SAR features are concatenated and convolved at the channel dimension, and global and local contextual information is extracted through multi-scale pooling. As shown in the formula: in, Input as RGB features As SAR feature input, It is a multilayer perceptron used to generate channel or spatial attention weights. Use the Sigmoid activation function; and These are the generated attention weights; A three-way parallel attention mechanism is introduced to evaluate feature importance from three levels: global statistics, channel correlation, and spatial distribution, generating adaptive fusion weights: in, Input feature map; For local convolution operations, extract local spatial attention. Use the Sigmoid activation function; For the fused spatial attention weights, For global channel attention weights; Local channel attention weights; In the spatial dimension, a dual-path spatial attention gating system is employed, combining global spatial awareness with local context enhancement, and the fused features are calibrated. in: The fused feature map and These are convolutional operations for the RGB and SAR branches respectively, used for feature dimensionality reduction. , These are spatial attention weight maps for the RGB and SAR branches, respectively, with shapes consistent with the input features. A composite output feature map is generated by normalized weighted summation.
6. The intelligent mapping method for rice distribution based on a multi-scale asymmetric fusion network according to claim 1, characterized in that, Also includes: The fused feature maps are input into the improved UPerNet structure for decoding and reconstruction, generating high-resolution rice distribution mapping results.