Rice information extraction method and device based on multi-source remote sensing and machine learning
By combining Sentinel-1 VH polarization data and Sentinel-2 reflectivity data, a rice index was constructed and a random forest model was used to solve the problems of low accuracy and high cost in existing rice extraction technologies, thus achieving high-precision and low-cost rice information extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGSU WATER CONSERVANCY SCI RES INST
- Filing Date
- 2026-02-06
- Publication Date
- 2026-04-17
AI Technical Summary
Existing rice extraction methods rely heavily on manual samples, and data quality is poor and computational costs are high in cloudy areas, making it difficult to achieve high-precision and efficient rice information extraction.
Using Sentinel-1 VH polarization data and Sentinel-2 reflectance data based on multi-source remote sensing, combined with the rice index and random forest model, a high-precision rice distribution map was obtained by constructing a rice sample union, training a classification model, and performing weighted fusion.
It achieves high-precision extraction of rice distribution, reduces computational costs, avoids unreasonable fragmentation in traditional methods, and improves the spatial distribution rationality of the extraction results.
Smart Images

Figure CN121661528B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of remote sensing data processing technology, specifically relating to a method and device for extracting rice information based on multi-source remote sensing and machine learning. Background Technology
[0002] Rice, as one of the major food crops, provides a stable food source for human society. Currently, the growing global population is placing greater demands on food supply, and the impacts of urbanization, environmental issues, and climate change are posing greater challenges to rice cultivation and planting areas. Furthermore, rice cultivation also affects farmland use intensity, agricultural water consumption, and greenhouse gas emissions, thereby impacting agricultural production management and the ecological environment. Therefore, timely access to agricultural information, such as rice planting area, plays a supporting role in regional agricultural regulation and environmental decision-making.
[0003] Remote sensing technology, with its advantages of large-area, high-timeliness, and low-cost Earth observation, has become an important technical means for extracting spatial distribution information of crops. Existing satellites in orbit can provide remote sensing data sources with different spatial, temporal, spectral resolutions and sensor types. This richness of remote sensing data sources has greatly promoted the development of refined remote sensing mapping for crops. In early rice remote sensing mapping research, low spatial resolution, high temporal resolution MODIS data was used to extract the spatial distribution of rice in southern China and Southeast Asia. Factors affecting the accuracy of rice identification were summarized from aspects such as the spatiotemporal resolution of remote sensing data, weather conditions, and land cover type (wetlands). Furthermore, a decision tree model driven by time-series MODIS data was used to extract the spatial distribution characteristics of single- and double-cropping rice in southern China. Subsequently, with the increasing richness of remote sensing data sources, researchers have developed a series of rice extraction methods based on satellite data such as Landsat, Sentinel-1, Sentinel-2, HJ-1, and GF-6, achieving good results in different application scenarios. Meanwhile, the rapid development of remote sensing cloud computing platforms such as Google Earth Engine (GEE), Amazon Web Services, PIE-Engine, and Planetary Computer has greatly improved the efficiency of large-scale remote sensing mapping of crops. Among them, the GEE platform is the most widely used in crop remote sensing mapping research.
[0004] Rice mapping primarily employs two approaches: phenological methods and machine learning methods. Phenological methods rely on the unique phenological characteristics of rice, particularly the flooding signal during transplanting. However, the flooding signal is short-lived, typically lasting only a few weeks, making these methods particularly sensitive to data quality issues. Missing values for key phenological stages can lead to misidentification, thus reducing classification accuracy. Therefore, existing high-resolution rice mapping studies either rely on SAR data or focus only on regions with less cloud cover, such as Northeast China. On the other hand, machine learning methods, including Support Vector Machines (SVM), Random Forests (RF), and deep learning methods, can achieve high accuracy. However, they typically require large amounts of training samples and face challenges in transferring models across years. These limitations highlight the urgent need for new methods that can effectively handle poor-quality optical data while ensuring reliable accuracy in long-term rice mapping. Summary of the Invention
[0005] The purpose of this invention is to overcome the problems of high dependence on manual samples, poor data quality in cloudy areas, and high computational costs in the existing technology for rice information extraction, and to provide a method and device for rice information extraction based on multi-source remote sensing and machine learning.
[0006] In a first aspect, the present invention provides a method for extracting rice information based on multi-source remote sensing and machine learning, the method comprising:
[0007] A method for extracting rice information based on multi-source remote sensing and machine learning, the method comprising:
[0008] The temporal VH polarization values of pixels are obtained based on Sentinel-1 VH polarization data, and the minimum value of the temporal VH polarization value of the pixel is calculated. Maximum value The ratio of variance to mean Constructing a rice index Combined pixels Value and threshold segmentation methods were used to extract the first rice sample from Sentinel-1 images;
[0009] A second rice sample was extracted from Sentinel-2 images based on Sentinel-2 reflectance data and remote sensing optical indices.
[0010] The union of the first rice sample and the second rice sample is taken to form a potential rice distribution map; pixels are extracted from the distribution map as rice sample labels, and pixels from the remaining area are extracted as non-rice sample labels.
[0011] VH time series values and SWIR1 band reflectance of each labeled sample point are extracted from Sentinel-1 and Sentinel-2 images to construct the first training dataset and the second training dataset.
[0012] The first classification model and the second classification model were trained using the first training dataset and the second training dataset, respectively.
[0013] Sentinel-1 and Sentinel-2 image data of the study area were acquired, and the first rice probability distribution and the second rice probability distribution were obtained using the first classification model and the second classification model, respectively.
[0014] The first rice probability distribution and the second rice probability distribution are linearly weighted and fused to obtain a probability distribution map. The probability distribution map is then constrained by the area using data from a statistical yearbook to obtain an optimal probability threshold. Based on the probability threshold, a rice distribution map is extracted from the probability distribution map.
[0015] In some embodiments of the present invention, the remote sensing optical index includes the LSWI index and the NDVI index;
[0016] The method for extracting the second rice sample from Sentinel-2 images is as follows:
[0017] The difference between the LSWI index value and the NDVI index value of each pixel during the flooding period of the paddy field is calculated. Based on the difference, a second rice sample is extracted using a threshold segmentation method.
[0018] In some embodiments of the present invention, the , , The rice index was calculated after normalization.
[0019] In some embodiments of the present invention, pixels are extracted from the distribution map by a random sampling method.
[0020] In some embodiments of the present invention, the first classification model and the second classification model are selected from the random forest model.
[0021] In some embodiments of the present invention, Sentinel-1 and Sentinel-2 image data of the study area are acquired, and the temporal VH polarization data of the Sentinel-1 image and the temporal SWIR1 band reflectance data of the Sentinel-2 image are used to generate a time series dataset using the median synthesis method, which is used as the input of the first classification model and the second classification model, respectively.
[0022] In some embodiments of the present invention, the first rice probability distribution and the second rice probability distribution are linearly weighted and fused with equal weights.
[0023] The present invention further provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to perform the steps of the above method.
[0024] The present invention further provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0025] The present invention further provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the steps of the above-described method.
[0026] The present invention has the following beneficial effects:
[0027] (1) This invention integrates Sentinel-1 SAR VH data and Sentinel-2 optical data, and uses rice sample optimization model and random forest model threshold optimization methods to finally obtain high-precision rice distribution in the test area;
[0028] (2) The present invention constructs a new rice index to achieve high-precision rice identification of Sentinel-1 images, thereby obtaining high-precision training data for classification model training;
[0029] (3) The rice extraction results of the present invention effectively avoid the unreasonable small fragments of traditional methods, and the spatial distribution of the extraction results is more reasonable.
[0030] (4) The present invention obtains high-precision extraction results based on free remote sensing data, which significantly reduces the application cost of the technical solution. Attached Figure Description
[0031] Figure 1 This is a flowchart of the rice information extraction method described in this invention.
[0032] Figure 2 This is a comparison chart of the extraction results from radar, optics, and fused SAR and optics. Detailed Implementation
[0033] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0034] Example 1
[0035] This embodiment specifically illustrates the implementation process of the method of the present invention, such as... Figure 1 As shown, it includes the following steps:
[0036] Step 1, screening of potential rice samples;
[0037] Sentinel-1 SAR VH data and Sentinel-2 optical data were collected during the rice growing season in the study area. Using the Sentinel-1 SAR VH polarization data, the VH time series for each pixel was calculated, and a rice index was constructed. Initial rice samples were selected, and the ratios of the minimum, maximum, variance, and mean of the time-series VH values were calculated and normalized. The rice index was then constructed as follows:
[0038]
[0039] In the formula, , , , These represent the normalized time-series minimum, maximum, variance, and mean of the VH value, respectively. A threshold was established based on experience and comparative experiments, and pixels with a rice index greater than the threshold were considered rice samples (hereinafter referred to as the first rice sample).
[0040] Using the reflectance of NIR, RED, and SWIR1 bands from Sentinel-2 optical data, the LSWI and NDVI of each pixel during the flooding period in paddy fields were calculated, and a rice sample selection method was constructed, wherein:
[0041]
[0042]
[0043] in , , These are the reflectivities of the NIR, RED, and SWIR1 bands of Sentinel-2, respectively.
[0044] If a pixel has at least one image during the flooding period that satisfies the difference between LSWI and NDVI greater than the threshold of 0.5, then the pixel is considered to be a rice sample (hereinafter referred to as the second rice sample).
[0045] Samples that simultaneously satisfy both Sentinel-1 and Sentinel-2 methods are considered to be rice samples with higher final accuracy, while the remaining samples are considered non-rice samples. A potential rice distribution map is generated by taking the union of the first and second rice samples. Pixels are randomly selected from this distribution map as rice sample labels, and pixels are randomly selected from the remaining areas as non-rice sample labels.
[0046] Step 2: Train random forest models for the two types of images based on the training samples obtained in Step 1.
[0047] Specifically, it includes:
[0048] Based on the rice sample labels and non-rice sample labels extracted in step 1, the corresponding radar time-series features and optical time-series features are extracted from the Sentinel-1 and Sentinel-2 images, respectively.
[0049] Among them, the radar timing feature is a 12-day Sentinel-1 time-series VH polarization sequence; the optical timing feature is an 8-day SWIR1 time-series sequence.
[0050] The first training dataset is constructed based on the extracted temporal VH polarization sequence and the corresponding sample labels; the second training dataset is constructed based on the extracted temporal SWIR1 sequence and the corresponding sample labels.
[0051] Using the first and second training datasets, a first classification model (Sentinel-1 classification model) and a second classification model (Sentinel-2 classification model) were constructed using the random forest method. The model parameters were uniformly configured as follows: the number of decision trees (n_estimators) was 100, the feature selection criterion was the Gini impurity coefficient, and the maximum number of features (max_features) was the square root of the total number of features.
[0052] Step 3: Use a classification model to obtain a probability image of rice.
[0053] Median synthesis was used to generate time-series datasets from Sentinel-1 SAR VH polarization data and Sentinel-2 SWIR1 band data, respectively. The first rice probability distribution was obtained using the first classification model for the time-series VH polarization data; the second rice probability distribution was obtained using the second classification model for the time-series SWIR1 band data, with values ranging from 0 to 1.
[0054] Step 4: Obtain the final rice distribution through weighted fusion and statistical yearbook constraints.
[0055] The probability distributions of the first and second rice varieties are linearly weighted and fused with the same weights to obtain a comprehensive probability distribution map.
[0056] By introducing the rice planting area data from the statistical yearbook of the corresponding year in the study area as a constraint, the optimal probability threshold is automatically found in the comprehensive probability distribution map through iterative search, so that the total area of rice pixels extracted under this threshold is closest to the actual planting area in the statistical yearbook.
[0057] The probability distribution map is binarized using the obtained optimal probability threshold, and finally a high-precision spatial distribution map of rice is generated.
[0058] Example 2
[0059] This embodiment uses the method shown in Embodiment 1 to extract rice information in a certain city. The city has complex land cover types and crop rotation, and the spectral confusion is relatively serious. At the same time, the optical data is greatly affected by clouds and rain, which greatly limits the usable data.
[0060] In this embodiment, based on measured rice samples from the city, the accuracy of the extracted rice distribution map is verified using overall accuracy (OA), user accuracy (UA), and producer accuracy (PA). The formula is as follows:
[0061]
[0062]
[0063]
[0064] Wherein, TP is the number of correctly classified rice samples, TN is the number of correctly classified non-rice samples, FP is the number of non-rice samples that were classified as rice samples, and FN is the number of rice samples that were classified as non-rice samples.
[0065] The accuracy evaluations of the extraction results of Sentinel-1, Sentinel-2, fusion extraction, and existing K-means methods are shown in Tables 1-4.
[0066] Table 1. Sentinel-1 SAR extraction results in rice
[0067]
[0068] Table 2. Results of Sentinel-2 optical extraction from rice
[0069]
[0070] Table 3. Results of rice extraction fused with SAR and optical methods
[0071]
[0072] Table 4. K-means Sample Selection Method and Rice Extraction Results
[0073]
[0074] As shown in Tables 1, 2, and 3, after fusing SAR and optical data, the extracted UA (Uniform Ability) of rice was 95.75%, higher than 95.61% for SAR data and 95.59% for optical data; the extracted PA (Potential Ability) was 96.21%, higher than 92.89% for SAR data and 92.42% for optical data; and the OA (Optical Ability) was 94.44%, also higher than 92.16% for SAR data and 91.83% for optical data. Comparing Tables 3 and 4, the method of extracting rice samples using the constructed rice index and Sentinel-2 optical data in this invention improves the accuracy of extracted UA, PA, and OA compared to the commonly used K-means rice sample selection method. Simultaneously, a comparison was made... Figure 2 The comparison distribution of radar, optical, and fused SAR and optical results shown in the diagram demonstrates that the spatial distribution of the extracted results from fused SAR and optical data is more reasonable, avoiding unreasonable small fragments and eliminating the influence of noise signals.
Claims
1. A method for extracting rice information based on multi-source remote sensing and machine learning, characterized in that, The method includes: The temporal VH polarization values of pixels are obtained based on Sentinel-1 VH polarization data, and the minimum value of the temporal VH polarization value of the pixel is calculated. Maximum value The ratio of variance to mean Constructing a rice index Combined pixels Value and threshold segmentation methods were used to extract the first rice sample from Sentinel-1 images; The second rice sample is extracted from Sentinel-2 images based on Sentinel-2 reflectance data combined with LSWI and NDVI indices, including: calculating the difference between the LSWI index value and NDVI index value of each pixel during the flooding period of the paddy field, and extracting the second rice sample based on the difference using a threshold segmentation method. The union of the first and second rice samples is taken to form a potential rice distribution map; pixels are extracted from the distribution map as rice sample labels, and pixels from the remaining area are extracted as non-rice sample labels. VH time series values and SWIR1 band reflectance of each labeled sample point are extracted from Sentinel-1 and Sentinel-2 images respectively to construct the first training dataset and the second training dataset. The first classification model and the second classification model were trained using the first training dataset and the second training dataset, respectively. Sentinel-1 and Sentinel-2 image data of the study area were acquired, and the first rice probability distribution and the second rice probability distribution were obtained using the first classification model and the second classification model, respectively. The first rice probability distribution and the second rice probability distribution are linearly weighted and fused to obtain a probability distribution map. The probability distribution map is then constrained by the area using data from a statistical yearbook to obtain an optimal probability threshold. Based on the optimal probability threshold, a rice distribution map is extracted from the probability distribution map.
2. The method according to claim 1, characterized in that, The , , The rice index was calculated after normalization.
3. The method according to claim 1, characterized in that, Pixels are extracted from the distribution map using a random sampling method.
4. The method according to claim 1, characterized in that, The first and second classification models are random forest models.
5. The method according to claim 1, characterized in that, Sentinel-1 and Sentinel-2 image data of the study area were acquired. The temporal VH polarization data of the Sentinel-1 image and the temporal SWIR1 band reflectance data of the Sentinel-2 image were used to generate time-series datasets using the median synthesis method, which were then used as inputs to the first classification model and the second classification model, respectively.
6. The method according to claim 1, characterized in that, The probability distributions of the first and second rice varieties are linearly weighted and fused with equal weights.
7. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.
9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Rice early-stage remote sensing identification method based on planting probability
CN115861844A
Large-scale regional rice remote sensing classification method under complex conditions
CN116109943A