A method for extracting spatial distribution information of Shanlan rice producing areas based on optical and radar imagery

By processing optical and radar image data on the GEE platform, and combining HSV color space denoising, surface temperature inversion, and random forest model classification, the problem of low accuracy in extracting spatial distribution information of *Rhizophora stylosa* was solved, and high-precision identification and distribution information acquisition of *Rhizophora stylosa* were achieved.

CN119251665BActive Publication Date: 2026-07-17INST OF AGRI ENVIRONMENT & SOIL HAINAN ACAD OF AGRI SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF AGRI ENVIRONMENT & SOIL HAINAN ACAD OF AGRI SCI
Filing Date
2024-08-06
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in extracting spatial distribution information of *Rhizophora stylosa* in complex planting areas such as low hills and mountains. Excessive noise in optical and radar image data leads to insufficient extraction accuracy.

Method used

Cloud masking was performed on the GEE platform to synthesize optical and radar image data. The data was then converted to the HSV color space for denoising. Feature categories were selected using surface temperature inversion and separation thresholding. Classification was performed by combining fully constrained least squares hybrid pixel decomposition and a random forest model.

Benefits of technology

It improves the extraction accuracy of spatial distribution information of mountain rice, enabling timely and accurate acquisition of planting area and growth information, enhancing the recognition of edge details, and improving recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251665B_ABST
    Figure CN119251665B_ABST
Patent Text Reader

Abstract

This invention provides a method for extracting spatial distribution information of *Rhizophora stylosa* producing areas based on optical and radar imagery, comprising the following steps: S11, acquiring optical and radar image data of *Rhizophora stylosa*; S12, converting the RGB images in the *Rhizophora stylosa* image data to the HSV color space for denoising; S13, retrieving the surface temperature of concentrated *Rhizophora stylosa* planting areas to improve the recognition accuracy of *Rhizophora stylosa*; S14, using a separation threshold method to select a preferred feature set with high separation degree from the feature categories of concentrated *Rhizophora stylosa* planting areas; S15, using a hybrid pixel decomposition method to analyze the preferred feature set to obtain a training feature set; S16, using a random forest model to classify the training feature set to obtain the spatial distribution information of *Rhizophora stylosa*. This invention, based on optical and radar imagery, improves the recognition accuracy of *Rhizophora stylosa* through image denoising and surface temperature retrieval, and then efficiently obtains the spatial distribution information of *Rhizophora stylosa* using a hybrid pixel decomposition method and a random forest model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mountain rice identification technology, and in particular to a method for extracting spatial distribution information of mountain rice production areas based on optical and radar images. Background Technology

[0002] Mountain rice, a unique type of dryland rice native to the Li ethnic group in Hainan, is a valuable asset preserved by the Li people through long-term production practices. Its spatial distribution varies dramatically due to natural conditions, agricultural development, and urbanization. Therefore, research on the spatiotemporal distribution and changes of mountain rice is of great significance. With the continuous advancement and widespread application of remote sensing technology, its application in spatial dynamic monitoring has been significantly improved. Compared with traditional crop monitoring methods, remote sensing technology has unique advantages such as shorter processing time, higher efficiency, wider monitoring range, and high temporal continuity and repeatability. However, for most tropical and subtropical regions, obtaining a sufficient number of images of mountain rice during its critical periods is extremely difficult.

[0003] Optical and radar imagery data are unaffected by weather conditions and can sensitively respond to changes in plant development and soil moisture in *Rhizophora stylosa*, making them important data sources for monitoring *Rhizophora stylosa* growth in cloudy and foggy areas. Although existing classification schemes can achieve high classification accuracy for *Rhizophora stylosa*, the excessive noise in the optical and radar imagery data itself still results in insufficient extraction accuracy. Summary of the Invention

[0004] To overcome the problem of low accuracy in extracting the area of ​​*Rhizophora stylosa* in structurally complex planting areas such as low hills and mountains, this invention provides a method for extracting spatial distribution information of *Rhizophora stylosa* production areas based on optical and radar imagery, comprising the following steps: S11. Perform cloud masking on the GEE (Google Earth Engine) platform, and synthesize images with cloud cover less than a preset cloud cover threshold to obtain optical and radar image data of Shanlandao. S12. Convert the RGB image in the image data obtained in step S11 to the HSV color space, use the 8-dimensional spatial vector information of a single pixel to form a polygonal pixel collection, and denoise the image data of Shanlan rice based on this polygonal pixel collection to preserve the edge detail features of Shanlan rice. S13. Surface temperature inversion: A universal single-channel algorithm is used to invert the surface temperature of the concentrated planting area of ​​*Syzygium sambac* after denoising in step S12, so as to improve the identification accuracy of *Syzygium sambac*. S14. Extract the feature categories of the concentrated planting area of ​​mountain orchid rice in step S13. The feature categories include spectral bands, vegetation index, water index and topographic and texture features. Then, use the separation threshold method to screen out the preferred feature set of the concentrated planting area of ​​mountain orchid rice with a feature category separation degree greater than the preset separation degree threshold. S15. Using the fully constrained least squares mixed pixel decomposition method, the selected feature set in step S14 is analyzed to obtain the abundance map of land types that are easily confused with the spectrum of *Oryza sativa*, and this map is used as a training feature set to be integrated into the classification system. S16. Construct a random forest model to classify the training feature set from step S15 and obtain the spatial distribution information of *Rhizophora stylosa*.

[0005] Furthermore, in step S12, the process of forming a polygonal pixel collection using the 8-dimensional spatial vector information of a single pixel includes the following steps: S21. Represent the 8-dimensional spatial vector of a single pixel as (u, v, h, s, l, r, g, b), which is used to express the pixel points of the image data of Shanlan rice. (u, v) represents the spatial position vector, (h, s, l) is the HSV color feature, and (r, g, b) is the RGB color feature. Based on the color features (h, s, l, r, g, b), calculate the number of local center seed points n. S22. The number of cluster centers is calculated from S21 as the number of local centers n of the pixels in the Shanlan rice image data. The number of sides of the polygon pixel set is taken, and the standard spacing of the set is calculated. S23. Using the cluster center as the local center, search for the color feature distance and spatial distance between all pixels within a range of twice the standard distance and the cluster center, and update the pixel label information; S24. Start iterative calculation for each pixel. After repeated iterations, when the distance error between each cluster center and its surrounding centers converges to 0.1, the iteration terminates and the polygon pixel set is determined.

[0006] Furthermore, in step S12, the denoising of the image data of *Symplocos stenoptera* based on polygon pixel collections includes the following steps: S31. Set the search box Similar to a fixed-size bounding box, the search box contains a collection of multiple polygonal pixels; S32. Traverse the polygon pixel set, compare its center pixel j with the pixel i to be processed, and perform a similarity comparison of all pixels between i and j to obtain the weight of pixel j. ; S33. Finally, perform a weighted average of all polygonal pixel sets within the search box to obtain the pixel estimate of the denoised image at point i. ; S34. Output the denoised image data of the mountain rice.

[0007] Furthermore, in step S14, the method for obtaining the preferred feature set by using the separation threshold method is as follows: S41. Use the Jeffries-Matusita (JM) distance of SEaTH to determine the separability of feature categories. Its value range is [0, 2]. A value close to 0 indicates that the two categories are almost indistinguishable on a certain feature; a value of 2 indicates that the two categories can be completely distinguished on a certain feature. S42. Visual interpretation is aided by combining field survey data with image data of concentrated planting areas of mountain rice; S43. Calculate the separation degree of the feature categories, retain the two feature categories with the highest separation degree, including duplicate feature categories, and obtain the preferred feature set of the concentrated planting area of ​​Shanlan rice.

[0008] Furthermore, in step S15, the step of analyzing the dataset of the concentrated planting area of ​​Shanlan rice is as follows: S51. Using the optimized feature set of the concentrated planting area of ​​Shanlan rice as the input dataset, principal component analysis is performed on the feature variables through minimum noise separation to achieve data dimensionality reduction and estimate image noise points. S52. Calculate the clean pixel index based on image noise; S53. Use n-dimensional visualization tools to select pure pixels for each land type; S54. Based on fully constrained least squares mixed pixel decomposition, vegetation abundance maps of Cymbidium goeringii, other crops, and vegetables are obtained. S55. Use vegetation abundance maps as an aid to distinguish land types that are prone to spectral confusion.

[0009] Furthermore, in step S16, the random forest model construction steps are as follows: S61. Feature selection during node splitting is performed using the Gini index, where the Gini index represents the probability of a random sample in the sample set being misclassified. A set is constructed. Gini index The expression is:

[0010] In the formula, The number of categories in the training samples; The category to which a randomly selected sample from set D belongs. The probability of; S62, If set Based on characteristics Whether a certain value α is taken is divided into and Two parts, then in terms of features Under the condition of set Gini index The expression is:

[0011] in, Represents a set The number of samples in Represents a set The number of samples in Represents a set The number of samples in; S63. The importance of the training feature set is evaluated using the method of reducing average impurity, and its expression is:

[0012] in, This represents the number of decision trees in the random forest model. , Let A and B be the Gini indexes of the sets D before and after the t-th decision tree is partitioned by feature A.

[0013] Furthermore, in step S43, the step of evaluating the importance of the training feature set using the method of reducing average impurity is as follows: S71. The importance of each feature in the training feature set of the random forest model is calculated by the method of reducing average impurity, and then sorted in order from high to low. S72. When sorting according to the importance of each feature, the first feature is selected in the first time, the first two features are selected in the second time, and so on, to obtain a random forest model with 20 different feature combinations for a single time phase and 100 different feature combinations for time series images. S73. Calculate the out-of-bag (OOB) data for each model, and determine the optimal feature combination after comprehensively considering the accuracy and complexity of the random forest model, thereby reducing the model complexity while ensuring classification accuracy.

[0014] Furthermore, in step S16, the method for extracting the spatial distribution information of *Rhizophora stylosa* is as follows: S81. Using a random forest model, N samples with replacement are randomly selected from the training feature set as training samples using the bootstrap method. S82, through Sub-sample extraction and training can obtain A decision tree model; S83. Randomly select at each node of each decision tree. Each feature is then divided into internal nodes; S84. Integrate all decision tree model results and use majority voting to determine the final classification result of the spatial distribution information of *Syzygium serratum*.

[0015] The beneficial effects of this invention are as follows: This invention relates to a method for extracting spatial distribution information of *Rhizophora stylosa* production areas based on optical and radar imagery. First, image data of *Rhizophora stylosa* from optical and radar images is acquired using a GEE platform, enabling timely and accurate acquisition of planting area and growth information. The RGB images in the *Rhizophora stylosa* image data are converted to the HSV color space for denoising, effectively preserving the edge details of the *Rhizophora stylosa*, thus improving inversion accuracy and providing a foundation for subsequent extraction and classification, thereby enhancing the accuracy of spatial distribution information extraction. Second, a universal single-channel algorithm is used to invert the surface temperature of concentrated *Rhizophora stylosa* planting areas after denoising, further improving the identification accuracy. Finally, the least squares mixed pixel decomposition method and a random forest model are used for classification to obtain the spatial distribution information of *Rhizophora stylosa*. This invention denoises the images before extracting and classifying *Rhizophora stylosa*, improving image recognition accuracy and thus enhancing the accuracy of spatial distribution information extraction. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort. Figure 1 This is a flowchart of a method for extracting spatial distribution information of Shanlan rice production areas based on optical and radar imagery, according to the present invention. Detailed Implementation

[0017] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0018] This invention provides a method for extracting spatial distribution information of Shanlan rice producing areas based on optical and radar imagery, comprising the following steps: S11. Cloud masking was performed on the GEE (Google Earth Engine) platform. Images with cloud cover below a preset threshold were synthesized. Combined with the main phenological stages of single-season rice in the study area, optical and radar image data of single-season rice were obtained. Optical remote sensing data is easily affected by weather factors such as precipitation and fog, making it difficult to extract the spatial distribution of single-season rice in rainy areas. Spaceborne synthetic aperture radar (SAR) achieves target observation by actively emitting electromagnetic waves and recording the characteristics of surface scattering echoes. It can continuously image the ground regardless of weather conditions and has all-day, all-weather observation capabilities. Therefore, optical and radar data were synthesized in time series according to the phenological stages of single-season rice. Based on the synthesized data, data fusion was performed to obtain synthesized image data, which can solve the problem of severe lack of optical data and low accuracy of single images for crop extraction in southern regions due to geographical location, fog, etc.

[0019] S12. Convert the RGB image in the image data obtained in step S11 to the HSV color space. Utilize the 8-dimensional spatial vector information of a single pixel to form a polygonal pixel collection. Based on this polygonal pixel collection, denoise the image data of *Rhizophora stylosa* while preserving its edge details. This algorithm uses 8-dimensional vectors as pixel-level information representation, which better preserves the edge details of the image in terms of noise resistance. This algorithm can achieve image denoising under different devices and different interference conditions.

[0020] In this embodiment, paddy fields filled with water and mixed with wild rice plants in the RGB image are enhanced and displayed as blue-purple. Water bodies are displayed as blue, forests as dark green, and features such as buildings and bare land as grayish-white or yellowish-brown. Using RGB images can highlight wild rice fields on a map, thus distinguishing them from other land cover types. In the RGB color space, changes in lighting conditions can cause variations in the values ​​of different land cover features on the same color channel, making threshold selection difficult. The HSV color space divides color information into three components: hue (H), saturation (S), and value (V). In image processing, hue corresponds to the color information of land cover features and can be used to distinguish different land cover types. Saturation reflects the purity and vividness of color and can be used to distinguish the material and characteristics of different land cover features. Value represents the brightness information of the image and can provide information about the reflectivity and radiance of land cover features. In the HSV color system, the V channel is relatively insensitive to changes in lighting, making threshold selection more stable. Separating value information from color information can better adapt to changes in lighting conditions and improve the robustness of recognition. Mapping the RGB channels in an image separately and converting them to the HSV color system for recognition can make full use of the information in the image and improve the accuracy of recognition.

[0021] In step S12, forming a polygonal pixel collection using the 8-dimensional spatial vector information of a single pixel includes the following steps: S21. Represent a single pixel's 8-dimensional spatial vector as (u, v, h, s, l, r, g, b), used to represent the pixel points in the image data of *Symplocos serrata*, where (u, v) represents the spatial location vector, (h, s, l) is the HSV color feature, and (r, g, b) is the RGB color feature. Based on the color features (h, s, l, r, g, b), the steps for calculating the number of local center seed points n are as follows: Create k points as initial centroids, where k is a random number in the image pixel count (width * height).

[0022] When the cluster assignment result of any point changes: For each data point in the feature set, calculate the Euclidean distance from the centroid to the color features of all data points, and assign the data point to the nearest cluster. For each cluster, calculate the mean of all points in the cluster and use the mean as the new centroid.

[0023] When the number of clusters is less than k: For each cluster, calculate the total error, and continue clustering on the given cluster; Calculate the total error of splitting the cluster into two, and select the cluster that minimizes the error for the partitioning operation.

[0024] S22. The number of cluster centers obtained from S21 is the number of local centers n of the pixels in the Shanlan rice image data. The number of pixels in the image is M=R×C. The number of sides of the polygon pixel set is taken, and the standard spacing of the set is calculated. The number of sides of the polygon pixel set ranges from 4 to 8, and can be set according to the final effect. When the value is less than 6, the pixels will not be smooth enough. When the value is greater than 6, the amount of calculation will increase. Based on this, this embodiment selects 6 for calculation, and the obtained effect is relatively good. The formula for calculating the standard spacing of the set is:

[0025] Where M represents the number of pixels in the image, R represents the number of rows in the image, and C represents the number of columns in the image. Represents the standard distance of a set. Indicates the total number of pixels. This represents the number of edges of a set.

[0026] S23. Using the cluster center as the local center, search for the color feature distance and spatial distance between all pixels within a range of twice the standard distance and the cluster center, and update the pixel label information; S24. The calculation is repeated for each pixel. After repeated iterations, the iteration terminates when the distance error between each cluster center and its surrounding centers converges to 0.1. This determines the polygonal pixel set, which serves as the basic image processing method. This not only improves the denoising effect but also enhances the computational efficiency of the enhancement algorithm. A convergence value greater than 0.1 leads to poor clustering results, while a convergence value less than 1 causes convergence to stop. Therefore, this embodiment terminates the iteration when the distance error converges to 0.1. After the iteration terminates, the final cluster centers, cluster labels, and cluster ranges are obtained.

[0027] Furthermore, in step S12, the denoising of the image data of *Symplocos stenoptera* based on polygon pixel collections includes the following steps: S31. Set the search box Similar to a fixed-size bounding box, the search box contains a collection of multiple polygonal pixels; S32. Traverse the polygon pixel set, compare its center pixel j with the pixel i to be processed, and perform a similarity comparison of all pixels between i and j to obtain the weight of pixel j. ; S33. Finally, perform a weighted average of all polygonal pixel sets within the search box to obtain the pixel estimate of the denoised image at point i. ; S34. Output the denoised image data of the mountain rice.

[0028] In this embodiment, the purpose of setting up a search box is to overcome the slow speed of global search. To improve image processing speed, the image to be processed is divided into several small blocks, and then a search is performed within the search box. After the search is completed, filtering is performed to change its current value, replacing the original value with the new value. The estimated value, so further research is needed. Update the search bar by retrieving values ​​from each search box.

[0029] S13. Surface temperature inversion: A universal single-channel algorithm is used to invert the surface temperature of the concentrated planting area of ​​*Syzygium sambac* after denoising in step S12, so as to improve the identification accuracy of *Syzygium sambac*. In this embodiment, the physiological and biochemical effects of different growth processes of *Rhizophora stylosa* can significantly increase or decrease the LST (land surface temperature) of paddy fields. LST data, especially LST data during the critical growth stages of *Rhizophora stylosa*, can effectively improve the identification accuracy of *Rhizophora stylosa*. Existing research shows that single-channel algorithms are more advantageous than split-window algorithms for LST data inversion. Therefore, this invention uses a universal single-channel algorithm to invert LST data.

[0030] S14. Extract the feature categories of the concentrated planting area of ​​*Rhizophora stylosa* in step S13. The feature categories include spectral bands, vegetation index, water index, and topographic and textural features. Then, use the separation threshold method to select the preferred feature set of concentrated planting areas of *Rhizophora stylosa* whose feature category separation degree is greater than a preset separation degree threshold. In the process of identifying *Rhizophora stylosa*, the feature categories used are selected. Considering that *Rhizophora stylosa* has plant characteristics and water-required growth characteristics, if all selected feature categories are used for classification, it will not only increase the classification time, but also affect the classification accuracy due to data redundancy. In order to reduce data redundancy, the separation threshold method is used for screening.

[0031] Furthermore, in step S14, the method for obtaining the preferred feature set by using the separation threshold method is as follows: S41. Use the Jeffries-Matusita (JM) distance of SEaTH to determine the separability of feature categories. Its value range is [0, 2]. A value close to 0 indicates that the two categories are almost indistinguishable on a certain feature; a value of 2 indicates that the two categories can be completely distinguished on a certain feature. S42. Visually interpret mountain rice, forest land, water area, construction land and other land types by combining field survey data with image data of concentrated mountain rice planting areas; S43. Calculate the separation degree of the feature categories, retain the two feature categories with the highest separation degree, including duplicate feature categories, and obtain the preferred feature set of the concentrated planting area of ​​Shanlan rice.

[0032] The formula for calculating the resolution is:

[0033] in, Resolution; This is the Bach distance, used to calculate the distance between factors; , It is the mean of a certain feature distribution of two different samples; Let be the standard deviation of a certain feature distribution of two different samples.

[0034] S15. Using the fully constrained least squares mixed pixel decomposition method, the selected feature set in step S14 is analyzed to obtain the abundance map of land types that are easily confused with the spectrum of *Oryza sativa*, and this map is used as a training feature set to be integrated into the classification system. Furthermore, in step S15, the step of analyzing the dataset of the concentrated planting area of ​​Shanlan rice is as follows: S51. Using the optimized feature set of the concentrated planting area of ​​Shanlan rice as the input dataset, principal component analysis is performed on the feature variables through minimum noise separation to achieve data dimensionality reduction and estimate image noise points. S52. Calculate the clean pixel index based on image noise. The hyperparameters in the calculation process include the number of iterations, the iteration unit, and the threshold coefficient, which are set to 10000, 250, and 2.5, respectively. S53. Use n-dimensional visualization tools to select pure pixels for each land type; S54. Based on fully constrained least squares mixed pixel decomposition, vegetation abundance maps of Cymbidium goeringii, other crops, and vegetables are obtained. S55. Use vegetation abundance maps as an aid to distinguish land types that are prone to spectral confusion.

[0035] In this embodiment, the abundance maps of land types (vegetables and other crops) that are easily confused with *Rhizophora stylosa* are obtained using the fully constrained least squares mixed pixel decomposition method. These maps are then used as training feature sets and incorporated into the classification system. The above method is used to obtain the abundance maps of *Rhizophora stylosa*, other crops, and vegetable bases at various periods. The vegetation abundance obtained through mixed pixel decomposition can effectively distinguish the three land types that are easily confused with each other.

[0036] S16. Construct a random forest model to classify the training feature set from step S15 and obtain the spatial distribution information of *Rhizophora stylosa*. The random forest algorithm is an ensemble learning method that uses decision trees as the basic classifier and combines Bagging ensemble learning theory with the random subspace method.

[0037] Furthermore, in step S16, the random forest model construction steps are as follows: S61. In random forests, the Gini index is used for feature selection during node splitting when constructing decision trees. The Gini index represents the probability that a random sample in the sample set will be misclassified. The smaller the Gini index, the higher the purity of the set and the lower the probability of misclassification; conversely, the more impure the set. Constructing the set... Gini index The expression is:

[0038] In the formula, The number of categories in the training samples; The category to which a randomly selected sample from set D belongs. The probability of; S62, If set Based on characteristics Whether a certain value α is taken is divided into and Two parts, then in terms of features Under the condition of set Gini index The expression is:

[0039] in, Represents a set The number of samples in Represents a set The number of samples in Represents a set The number of samples in; S63. This shows that when constructing a random forest model, the greater the reduction in the Gini index after partitioning by a certain feature, the purer the resulting set becomes, and the more important that feature is in the model. Therefore, the importance of the training feature set can be evaluated using the method of reducing average impurity, expressed as:

[0040] in, This represents the number of decision trees in the random forest model. , Let A and B be the Gini indexes of the sets D before and after the t-th decision tree is partitioned by feature A.

[0041] Furthermore, in step S43, the step of evaluating the importance of the training feature set using the method of reducing average impurity is as follows: S71. The importance of each feature in the training feature set of the random forest model is calculated by the method of reducing average impurity, and then sorted in order from high to low. S72. When sorting according to the importance of each feature, the first feature is selected in the first time, the first two features are selected in the second time, and so on, to obtain a random forest model with 20 different feature combinations for a single time phase and 100 different feature combinations for time series images. S73. Calculate the out-of-bag (OOB) data for each model, and determine the optimal feature combination after comprehensively considering the accuracy and complexity of the random forest model, thereby reducing the model complexity while ensuring classification accuracy.

[0042] Furthermore, in step S16, the method for extracting the spatial distribution information of *Rhizophora stylosa* is as follows: S81. The random forest model uses the bootstrap method to randomly select N samples with replacement from the training feature set as training samples. Therefore, about 37% of the samples will not be selected. These samples are called out-of-bag (OOB) data, which can be used to evaluate the performance of the random forest model. S82, through Sub-sample extraction and training can obtain A decision tree model; S83. Randomly select at each node of each decision tree. indivual( < , (Total number of features) and perform internal node partitioning; S84. Integrate all decision tree model results and use majority voting to determine the final classification result of the spatial distribution information of *Syzygium serratum*.

[0043] In this embodiment, the OOB error is most stable when the number of decision trees is 100. Therefore, this paper ultimately sets Ntree to 100 and Mtry to the square root of the total number of input features.

[0044] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. This invention provides a method for extracting spatial distribution information of Shanlan rice producing areas based on optical and radar imagery, characterized in that, Includes the following steps: S11. Perform cloud masking on the GEE (Google Earth Engine) platform, and synthesize images with cloud cover less than a preset cloud cover threshold to obtain optical and radar image data of Shanlandao. S12. Convert the RGB image in the image data obtained in step S11 to the HSV color space, use the 8-dimensional spatial vector information of a single pixel to form a polygonal pixel collection, and denoise the image data of Shanlan rice based on this polygonal pixel collection to preserve the edge detail features of Shanlan rice. S13. Surface temperature inversion: A universal single-channel algorithm is used to invert the surface temperature of the concentrated planting area of ​​*Syzygium sambac* after denoising in step S12, so as to improve the identification accuracy of *Syzygium sambac*. S14. Extract the feature categories of the concentrated planting area of ​​mountain orchid rice in step S13. The feature categories include spectral bands, vegetation index, water index and topographic and texture features. Then, use the separation threshold method to screen out the preferred feature set of the concentrated planting area of ​​mountain orchid rice with a feature category separation degree greater than the preset separation degree threshold. S15. Using the fully constrained least squares mixed pixel decomposition method, the selected feature set in step S14 is analyzed to obtain the abundance map of land types that are easily confused with the spectrum of *Oryza sativa*, and this map is used as a training feature set to be integrated into the classification system. S16. Construct a random forest model to classify the training feature set from step S15 and obtain the spatial distribution information of *Rhizophora stylosa*.

2. The method for extracting spatial distribution information of Shanlan rice producing areas based on optical and radar imagery according to claim 1, characterized in that, in step S12, the step of forming a polygonal pixel collection using the 8-dimensional spatial vector information of a single pixel includes the following steps: S21. Represent the 8-dimensional spatial vector of a single pixel as (u, v, h, s, l, r, g, b), which is used to express the pixel points of the image data of Shanlan rice. (u, v) represents the spatial position vector, (h, s, l) is the HSV color feature, and (r, g, b) is the RGB color feature. Based on the color features (h, s, l, r, g, b), calculate the number of local center seed points n. S22. The number of cluster centers is calculated from S21 as the number of local centers n of the pixels in the Shanlan rice image data. The number of sides of the polygon pixel set is taken, and the standard spacing of the set is calculated. S23. Using the cluster center as the local center, search for the color feature distance and spatial distance between all pixels within a range of twice the standard distance and the cluster center, and update the pixel label information; S24. Start iterative calculation for each pixel. After repeated iterations, when the distance error between each cluster center and its surrounding centers converges to 0.1, the iteration terminates and the polygon pixel set is determined.

3. The method for extracting spatial distribution information of Shanlan rice production areas based on optical and radar imagery according to claim 1, characterized in that, in step S12, the denoising of the Shanlan rice image data based on polygon pixel ensembles includes the following steps: S31. Set the search box Similar to a fixed-size bounding box, the search box contains a collection of multiple polygonal pixels; S32. Traverse the polygon pixel set, compare its center pixel j with the pixel i to be processed, and perform a similarity comparison of all pixels between i and j to obtain the weight of pixel j. ; S33. Finally, perform a weighted average of all polygonal pixel sets within the search box to obtain the pixel estimate of the denoised image at point i. ; S34. Output the denoised image data of the mountain rice.

4. The method for extracting spatial distribution information of Shanlan rice producing areas based on optical and radar imagery according to claim 1, characterized in that, in step S14, the method for obtaining the preferred feature set by the separation threshold method is as follows: S41. Use the Jeffries-Matusita (JM) distance of SEaTH to determine the separability of feature categories. Its value range is [0, 2]. A value close to 0 indicates that the two categories are almost indistinguishable on a certain feature; a value of 2 indicates that the two categories can be completely distinguished on a certain feature. S42. Visual interpretation is aided by combining field survey data with image data of concentrated planting areas of mountain rice; S43. Calculate the separation degree of the feature categories, retain the two feature categories with the highest separation degree, including duplicate feature categories, and obtain the preferred feature set of the concentrated planting area of ​​Shanlan rice.

5. The method for extracting spatial distribution information of Shanlan rice production areas based on optical and radar imagery according to claim 1, characterized in that, in step S15, the step of analyzing the dataset of concentrated Shanlan rice planting areas is as follows: S51. Using the optimized feature set of the concentrated planting area of ​​Shanlan rice as the input dataset, principal component analysis is performed on the feature variables through minimum noise separation to achieve data dimensionality reduction and estimate image noise points. S52. Calculate the clean pixel index based on image noise; S53. Use n-dimensional visualization tools to select pure pixels for each land type; S54. Based on fully constrained least squares mixed pixel decomposition, vegetation abundance maps of Cymbidium goeringii, other crops, and vegetables are obtained. S55. Use vegetation abundance maps as an aid to distinguish land types that are prone to spectral confusion.

6. The method for extracting spatial distribution information of Shanlan rice producing areas based on optical and radar imagery according to claim 1, characterized in that, in step S16, the random forest model construction step is as follows: S61. Feature selection during node splitting is performed using the Gini index, where the Gini index represents the probability of a random sample in the sample set being misclassified. A set is constructed. Gini index The expression is: In the formula, The number of categories in the training samples; The category to which a randomly selected sample from set D belongs. The probability of; S62, If set Based on characteristics Whether a certain value α is taken is divided into and Two parts, then in terms of features Under the condition of set Gini index The expression is: in, Represents a set The number of samples in Represents a set The number of samples in Represents a set The number of samples in; S63. The importance of the training feature set is evaluated using the method of reducing average impurity, and its expression is: in, This represents the number of decision trees in the random forest model. , Let A and B be the Gini indexes of the sets D before and after the t-th decision tree is partitioned by feature A.

7. The method for extracting spatial distribution information of Shanlan rice producing areas based on optical and radar imagery according to claim 6, characterized in that, in step S43, the step of evaluating the importance of the training feature set using the method of reducing average impurity is as follows: S71. The importance of each feature in the training feature set of the random forest model is calculated by the method of reducing average impurity, and then sorted in order from high to low. S72. When sorting according to the importance of each feature, the first feature is selected in the first time, the first two features are selected in the second time, and so on, to obtain a random forest model with 20 different feature combinations for a single time phase and 100 different feature combinations for time series images. S73. Calculate the out-of-bag (OOB) data for each model, and determine the optimal feature combination after comprehensively considering the accuracy and complexity of the random forest model, thereby reducing the model complexity while ensuring classification accuracy.

8. The method for extracting spatial distribution information of Shanlan rice production areas based on optical and radar imagery according to claim 7, characterized in that, in step S16, the method for extracting the spatial distribution information of Shanlan rice is as follows: S81. Using a random forest model, N samples with replacement are randomly selected from the training feature set as training samples using the bootstrap method. S82, through Sub-sample extraction and training can obtain A decision tree model; S83. Randomly select at each node of each decision tree. Each feature is then divided into internal nodes; S84. Integrate all decision tree model results and use majority voting to determine the final classification result of the spatial distribution information of *Syzygium serratum*.