Spartina alterniflora detection method and device
Through hyperspectral remote sensing technology and random forest model, the problem that traditional monitoring methods cannot monitor the distribution of mutual flower rice grass in real time is solved, and accurate identification and management of mutual flower rice grass is achieved.
Patent Information
- Application Number
- CN202510427968.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional artificial ground investigations cannot comprehensively and in real time monitor the distribution and spread of mutual flower rice and grass, resulting in the invasion of ecological threats to their invasion being unable to be effectively managed.
Through hyperspectral remote sensing technology, low-dimensional spectral data are extracted using principal component analysis method, random forest model is constructed, and the distribution and diffusion of mutual flower rice grass is identified in combination with decision tree algorithm.
Accurate monitoring of mutual flower rice grass is achieved, the accuracy and reliability of identification is improved, and reliable data support is provided to help manage and control its growth and diffusion.
Smart Images

Figure CN120339838A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of remote sensing image monitoring, and particularly to a method and device for detecting Spartina alterniflora. Background Art
[0002] Spartina alterniflora is a perennial herbaceous plant native to South America. It is commonly found in water environments such as rivers, lakes, and paddy fields, and belongs to invasive species in China. Spartina alterniflora has the ability to grow and reproduce rapidly, and can quickly occupy waters and wetlands, posing a serious threat to China's ecological system. Therefore, it is necessary to monitor the invasion of Spartina alterniflora. Traditional monitoring methods mainly rely on manual ground surveys, and cannot comprehensively and real-time monitor the distribution and spread of Spartina alterniflora. Summary of the Invention
[0003] In view of this, this application provides a method and device for detecting Spartina alterniflora, so as to accurately monitor Spartina alterniflora through hyperspectral.
[0004] Specifically, this application is implemented through the following technical solutions:
[0005] The first aspect of this application provides a method for detecting Spartina alterniflora, and the method includes:
[0006] Obtain hyperspectral data of a target area, and extract first spectral data from the hyperspectral data based on the principal component analysis method. The spectral information dimension of the first spectral data is lower than that of the hyperspectral data;
[0007] Determine a projection direction according to the first spectral data, and remove spectral data other than the first spectral data from the hyperspectral data to form second spectral data;
[0008] Generate a spectral matrix based on the second spectral data, and each row in the matrix is the characteristic spectral data of a band;
[0009] Traverse each row of spectral data in the spectral matrix, project the selected spectral data into the one-dimensional space of the first spectral data according to the projection direction, and calculate the projection length after projection of each row of spectral data;
[0010] Select the target row spectral data corresponding to the minimum projection length from the spectral matrix, and add it to the first spectral data;
[0011] Judge whether the first spectral data meets a preset condition. If not, return to the step of traversing each row of spectral data in the spectral matrix;
[0012] Use the first spectral data as a sample to construct a random forest model; wherein, the random forest model is composed of multiple decision trees;
[0013] Identify the image to be detected in the target area based on the random forest model to obtain the identification result of Spartina alterniflora.
[0014] The second aspect of the present application provides a Spartina alterniflora detection device, which includes a generation module, a calculation module, and a construction module; wherein,
[0015] The generation module is used to obtain the hyperspectral data of the target area, and extract the first spectral data from the hyperspectral data based on the principal component analysis method. The spectral information dimension of the first spectral data is lower than that of the hyperspectral data;
[0016] The generation module is further used to determine the projection direction according to the first spectral data, and remove the spectral data other than the first spectral data from the hyperspectral data to form the second spectral data;
[0017] The generation module is further used to generate a spectral matrix based on the second spectral data, and each row in the matrix is the characteristic spectral data of a band;
[0018] The calculation module is used to traverse each row of spectral data in the spectral matrix, project the selected spectral data into the one-dimensional space of the first spectral data according to the projection direction, and calculate the projection length after the projection of each row of spectral data;
[0019] The calculation module is further used to select the target row spectral data corresponding to the minimum projection length from the spectral matrix and add it to the first spectral data;
[0020] The calculation module is further used to determine whether the first spectral data meets the preset conditions. If not, return to the step of traversing each row of spectral data in the spectral matrix;
[0021] The construction module is used to use the first spectral data as a sample to construct a random forest model; wherein, the random forest model is composed of multiple decision trees;
[0022] The construction module is further used to identify the image to be detected in the target area based on the random forest model to obtain the identification result of Spartina alterniflora.
[0023] The Spartina alterniflora detection method and device provided by this application, after obtaining the hyperspectral data of the target area, use the principal component analysis method to extract the low-dimensional first spectral data to reduce the data complexity and improve the processing efficiency. Based on this, determine the projection direction and construct the second spectral data and the spectral matrix. Through the projection operation, screen the key spectral information and supplement it into the first spectral data, making it more comprehensively represent the spectral characteristics of the target area, which helps to more accurately capture the unique spectral information of Spartina alterniflora and improve the accuracy of recognition. Construct a random forest model with the processed first spectral data, so that the random forest model can learn more representative and discriminative feature patterns. This helps to improve the generalization ability of the model, reduce the risk of overfitting, and improve the classification accuracy and reliability of the model for new data (the image to be detected). Based on the above optimized classification model, identify the image to be detected in the target area, and can more accurately judge whether there is Spartina alterniflora in the image and its specific location and scope. The accurate recognition results are of great significance for the monitoring, prevention and control, and ecological research of Spartina alterniflora, etc., can provide reliable data support for relevant decisions, and help to take more effective measures to manage and control the growth and spread of Spartina alterniflora. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of the first embodiment of the Spartina alterniflora detection method provided by this application;
[0025] Figure 2 It is a schematic diagram of a remote sensing image exemplarily shown by this application;
[0026] Figure 3 It is a schematic structural diagram of the first embodiment of the Spartina alterniflora detection device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0028] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0029] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0030] Specific embodiments are given below to introduce the technical solution of the present application in detail.
[0031] Figure 1 This is a flow chart of Example 1 of the method for detecting Spartina alterniflora provided in this application. Figure 1 , the method provided in this embodiment may include:
[0032] S101 , obtaining hyperspectral data of a target area, and extracting first spectral data from the hyperspectral data based on a principal component analysis method, wherein the spectral information dimension of the first spectral data is lower than that of the hyperspectral data.
[0033] Specifically, the target area refers to the area where the invasion of Spartina alterniflora needs to be monitored, that is, the area where Spartina alterniflora may grow. The hyperspectral data of the target area can be obtained by drone hyperspectral remote sensing equipment. In specific implementation, the collection site is located in Chashan Island, Dinghai District, Zhoushan City, Zhejiang Province (29°58′N, 122°11′E), which belongs to the marine climate, with a coastline of 2.12km, a beach area of 1.099 square kilometers, and the highest point is 33m above sea level. The intertidal zone soil is mainly muddy mud of coastal saline soil, and the terrestrial soil is brown sand of coarse bone soil. The vegetation is mainly black pine forest and white grass.
[0034] The drone hyperspectral image is obtained by the FS60-UAV hyperspectral measurement system carried by the DJI M300 RTK drone. The spectral range is 400-1000nm, the spectral resolution is 2.5nm, it contains 1200 bands, the slit width is 25um, and the spectral resolution is 2.5nm. The image was collected on May 18, 2024. The drone flight mission was between 8:30 and 17:00. The weather was clear and windless. The drone was calibrated with a white board before flight. The flight altitude was set to 80m, the flight speed was 5m / s, the overlap was 80% for the route, 30% for the sideways, and the flight speed was 5m / s. After obtaining the drone hyperspectral image, image stitching and radiation correction were performed to extract the spectral reflectance of each band.
[0035] Figure 2 For a schematic diagram of remote sensing images exemplified in this application, please refer to Figure 2, The picture image is obtained through the visible light camera carried by the DJI M300 RTK drone for hyperspectral images. Its pixel count is 45 million pixels. During flight, the overlap rate is 80% in the forward direction and 70% in the lateral direction, with a flight speed of 12 m / s, collecting orthorectified two-dimensional images and three-dimensional oblique models.
[0036] Furthermore, the specific implementation steps for obtaining the hyperspectral data of the target area include:
[0037] (1) Determine the interfering objects in the target area, where the interfering objects are features with a similarity greater than a threshold to the appearance characteristics of Spartina alterniflora.
[0038] Specifically, the interfering objects refer to plants in the target area with a similarity greater than the threshold to the characteristics of Spartina alterniflora. Since Spartina alterniflora is an invasive species, other plants existing in the local ecosystem may become interfering objects. For example, on Chashan Island, there are other vegetation on the tidal flats in addition to Spartina alterniflora. If the spectral characteristics of these vegetation are more similar than the set threshold to the spectral characteristics of Spartina alterniflora, they will be determined as interfering objects. It should be noted that the threshold is set according to actual needs and is not limited in this embodiment. For example, in one embodiment, the threshold is 80%. When judging the similarity between the interfering objects and the appearance characteristics of Spartina alterniflora, it can be calculated and compared according to characteristics such as the shape of the spectral curve and the reflectance of specific bands using relevant statistical methods or machine learning algorithms. The specific implementation process for comparing the similarity between the interfering objects and the appearance characteristics of Spartina alterniflora can refer to the descriptions in related technologies and will not be elaborated here.
[0039] (2) Extract the characteristic data of the interfering objects and Spartina alterniflora under each spectrum.
[0040] Specifically, after obtaining the hyperspectral image of the target area, the characteristic data of the interfering objects and Spartina alterniflora in each spectrum in the hyperspectral image can be extracted through a hyperspectral imager. Different objects or plants will exhibit different characteristics under different spectra, such as reflectance, absorptance, etc.
[0041] (3) Determine the target spectrum based on the difference in the characteristic data of the two objects under the same spectrum.
[0042] Specifically, for each identical spectrum (such as the red light spectrum), calculate the difference (such as the difference value or ratio) between the interference object and the characteristic data of Spartina alterniflora under this spectrum, compare the magnitudes of this difference value under different spectra, sort the different spectra in descending order according to the characteristic data difference, and select the spectra ranked at the top from the sorting results as the target spectra. For example, in a possible implementation manner, the spectra include the infrared spectrum and the ultraviolet spectrum, etc. Among them, the value of the reflectance difference between the interference object and Spartina alterniflora under the infrared spectrum is larger than that under the ultraviolet spectrum, then the infrared spectrum is the target spectrum. It should be noted that the number of target spectra is set according to actual needs, and it is not limited in this embodiment.
[0043] (4) Obtain the hyperspectral data of the target area under the target spectrum.
[0044] Specifically, by using technical means such as hyperspectral imaging, collect the hyperspectral data of the target spectrum in the target area. These data contain the detailed information of various objects and plants in the target area under the target spectrum, and can be used for further analysis and processing in the future, such as accurately identifying Spartina alterniflora in the target area and excluding the influence of interference objects.
[0045] Furthermore, the principal component analysis method is used to perform dimensionality reduction processing on the hyperspectral data, and extract the low-dimensional first spectral data from the hyperspectral data. As much as possible, the main information in the hyperspectral data is retained in the first spectral data. For the specific implementation steps of extracting the first spectral data from the hyperspectral data using the principal component analysis method, please refer to the descriptions in related technologies and will not be elaborated here.
[0046] S102. Determine the projection direction according to the first spectral data, and remove the spectral data other than the first spectral data from the hyperspectral data to form the second spectral data.
[0047] Specifically, since the first spectral data is the first component extracted from the hyperspectral data by the principal component analysis method, and the main information in the hyperspectral data is retained in the first spectral data, the projection direction can be determined according to the first principal component direction in the first spectral data. Because the first principal component often contains the largest variance information in the original data and can reflect the main feature change direction of the data to the greatest extent.
[0048] Furthermore, the second spectral data is the other part of the hyperspectral data except for the spectral dimensions covered by the first spectral data. Constituting the second spectral data can separate the information other than the first spectral data alone, which helps to further explore the spectral feature differences of ground objects that are not fully reflected in the first spectral data. Especially for targets such as Spartina alterniflora that require precise identification, these additional features may be the key to distinguishing it from other similar ground objects.
[0049] S103. Generate a spectral matrix based on the second spectral data, where each row in the matrix is the characteristic spectral data of a band.
[0050] Specifically, organize and arrange the second spectral data to generate a spectral matrix. In this spectral matrix, each row represents the characteristic spectral data of a band. For example, if the second spectral data contains information on 10 different bands, then the spectral matrix will have 10 rows, and each row corresponds to the specific spectral characteristics (such as reflectivity, absorptivity, etc.) of a band.
[0051] In specific implementation, determine the band range covered by the second spectral data, clarify how many bands there are in total in the second spectral data, sort the bands in ascending or descending order of wavelength to ensure that each band has a fixed position and order in the matrix, and for each band, extract all the characteristic data corresponding to that band from the second spectral data.
[0052] Furthermore, create an empty matrix according to the number of bands and the number of data points per band. For example, if there are n bands and each band has m data points, then create an n×m matrix, and use the characteristic data of each band as a row and fill it into the spectral matrix in sequence. That is, the first row is filled with the characteristic data of the first band, the second row is filled with the characteristic data of the second band, and so on, finally obtaining a complete spectral matrix, where each row is the characteristic spectral data of a band.
[0053] S104. Traverse each row of spectral data in the spectral matrix, project the selected spectral data onto the one-dimensional space of the first spectral data in the projection direction, and calculate the projection length of each row of spectral data after projection.
[0054] Specifically, for each row of selected spectral data, project it onto the one-dimensional space formed by the first spectral data in the previously determined projection direction. The projection operation is a process of mapping high-dimensional data to a low-dimensional space (here it is a one-dimensional space). After the projection is completed, calculate the projection length of each row of spectral data after projection. This projection length reflects the characteristics of the spectral data of this row in the projection direction, and different projection lengths indicate the differences of different spectral data in this projection direction.
[0055] S105. Select the target row of spectral data corresponding to the minimum projection length from the spectral matrix and add it to the first spectral data.
[0056] Specifically, after calculating the projection lengths of each row of spectral data, compare these projection lengths to find the row of spectral data corresponding to the minimum projection length, that is, the target row spectral data, and add the target row spectral data to the first spectral data. The purpose of doing this is to gradually increase the information of the first spectral data so that it can more comprehensively represent the characteristics of the original hyperspectral data.
[0057] S106. Determine whether the first spectral data meets the preset conditions. If not, return to the step of traversing each row of spectral data in the spectral matrix.
[0058] Specifically, judge the first spectral data after adding the new target row spectral data to see if it meets the preset conditions. If not, return to step S104 and continue to repeat the above process, that is, select the spectral data corresponding to the minimum projection length again and add it to the first spectral data until the preset conditions are met. It should be noted that the preset conditions are set in advance according to actual needs and are not limited in this embodiment. For example, in a possible implementation, the preset condition is that the similarity between the first spectral data and the hyperspectral data reaches 90%.
[0059] S107. Use the first spectral data as a sample to construct a random forest model; wherein, the random forest model is composed of multiple decision trees.
[0060] Specifically, when the first spectral data meets the preset conditions, use it as sample data. The random forest model is composed of multiple decision tree models. A decision tree is a commonly used machine learning classification algorithm. It constructs a tree-shaped model through learning the sample data for classifying and predicting new data. Here, taking the construction of a decision tree model as an example, the steps for constructing the decision tree model are described as follows. The specific implementation steps include:
[0061] (1) Extract the features of the first spectral data to obtain the first feature set;
[0062] Specifically, the features of the first spectral data can be extracted through deep learning algorithms. These features can be various statistics in the spectral data (such as mean, variance, standard deviation, etc.), reflectance values of specific bands, ratios between different bands, etc. By analyzing and processing the first spectral data, these features are extracted and combined into a set, that is, the first feature set.
[0063] (2) Calculate the discrimination degree between each sub-feature in the first feature set;
[0064] Specifically, the discrimination degree is an index to measure the ability of a feature to distinguish Spartina alterniflora from interfering objects, and the discrimination degree can be calculated through information gain, information gain ratio, Gini index, etc. Taking information gain as an example, it measures the discrimination ability of a feature by calculating the change in the information entropy of the data set before and after using a certain feature for partitioning. Information entropy is an index to measure the uncertainty in the data set. The greater the information gain, the more the uncertainty of the data set is reduced after using this feature for partitioning, that is, the higher the discrimination degree of this feature.
[0065] When specifically implemented, it includes:
[0066] 2.1. Calculate the class probabilities of each sub-feature based on conditional entropy, and cluster the first feature set;
[0067] Specifically, conditional entropy is used to measure the uncertainty of a random variable under certain known conditions. For each sub-feature in the first feature set, its conditional entropy is calculated to determine the probability distribution of this sub-feature under different classes. Further, based on the calculated class probabilities of each sub-feature, the first feature set is clustered. Clustering is a process of grouping similar data points in a data set into one class. Various clustering algorithms can be used here, such as the K-Means clustering algorithm, hierarchical clustering algorithm, etc. Taking the K-Means algorithm as an example, it randomly initializes K clustering centers, and then assigns each data point (i.e., sub-feature) to the nearest cluster according to the distance between the data point and the clustering center, and then updates the clustering centers. This process is repeated until the clustering centers no longer change significantly. Through clustering, the sub-features in the first feature set are divided into different classes.
[0068] 2.2. Select the centers of each class after clustering the first feature set as the second feature set;
[0069] Specifically, the centers of each class after clustering represent the average features of all sub-features in this cluster. Select the center of each cluster as a new feature, and these features form the second feature set. The purpose of doing this is to replace the original first feature set with fewer representative features, reduce the number of features, and at the same time retain the main feature information of the data for more efficient subsequent processing.
[0070] 2.3. Select the optimal feature discrimination degree calculation method according to the number of features in the second feature set. Among them, the calculation amounts of different feature discrimination degree calculation methods are different, and the number of features is inversely proportional to the calculation amount of the selected optimal feature discrimination degree calculation method;
[0071] Specifically, different feature discrimination calculation methods (such as information gain, information gain ratio, Gini index, etc.) vary in computational complexity. Some methods have a large computational amount, while some have a small computational amount. Therefore, an appropriate feature discrimination calculation method can be selected according to the number of features in the second feature set. If the number of features in the second feature set is large, in order to improve computational efficiency, a method with a small computational amount will be chosen; if the number of features is small, a method with relatively higher computational accuracy but a relatively large computational amount can be selected. For example, when there are only a few features in the second feature set, the information gain ratio, which is relatively complex but can more accurately measure feature discrimination, can be selected; when the number of features is large, the Gini index, which is relatively simple to calculate, can be selected.
[0072] (1) Conditional entropy
[0073]
[0074] where k is the number of element categories in the set;
[0075] p i is the probability density of the element category;
[0076] b. Definition of conditional entropy:
[0077]
[0078] H(Y|X = x i ) is the entropy of feature x under condition Y i ;
[0079] (2) Information gain:
[0080] g(D, X) = H(D) - H(D|X);
[0081] where g(D, X) is the information gain;
[0082] H(D) is the entropy of set D;
[0083] H(D|X) is the conditional entropy of D under the condition of feature X.
[0084] (3) Information gain ratio:
[0085] When using information gain as the feature for dividing the data set, it tends to select features with more values
[0086]
[0087] where, g R (D, X) is the information gain ratio;
[0088] g(D, X) is the information gain
[0089] H X (D) is the entropy of set D, denoted as H(D).
[0090] It can be seen from the above formula that when the number of eigenvalue value categories is large, the denominator will become larger, which can significantly reduce the information gain and tend to select features with more values.
[0091] (4) Gini coefficient
[0092] When calculating the information gain, a large number of logarithmic operations are involved. Therefore, in order to simplify the model without losing the advantages of the information gain model, the Gini coefficient can be used to replace it in part.
[0093]
[0094] Among them, K is the number of categories of the target variable.
[0095] p i is the proportion of the i-th type of sample in the dataset.
[0096] (5) Gini gain
[0097] Similar to the information gain, if the Gini coefficient of the dataset is subtracted from the Gini coefficient obtained after dividing the dataset according to the feature, the Gini gain (coefficient) is obtained. Obviously, the better the feature used for division, the greater the Gini gain. Based on the division of the dataset by each previous feature, the corresponding Gini gain can be obtained.
[0098] G(D,X) = Gini(D) - Gini(D,X);
[0099] Among them, Gini(D) is the initial impurity of dataset D, which measures the degree of chaos in the class distribution of the dataset.
[0100] Gini(D,X) is the weighted average of the Gini impurities of each child node after dividing dataset D using feature X.
[0101] 2.4. Calculate the discrimination degrees of each sub-feature in the second feature set based on the optimal feature discrimination degree calculation method.
[0102] Specifically, after determining the optimal feature discrimination degree calculation method, apply this method to each sub-feature in the second feature set to calculate its discrimination degree. By calculating the discrimination degree, it is possible to understand the ability of each sub-feature to distinguish Spartina alterniflora from other objects. For example, in one embodiment, the information gain is selected as the optimal feature discrimination degree calculation method. Then, according to the calculation formula of the information gain, combined with the sample data and corresponding class labels of the second feature set, calculate the information gain of each sub-feature, and this information gain value is the discrimination degree of the sub-feature. The calculated discrimination degrees will be used for subsequent feature selection and model construction operations.
[0103] (3) Select node features according to the discrimination degree, where the discrimination degree of the node features is greater than that of non-node features;
[0104] Specifically, according to the calculated discrimination degrees of each sub-feature, select features with higher discrimination degrees from the first feature set as node features. Node features are the features used to divide the data set when constructing a decision tree. Selecting features with a discrimination degree greater than that of non-node features as node features is to ensure that at each node of the decision tree, using this feature for division can minimize the uncertainty of the data set, so that the decision tree can classify the data more effectively.
[0105] (4) Use the node features as the nodes of the tree to construct multiple decision tree models;
[0106] Specifically, in a decision tree, each node represents a feature, and the branches extending from this node represent different values or value ranges of this feature. Use the selected node features as the nodes of the decision tree, and divide the data set according to the value conditions of the node features to generate sub-nodes. For example, if the node feature is the reflectivity of a certain wavelength band, and the value range of this reflectivity can be divided into three intervals: high, medium, and low, then three branches are extended from this node, corresponding to these three value intervals respectively. By continuously selecting node features and performing divisions, a complete decision tree model is gradually constructed.
[0107] (5) Use the first spectral data as samples to train multiple decision tree models to obtain the random forest model.
[0108] Specifically, after constructing the structure of the decision tree model, use the first spectral data as training samples to train the decision tree. During the training process, the decision tree adjusts its own parameters (such as the division rules of nodes, etc.) according to the features and corresponding class labels in the sample data (in the task of identifying Spartina alterniflora, the class labels may be "Spartina alterniflora" or "non-Spartina alterniflora") to minimize the classification error. Through the training of the first spectral data, the decision tree model can learn the patterns and rules in the data, thus becoming a random forest model that can classify new data. When new spectral data is input, this random forest model can judge whether this data belongs to Spartina alterniflora according to the rules obtained from the training.
[0109] Furthermore, the implementation steps of constructing a random forest model may further include:
[0110] (1) Determine the current level of decision tree calculation, and match the target indicators sorted in the current level matching stage;
[0111] Specifically, a decision tree is a model with a tree structure, consisting of multiple levels, and each level contains several nodes. In the process of constructing a decision tree, it is first necessary to clarify which level of the decision tree calculation is currently in. For example, at the beginning, it is the first level where the root node is located. As the tree is constructed and extends downward, the number of levels gradually increases.
[0112] Different levels may correspond to different stages, and each stage has a target metric for sorting. These target metrics are the criteria for measuring the effect of node features on dataset partitioning. Common target metrics include information gain, information gain ratio, Gini index, etc. For example, in the initial level of decision tree construction, information gain may be selected as the target metric because it can quickly perform a preliminary partitioning of the data; while in deeper levels, the information gain ratio may be selected to avoid overfitting problems. According to the current level, the corresponding target metric is matched in order to evaluate the node features subsequently.
[0113] (2) Calculate the target metric values of each node among the candidate nodes at the current level, and sort them according to the magnitude of the target metric values to obtain the optimal partitioning feature; where the optimal partitioning feature is the feature that makes the target metric perform optimally in the hyperspectral dataset.
[0114] Specifically, at the current level, there are multiple candidate nodes, and each candidate node corresponds to a feature in the hyperspectral dataset. For each candidate node, according to the selected target metric (such as information gain), its target metric value is calculated. Taking information gain as an example, it is necessary to calculate the change in information entropy of the dataset before and after partitioning using this feature based on the current dataset and the value situation of this feature, so as to obtain the information gain value. After calculating the target metric values of all candidate nodes, they are sorted according to the magnitude of these values. The feature with the optimal target metric value (such as the largest information gain) is determined as the optimal partitioning feature. This optimal partitioning feature is the feature that can make the target metric perform best at the current level, that is, the feature with the best partitioning effect on the dataset, and it will be used to partition the dataset of the current node.
[0115] (3) Create a node according to the optimal partitioning feature, create nodes for the next-level branches based on the sorting result, and return to the step of determining the current level of decision tree calculation until the number of levels of the decision tree reaches the threshold.
[0116] Specifically, after determining the optimal partitioning feature, a new node is created based on this feature. This node will serve as the partitioning basis for the current level. According to the different values or value ranges of this feature, the data set corresponding to the current node is divided into different subsets. Based on the previous sorting result of the target index values of the candidate nodes, nodes for the next-level branches are created. For example, if there are multiple candidate nodes, they are sequentially used as nodes for the next-level branches in the sorting order. Each branch corresponds to a subset after the upper-level node is partitioned according to the optimal partitioning feature. After creating the next-level branch nodes, the steps for determining the calculation of the current level of the decision tree are returned, and the next level is continued to be processed. Repeat the above process, continuously determining the current level, calculating the target index value, selecting the optimal partitioning feature, creating nodes and branches, until the number of levels of the decision tree reaches a pre-set threshold. It should be noted that the threshold is set according to actual needs and is not limited in this embodiment. For example, in one embodiment, the threshold is 10 levels.
[0117] Specifically, the cut-off degree of the model is to control the depth of the decision tree. No matter which classification method is used, it is necessary to control the depth of the tree to avoid overfitting. Pruning needs to determine the optimal subtree through the validation sample. Pruning control is also a shortcoming of the decision tree and needs to be manually controlled. It can be controlled by the following formula:
[0118] L α = Gini(t) * |T| → α|T leaf |;
[0119] Where, L α is the final loss;
[0120] Gini(t) is the Gini coefficient of the current node;
[0121] |T| is the number of samples contained in the current node;
[0122] |T leaf | is the number of nodes generated after partitioning;
[0123] α is the set preference coefficient.
[0124] Here, the larger α is, the greater the "penalty" degree for dividing into more sub-nodes, and vice versa. It is used to control the model to prevent overfitting. α is obtained by making the loss functions before and after pruning equal.
[0125] S108. Identify the to-be-detected image of the target area based on the random forest model to obtain the identification result of Spartina alterniflora.
[0126] Specifically, before identifying the to-be-detected image of the target area based on the random forest model, it further includes:
[0127] (1) Set the pruning index and pruning coefficient of the decision tree;
[0128] Specifically, the pruning index is a criterion used to measure whether a decision tree node or branch should be pruned. Common pruning indices include error rate, Gini index, information gain, etc. For example, when using the error rate as the pruning index, it is considered whether the classification error of the decision tree for the training data or validation data will be within an acceptable range after cutting off a certain node or branch.
[0129] Furthermore, the pruning coefficient is a parameter used to control the degree of pruning. It plays a regulating role in the pruning process and is usually a numerical value. A larger pruning coefficient will make the decision tree pruning more strict, cutting off more nodes and branches, thus obtaining a simpler sub-decision tree; a smaller pruning coefficient results in a relatively lighter pruning degree, retaining more nodes and branches. For example, in pruning based on the error rate, the pruning coefficient can be used to balance the relationship between the increase in error caused by cutting off nodes and the reduction in model complexity.
[0130] In this step, the pruning index can be set based on training error, information gain, Gini index, number of nodes, and depth of the decision tree, etc. The pruning coefficient can also be determined based on empirical values. It should be noted that the pruning index and pruning coefficient are set according to actual needs and are not limited in this embodiment.
[0131] (2) Prune the decision tree according to the pruning index and the pruning coefficient to obtain multiple sub-decision trees;
[0132] Specifically, according to the set pruning index and pruning coefficient, starting from the leaf nodes of the decision tree, each node is recursively checked upward to perform pruning operations on the original decision tree. For each node, calculate the impact on the performance of the decision tree if the node is cut off (i.e., turned into a leaf node) according to the pruning index. If this impact is within the range allowed by the pruning coefficient (for example, the increase in error after cutting off the node does not exceed the threshold specified by the pruning coefficient), then cut off the node. Further, by changing the value of the pruning coefficient, multiple sub-decision trees with different degrees of pruning can be obtained. For example, first set a smaller pruning coefficient for one pruning to obtain a sub-decision tree; then increase the pruning coefficient and perform pruning again to obtain another sub-decision tree. In this way, a series of sub-decision trees with different complexities can be obtained.
[0133] For example, in a possible implementation, use the Gini coefficient for pruning. For the first group with L = 0 and L = 1, each having 5, and the probability values are all 0.5. The Gini coefficient of this group is obtained as 0.5 through the Gini coefficient calculation formula. Similarly, the Gini coefficient of the second group is 0.32, and the Gini coefficient of the third group is 0. It can be seen from this that the smaller the Gini coefficient, the lower the degree of chaos.
[0134] Next, in the branches of the tree, by arranging in different orders, the Gini coefficient can be reduced. As shown in the figure below, it is reduced from 0.5 to 0.12. There are 10 original data points, and the Gini coefficient is 0.5, which is in a chaotic state. Through the judgment of a certain attribute, the data is divided into two categories. There are 4 on the left and 6 on the right. A new Gini coefficient is obtained through weighted average. Compared with the original state, the purity is improved, so this attribute is selected as excellent.
[0135] (3) Comprehensively evaluate the multiple sub-decision trees, and use the sub-decision trees whose evaluation results meet the preset conditions as the final random forest model.
[0136] Specifically, indicators such as accuracy, recall rate, F1 value, and mean squared error can be used to evaluate the multiple sub-decision trees. The data set is divided into a training set and a validation set. The training set is used to train the sub-decision trees, and then the above evaluation indicators are calculated on the validation set. By comparing the comprehensive evaluation results of the multiple sub-decision trees, select the sub-decision trees whose evaluation results meet the preset conditions as the final random forest model. The random forest model achieves a good balance between the generalization ability and the fitting degree of the data, and can classify or predict new data more accurately. For example, in the task of identifying Spartina alterniflora, it can more accurately judge whether there is Spartina alterniflora in the image and its location and other information. It should be noted that the preset conditions are set according to actual needs, and in this embodiment, they are not limited.
[0137] Furthermore, the specific implementation steps for identifying the to-be-detected image of the target area based on the random forest model to obtain the identification result of Spartina alterniflora include:
[0138] (1) Obtain the first classification result of each pixel in the to-be-detected image based on the random forest model;
[0139] Specifically, for the to-be-detected image, the spectral data of each pixel of it (these spectral data are usually obtained through hyperspectral imaging technology and contain information such as the reflectance of the pixel in multiple bands) are input into the random forest model. The random forest model analyzes and judges the spectral data of each pixel according to the rules and patterns obtained from its training, and outputs the classification result of whether the pixel belongs to Spartina alterniflora or other objects, that is, the first classification result. After obtaining the first classification result, it also includes:
[0140] 1.1 Extract the boundary pixel points of Spartina alterniflora according to the first classification result;
[0141] Specifically, traverse all the pixels in the first classification result to find those pixel points whose classification results are different from those of the surrounding pixels. Specifically, if a pixel marked as Spartina alterniflora has pixels marked as non-Spartina alterniflora among its adjacent pixels (usually the four adjacent pixels above, below, left, and right, and the eight adjacent pixels including the diagonal ones can also be considered), then this pixel is identified as a boundary pixel point of Spartina alterniflora; vice versa. In this way, all the boundary pixel points of Spartina alterniflora are extracted from the first classification result.
[0142] 1.2. Obtain the pixel points in the adjacent areas within a preset distance on both sides of the boundary pixel points, and combine them with the boundary pixel points to form a region to be confirmed. The preset distance is determined by the appearance difference degree between the interference objects and Spartina alterniflora in the image to be detected;
[0143] Specifically, for each extracted boundary pixel point, obtain the pixel points in the adjacent areas on both sides (i.e., the Spartina alterniflora side and the non-Spartina alterniflora side) according to the preset distance, and combine these boundary pixel points with the obtained pixel points in the adjacent areas on both sides to form a region to be confirmed. This region to be confirmed contains the pixel points that may have inaccurate classifications in the first classification result. It should be noted that the determination of the preset distance is related to the appearance difference degree between the interference objects and Spartina alterniflora in the image to be detected. If the appearance difference between the interference objects and Spartina alterniflora is small, in order to more accurately identify the boundary, a larger preset distance may be set to include more potentially confusing pixel points for further confirmation; on the contrary, if the appearance difference is large, the preset distance can be set smaller. In this embodiment, it is not limited.
[0144] 1.3. Use the first classification results of the adjacent pixel points of the region to be confirmed as the recognition prompt;
[0145] Specifically, in the region to be confirmed, the adjacent pixel points already have classification marks in the first classification result, and the first classification results of these adjacent pixel points can be used as the recognition prompt. For example, if most of the adjacent pixel points on one side of the region to be confirmed are marked as Spartina alterniflora, then when re-identifying this region to be confirmed, this information can be referred to, and the pixel points with similar features in this region tend to be identified as Spartina alterniflora.
[0146] 1.4. Based on the random forest model and the recognition prompt, re-identify the classification result of the region to be confirmed, and fuse the recognition results of the non-region-to-be-confirmed area in the first classification result and the recognition result of the region to be confirmed as the recognition result of Spartina alterniflora.
[0147] Specifically, for the area to be confirmed, the spectral data of its pixel points are input into the classification model again for classification. In the classification process, the output result of the classification model is adjusted in combination with the recognition prompt information obtained previously. If the recognition prompt shows that a part of the area to be confirmed should be inclined to the classification of Spartina alterniflora, and the preliminary output result of the classification model does not match it, then the result can be corrected according to certain rules (such as setting a weight to include the influence of the recognition prompt in the calculation of the classification result).
[0148] Further, after obtaining the re-recognition result of the area to be confirmed, it is fused with the recognition result of the non-to-be-confirmed area in the first classification result. Specifically, for the non-to-be-confirmed area, the first classification result is directly used; for the area to be confirmed, the re-recognition result is used. The fusion result finally obtained is the recognition result of Spartina alterniflora in the entire image to be detected, which can improve the accuracy of recognition, especially in the recognition of the boundary area of Spartina alterniflora.
[0149] (2) determining the area of Spartina alterniflora according to the first classification result;
[0150] Specifically, the number of pixels marked as Spartina alterniflora in the first classification result is counted, the ground area corresponding to each pixel in the actual scene is determined according to the spatial resolution of the first classification result, and the area of Spartina alterniflora in the area covered by the image to be detected can be obtained by multiplying the number of pixels marked as Spartina alterniflora by the actual area corresponding to each pixel.
[0151] (3) Determine the invasion status of the Spartina alterniflora according to the area and the shooting time of the image to be detected.
[0152] Specifically, compare the area of Spartina alterniflora in the current time shooting image with the area of Spartina alterniflora in the previous time shooting image. If the area increases, it means that Spartina alterniflora is expanding, and there may be a situation where the invasion situation is aggravated; if the area decreases, it may indicate that the control measures taken are effective and the invasion situation is alleviated; if the area remains basically unchanged, it means that the distribution of Spartina alterniflora is relatively stable. At the same time, the severity of the invasion situation can also be more accurately assessed in combination with the speed of area change (e.g., the amount of area increase or decrease per unit time). For example, if the area of Spartina alterniflora increases significantly in a short period of time, it indicates that the invasion situation is more severe.
[0153] Furthermore, in a possible implementation, identifying the image to be detected in the target area based on the random forest model further includes:
[0154] (1) acquiring an image to be detected, and determining a distance difference between a geographical location of the image to be detected and the target area;
[0155] Specifically, the image to be detected can be obtained through satellite remote sensing, drone photography, etc., and the geographical location information of the image to be detected can be determined. The distance difference between the geographical location of the image to be detected and the target area can be calculated using algorithms or tools related to the Geographic Information System (GIS).
[0156] (2) Predict the difference in the appearance characteristics of Spartina alterniflora in the image to be detected and the appearance characteristics of Spartina alterniflora in the target area based on the distance difference;
[0157] Specifically, since the environmental conditions (such as climate, soil, water, etc.) in different regions may vary with distance, and these environmental factors will affect the growth of Spartina alterniflora, which in turn leads to differences in its appearance characteristics. A relationship model between the distance difference and the difference in the appearance characteristics of Spartina alterniflora can be established based on existing knowledge or experience. For example, through field observations and studies of Spartina alterniflora in multiple different regions, it is found that the farther away from the target area, the greater the changes in the appearance characteristics such as plant height, leaf shape, and color of Spartina alterniflora. Using the established relationship model, based on the calculated distance difference, predict the difference in the appearance characteristics of Spartina alterniflora in the image to be detected and the appearance characteristics of Spartina alterniflora in the target area. For example, it may be predicted that the leaves of Spartina alterniflora in the image to be detected may be narrower and the color may be lighter than those in the target area.
[0158] (3) If the difference is greater than the recognition interval of the random forest model, adjust the number of nodes of the random forest model based on multiple images at the same location of the image to be detected;
[0159] Specifically, the recognition interval refers to the range of the appearance characteristics of Spartina alterniflora that the random forest model can accurately recognize. Compare the predicted difference in appearance characteristics with the recognition interval of the random forest model. If the difference is greater than the recognition interval, it means that the appearance characteristics of Spartina alterniflora in the image to be detected are quite different from those in the target area, and the original random forest model may not be able to accurately recognize. To improve the adaptability of the random forest model, obtain multiple images at the same location of the image to be detected (which can be taken at different times or from different angles). By analyzing the characteristics of Spartina alterniflora in these multiple images, adjust the number of nodes of the random forest model. The specific method of adjusting the number of nodes can be carried out according to the construction principle of the decision tree. For example, if it is found that the original random forest model is too simple to distinguish these quite different characteristics of Spartina alterniflora, the number of nodes can be increased to make the decision tree more complex, thereby improving its recognition ability for different appearance characteristics; conversely, if the random forest model is too complex resulting in overfitting, and the characteristics shown in these images have relatively small differences, the number of nodes can be appropriately reduced.
[0160] (4) Identify the image to be detected based on the adjusted random forest model.
[0161] Specifically, after adjusting the number of nodes of the random forest model in the previous steps, a random forest model that is more adaptable to the Spartina alterniflora characteristics in the image to be detected is obtained. The image to be detected is input into the adjusted random forest model. According to its adjusted structure and parameters, the random forest model analyzes and judges the input data, and outputs the recognition result of Spartina alterniflora in the image to be detected, that is, judges whether there is Spartina alterniflora in the image and its specific position and range and other information.
[0162] The Spartina alterniflora detection method provided in this embodiment realizes dimensionality reduction by extracting the first spectral data from the hyperspectral data through the principal component analysis method, reduces the data processing complexity while retaining the main information, and improves the calculation efficiency. Then, according to the first spectral data, the projection direction is determined, the second spectral data and the spectral matrix are constructed, and valuable spectral data is gradually screened and added to the first spectral data through the projection operation, so that the first spectral data can more comprehensively and accurately represent the spectral characteristics of the target area. Then, a classification model based on the decision tree is constructed with the first spectral data as a sample, and a random forest is constructed by using the classification ability of the decision tree, so that it can learn the patterns and rules in the data. Finally, the recognition result of Spartina alterniflora is obtained by recognizing the image to be detected based on this classification model. Due to the effectiveness of the previous data processing, the classification model has high accuracy and reliability, and can more accurately identify the Spartina alterniflora in the target area, providing strong support for subsequent monitoring, prevention and control and other work.
[0163] Corresponding to the foregoing embodiment of a Spartina alterniflora detection method, the present application also provides an embodiment of a Spartina alterniflora detection device.
[0164] Figure 3 It is a schematic structural diagram of Embodiment 1 of the Spartina alterniflora detection device provided by the present application. Please refer to Figure 3 , the device provided in this embodiment includes a generation module 310, a calculation module 320, and a construction module 330; wherein,
[0165] The generation module 310 is configured to obtain hyperspectral data of a target area, and extract first spectral data from the hyperspectral data based on the principal component analysis method, and the spectral information dimension of the first spectral data is lower than that of the hyperspectral data;
[0166] The generation module 310 is further configured to determine a projection direction according to the first spectral data, and remove spectral data other than the first spectral data from the hyperspectral data to form second spectral data;
[0167] The generation module 310 is further configured to generate a spectral matrix based on the second spectral data, and each row in the matrix is characteristic spectral data of a band;
[0168] The calculation module 320 is configured to traverse each row of spectral data in the spectral matrix, project the selected spectral data onto the one-dimensional space of the first spectral data in accordance with the projection direction, and calculate the projection length of each row of spectral data after projection;
[0169] The calculation module 320 is further configured to select the target row of spectral data corresponding to the minimum projection length from the spectral matrix and add it to the first spectral data;
[0170] The calculation module 320 is further configured to determine whether the first spectral data meets a preset condition. If not, it returns to the step of traversing each row of spectral data in the spectral matrix;
[0171] The construction module 330 is configured to use the first spectral data as a sample to construct a random forest model; wherein, the random forest model is composed of multiple decision trees;
[0172] The construction module 330 is further configured to identify the to-be-detected image of the target area based on the random forest model to obtain the recognition result of Spartina alterniflora.
[0173] The device of this embodiment can be used to execute Figure 1 the steps of the method embodiment shown. The specific implementation principle and process are similar and will not be elaborated here.
[0174] For the implementation process of the functions and roles of each unit in the above device, refer to the implementation process of the corresponding steps in the above method for details. It will not be elaborated here.
[0175] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0176] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the scope of protection of this application.
Claims
1. A method for detecting Spartina alterniflora, characterized in that, The method includes: Obtaining hyperspectral data of a target area, and extracting first spectral data from the hyperspectral data based on the principal component analysis method, where the spectral information dimension of the first spectral data is lower than that of the hyperspectral data; Determining a projection direction according to the first spectral data, and removing spectral data other than the first spectral data from the hyperspectral data to form second spectral data; Generating a spectral matrix based on the second spectral data, where each row in the matrix is characteristic spectral data of a band; Traversing each row of spectral data in the spectral matrix, projecting the selected spectral data onto the one-dimensional space of the first spectral data according to the projection direction, and calculating the projection length after projection of each row of spectral data; Selecting the target row of spectral data corresponding to the minimum projection length from the spectral matrix and adding it to the first spectral data; Judging whether the first spectral data meets a preset condition, if not, returning to the step of traversing each row of spectral data in the spectral matrix; Using the first spectral data as a sample to construct a random forest model; where the random forest model consists of multiple decision trees; Identifying a to-be-detected image of the target area based on the random forest model to obtain an identification result of Spartina alterniflora.
2. The method according to claim 1, characterized in that The constructing a random forest model using the first spectral data as a sample; where the random forest model consists of multiple decision trees; the construction process of the decision tree includes: Extracting features of the first spectral data to obtain a first feature set; Calculating the discrimination degree between each sub-feature in the first feature set; Selecting a node feature according to the discrimination degree, where the discrimination degree of the node feature is greater than that of non-node features; Using the node feature as a node of the tree to construct multiple decision tree models; Using the first spectral data as a sample to train multiple decision tree models to obtain the random forest model.
3. The method according to claim 2, wherein The calculating the discrimination degree between each sub-feature in the first feature set; includes: Calculating the class probability of each sub-feature based on conditional entropy and clustering the first feature set; Selecting the center of each category after clustering of the first feature set as a second feature set; Selecting an optimal feature discrimination degree calculation method according to the number of features in the second feature set, where different feature discrimination degree calculation methods have different calculation amounts, and the number of features is inversely proportional to the calculation amount of the selected optimal feature discrimination degree calculation method; Calculating the discrimination degree between each sub-feature in the second feature set based on the optimal feature discrimination degree calculation method.
4. The method according to claim 1, characterized in that, The constructing a random forest model using the first spectral data as a sample; further includes: Determining the current level of decision tree calculation and matching a target index sorted in the current level matching stage; Calculating the target index values of each node in the candidate nodes at the current level, sorting them according to the size of the target index values to obtain an optimal partitioning feature; where the optimal partitioning feature is the feature that makes the target index perform optimally in the hyperspectral dataset; Create nodes according to the optimal partitioning features, create nodes for the next-level branches based on the sorting results, and return the step of determining the current level of the decision tree calculation until the number of levels of the decision tree reaches the threshold.
5. The method according to claim 1, characterized in that, The obtaining of the hyperspectral data of the target area includes: Determine the interfering objects in the target area, where the interfering objects are features with a similarity greater than the threshold to the appearance features of Spartina alterniflora; Extract the feature data of the interfering objects and Spartina alterniflora under each spectrum; Determine the target spectrum based on the difference in the feature data of the two objects under the same spectrum; Obtain the hyperspectral data of the target area under the target spectrum.
6. The method according to claim 1, wherein Before identifying the image to be detected in the target area based on the random forest model, it includes: Set the pruning index and pruning coefficient of the decision tree; Prune the decision tree according to the pruning index and the pruning coefficient to obtain multiple sub-decision trees; Comprehensively evaluate the multiple sub-decision trees, and use the sub-decision trees whose evaluation results meet the preset conditions as the final decision tree model.
7. The method according to claim 1, wherein Identifying the image to be detected in the target area based on the random forest model to obtain the identification result of Spartina alterniflora includes: Obtain the first classification result of each pixel in the image to be detected based on the random forest model; Determine the area of Spartina alterniflora according to the first classification result; Determine the invasion situation of Spartina alterniflora according to the area and the shooting time of the image to be detected.
8. The method according to claim 7, wherein After obtaining the first classification result of each pixel in the image to be detected based on the random forest model, it includes: Extract the boundary pixel points of Spartina alterniflora according to the first classification result; Obtain the pixel points in the adjacent area within a preset distance on both sides of the boundary pixel points, and combine them with the boundary pixel points to form an area to be confirmed, where the preset distance is determined by the appearance difference between the interfering objects and Spartina alterniflora in the image to be detected; Use the first classification result of the adjacent pixel points of the area to be confirmed as the identification prompt; Based on the random forest model and the identification prompt, re-identify the classification result of the area to be confirmed, and fuse the identification result of the non-area to be confirmed in the first classification result and the identification result of the area to be confirmed as the identification result of Spartina alterniflora.
9. The method according to claim 1, wherein Identifying the image to be detected in the target area based on the random forest model to obtain the identification result of Spartina alterniflora includes: Obtain the image to be detected, and judge the distance difference between the geographical location of the image to be detected and the target area; Predict the difference between the appearance features of Spartina alterniflora in the image to be detected and the appearance features of Spartina alterniflora in the target area based on the distance difference; If the difference is greater than the identification interval of the random forest model, adjust the number of nodes of the random forest model based on multiple images at the same position of the image to be detected; Identify the image to be detected based on the adjusted random forest model.
10. A Spartina alterniflora monitoring device, characterized in that, The device includes a generation module, a calculation module, and a construction module; where The generation module is configured to obtain hyperspectral data of a target area, and extract first spectral data from the hyperspectral data based on the principal component analysis method, where the spectral information dimension of the first spectral data is lower than that of the hyperspectral data; The generation module is further configured to determine a projection direction according to the first spectral data, and remove spectral data other than the first spectral data from the hyperspectral data to form second spectral data; The generation module is further configured to generate a spectral matrix based on the second spectral data, where each row in the matrix is the characteristic spectral data of a band; The calculation module is configured to traverse each row of spectral data in the spectral matrix, project the selected spectral data into the one-dimensional space of the first spectral data according to the projection direction, and calculate the projection length after projection of each row of spectral data; The calculation module is further configured to select the target row of spectral data corresponding to the minimum projection length from the spectral matrix and add it to the first spectral data; The calculation module is further configured to determine whether the first spectral data meets a preset condition. If not, return to the step of traversing each row of spectral data in the spectral matrix; The construction module is configured to use the first spectral data as a sample to construct a random forest model; where the random forest model is composed of multiple decision trees; The construction module is further configured to identify the to-be-detected image of the target area based on the random forest model to obtain the identification result of Spartina alterniflora.