Hyperspectral data-based automatic extraction method and system for tetragol mongolica

By utilizing hyperspectral data technology, an automatic extraction method and system for *Tetraphyta brevicornu* was developed, which solved the problem that traditional methods struggle to obtain spatial distribution data of *Tetraphyta brevicornu*. This enabled rapid and accurate identification and spatial distribution mapping of *Tetraphyta brevicornu*, supporting scientific conservation decisions.

CN120976799APending Publication Date: 2025-11-18INNER MONGOLIA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511103477.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional ecological survey methods are insufficient for efficiently obtaining spatial distribution data of the rare and protected plant Tetracentron sinense, affecting the accuracy of project feasibility studies and environmental impact assessment reports.

Method used

Using hyperspectral data technology, a data acquisition platform was built to collect hyperspectral reflectance data of the ground canopy and hyperspectral imagery data from UAVs. First-order derivative and continuum removal transformations were performed to construct classification feature variables for *Tetraphyta tenuifolia*. Combined with vegetation index threshold classification methods and machine learning algorithms, the optimal classification model was constructed for the automatic extraction of *Tetraphyta tenuifolia*.

Benefits of technology

It enables rapid and accurate extraction of spatial distribution data of *Tetramorpha fruticosa*, significantly improving the efficiency and accuracy of identification and extraction, reducing the manpower and material resources required for on-site monitoring, and supporting scientific conservation decisions and spatial distribution mapping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976799A_ABST
    Figure CN120976799A_ABST
Patent Text Reader

Abstract

The invention provides a hyperspectral data-based automatic extraction method and a hyperspectral data-based automatic extraction system for tetracalamus mongolicus, and the method comprises the steps: building a data collection platform, collecting the ground canopy hyperspectral reflectivity of each species and the hyperspectral image data of an unmanned aerial vehicle in a target region, and synchronously collecting the geographic coordinates of each shrub; carrying out dimension reduction processing on the ground canopy high-spectral reflectivity data and constructing classification characteristic variables of the tetragos, wherein the classification characteristic variables comprise a vegetation index variable, a continuum removal parameter variable and a high-spectral characteristic parameter variable; screening key feature variables of the classification feature variables of the tetrazygote; a vegetation index threshold value classification method and a classification method based on machine learning are adopted to construct classification models to extract the tetragos in the target area, and an optimal classification model is selected; and based on the optimal classification model, combining the hyperspectral image data of the unmanned aerial vehicle to carry out identification and classification on the quadrangle of the target area, and obtaining space distribution data of the quadrangle. By means of the method, rapid, accurate and automatic extraction of the quadrangle tree can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic plant extraction, and in particular to a four-leaf clover automatic extraction method and system based on hyperspectral data. BACKGROUND

[0002] Four-leaf clover is a rare protected plant with extremely high ecological value. In the site selection and planning process of development projects such as mining areas and artificial buildings, it is difficult for traditional ecological investigation methods to efficiently obtain the spatial distribution data of the rare protected plant four-leaf clover, which restricts the accuracy of project feasibility studies and environmental impact assessment reports.

[0003] However, hyperspectral data, with its hundreds of continuous spectral channels, demonstrates high-precision resolution of target objects. Due to the selective absorption, reflection and scattering of incident electromagnetic waves by different shrubs based on their own characteristics, different spectral classification characteristics are generated on the hyperspectral reflectance curve of each species.

[0004] Therefore, if a fast and accurate automatic identification method based on hyperspectral technology can be used for four-leaf clover, it will have important significance for scientific protection decision and spatial distribution mapping of four-leaf clover. SUMMARY

[0005] Based on the above background, the purpose of the present application is to provide a four-leaf clover automatic extraction method based on hyperspectral data to solve the problem of fast and accurate extraction of four-leaf clover spatial distribution data.

[0006] To achieve the above purpose, the present application adopts the following technical solutions:

[0007] In a first aspect, the present application provides a four-leaf clover automatic extraction method based on hyperspectral data, comprising,

[0008] A data acquisition platform is built to collect ground canopy hyperspectral reflectance of each species and unmanned aerial vehicle hyperspectral image data in the target area, and the geographic coordinates of each shrub are collected synchronously;

[0009] The ground canopy hyperspectral reflectance data is subjected to first derivative transformation and continuum removal transformation processing;

[0010] Different transformed spectral curves are analyzed and four-leaf clover classification characteristic variables are constructed, the four-leaf clover classification characteristic variables including vegetation index variables, continuum removal parameter variables and hyperspectral characteristic parameter variables;

[0011] The four-leaf clover classification characteristic variables are subjected to importance analysis and sensitivity analysis, and key characteristic variables are screened;

[0012] The vegetation index threshold classification method and the machine learning-based classification method are respectively used to construct a classification model to extract Tetraena mongolica in the target region, and the classification effect of the classification model is evaluated according to preset evaluation indexes, and the best classification model is selected according to the evaluation result.

[0013] Based on the best classification model of ground canopy spectral data, the Tetraena mongolica in the target region is identified and classified by combining the unmanned aerial vehicle hyperspectral image data, and the spatial distribution data of the Tetraena mongolica is obtained.

[0014] In the second aspect, the application provides a Tetraena mongolica automatic extraction system based on hyperspectral data, comprising,

[0015] A data acquisition platform is used to acquire ground canopy hyperspectral reflectance of each species and unmanned aerial vehicle hyperspectral image data in the target region, and simultaneously acquire the geographic coordinates of each shrub.

[0016] A hyperspectral data processing module is used to perform first derivative transformation and continuum removal transformation processing on the ground canopy hyperspectral reflectance data.

[0017] A feature variable construction module is used to analyze different transformed spectral curves and construct Tetraena mongolica classification feature variables, wherein the Tetraena mongolica classification feature variables include vegetation index variables, continuum removal parameter variables and hyperspectral feature parameter variables.

[0018] A feature variable screening module is used to perform importance analysis and sensitivity analysis on the Tetraena mongolica classification feature variables, and screen key feature variables.

[0019] A classification model construction module is used to respectively adopt the vegetation index threshold classification method and the machine learning-based classification method to construct a classification model to extract Tetraena mongolica in the target region, and the classification effect of the classification model is evaluated according to preset evaluation indexes, and the best classification model is selected according to the evaluation result.

[0020] A regional Tetraena mongolica identification module is used to identify and classify Tetraena mongolica in the target region based on the best classification model of ground canopy spectral data, and obtain the spatial distribution data of the Tetraena mongolica by combining the unmanned aerial vehicle hyperspectral image data.

[0021] The application has the following beneficial effects:

[0022] The present application takes the rare protected plant Tetraena mongolica as the research object, uses a ground hyperspectral instrument and a UAV hyperspectral imaging technology to obtain ground canopy hyperspectral data of each species, uses a spectral feature selection method to reduce dimension of original data, selects classification characteristic variables from the original data, and constructs a Tetraena mongolica classification characteristic model based on a vegetation index threshold classification and a machine learning algorithm. Then, the performance indicators of each classification model are evaluated, and the best ground canopy classification model is selected to automatically identify and extract Tetraena mongolica. In addition, based on the best classification model of the ground canopy spectral data, the Tetraena mongolica in the target region is identified and classified by combining the UAV hyperspectral image data, and the spatial distribution data of Tetraena mongolica is obtained.

[0023] The method is suitable for quickly and accurately automatically extracting the spatial distribution information of Tetraena mongolica in a region, greatly reduces the manpower, material resources and investigation time in field ecological monitoring, significantly improves the identification and extraction efficiency and accuracy of Tetraena mongolica, and has important significance for scientific protection decision and spatial distribution mapping of Tetraena mongolica. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0025] Figure 1 The flow chart of the automatic extraction method of Tetraena mongolica based on hyperspectral data provided by the embodiment of the present application is shown in the figure.

[0026] Figure 2 is the effect diagram of hyperspectral reflectance data after SG filtering;

[0027] Figure 3 is the sampling route diagram of the embodiment of the present application;

[0028] Figure 4 is the ground object hyperspectral instrument reflectance collection flow chart;

[0029] Figure 5 is the UAV hyperspectral image collection flow chart;

[0030] Figure 6 is the shrub RTK geographic coordinate collection flow chart;

[0031] Figure 7 is the hyperspectral reflectance spectral curve of each species and the reflectance curve after different spectral transformation;

[0032] Figure 8 is the flow chart of two Tetraena mongolica classification methods of the embodiment of the present application;

[0033] Figure 9 is a spectral transformation vegetation index classification rule;

[0034] Figure 10 is a comparison of classification accuracies of different models of poplar;

[0035] Figure 11 is a spatial distribution map of poplar;

[0036] Figure 12 is a schematic diagram of a poplar automatic extraction system based on hyperspectral data provided by an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to further understand the present application, the preferred embodiments of the present application are described below in conjunction with the embodiments, but it should be understood that these descriptions are only for further illustrating the features and advantages of the present application, and are not limitations on the claims of the present application.

[0038] Embodiment 1

[0039] Referring to Figure 1 , the present embodiment provides a poplar automatic extraction method based on hyperspectral data, comprising,

[0040] S1: building a data acquisition platform, collecting ground canopy hyperspectral reflectance of each species and unmanned aerial vehicle hyperspectral image data in a target area, and synchronously collecting geographic coordinates of each shrub;

[0041] S2: performing first derivative transformation and continuum removal transformation processing on the ground canopy hyperspectral reflectance data;

[0042] S3: analyzing different transformed spectral curves and constructing poplar classification characteristic variables, the poplar classification characteristic variables including vegetation index variables, continuum removal parameter variables and hyperspectral characteristic parameter variables;

[0043] S4: performing importance analysis and sensitivity analysis on the poplar classification characteristic variables, and screening key characteristic variables;

[0044] S5: respectively using a vegetation index threshold classification method and a classification method based on machine learning to construct a classification model to extract poplar in the target area, combining a preset evaluation index to evaluate the classification effect of the classification model, and selecting the best classification model according to the evaluation result;

[0045] S6: based on the best classification model of ground canopy spectral data, combining unmanned aerial vehicle hyperspectral image data to identify and classify the poplar in the target area, and obtaining spatial distribution data of the poplar.

[0046] Specifically,

[0047] In step S1, a data acquisition platform is built, and ground canopy hyperspectral reflectance of each species and UAV hyperspectral image data in the target area are acquired, and the geographic coordinates of each shrub are synchronously acquired, including,

[0048] The data acquisition platform is built, including using a ground object hyperspectral instrument (ASD HandHeld 2, ASD HH2) to acquire ground canopy hyperspectral reflectance data, using a UAV hyperspectral platform (Unmanned Aerial Vehicle, UAV) to acquire hyperspectral image data, and using a South Surveying and Mapping polar point RTK (Real-Time Kinematic, RTK) to acquire geographic coordinates of each shrub.

[0049] The acquired data acquisition platform is processed, including performing reflectance conversion, outlier rejection, SG filtering and other pretreatments on the ground canopy reflectance data, the SG filtering being a mathematical filtering method using local polynomial least square fitting, and specific filtering effects being referred to Figure 2 ; the hyperspectral image data are processed through radiation, reflectance conversion, geometric correction, image stitching and low-pass filtering; when training a classification model, the samples are set as: the test set accounts for 1 / 3, and the training set accounts for 2 / 3.

[0050] The specific process is as follows:

[0051] S101: Using a ground object hyperspectral instrument (ASD HandHeld 2, ASD HH2) to acquire ground canopy hyperspectral reflectance data of each species, including,

[0052] S1011: Dividing a target area and arranging sampling lines in the divided area;

[0053] Referring to Figure 3 , in order to ensure the uniformity of sampling and the reliability of results, in the embodiment, due to the distribution characteristics of the sample vegetation, in order to ensure uniform sampling in the sample plot area, the target area is divided into three small areas of the same size, and four parallel sample lines are arranged in each small area in the same direction and at equal intervals, and sampling is performed along the sample lines in a systematic manner;

[0054] S1012: Using the ASD HH2 ground object hyperspectral instrument to acquire canopy hyperspectral reflectance data along the sampling route;

[0055] Referring to Figure 4 , the sampling process is specifically as follows:

[0056] 1) Equipment preheating: turning on the equipment 30 minutes before starting measurement to preheat the equipment, and starting acquisition after the equipment tends to be stable.

[0057] 2) Sampling preparation: The sampling time requirement is between 10 am and 2 pm to ensure sufficient solar elevation angle, and the specific scene also requires high visibility and less cloud cover to ensure light stability. In addition, the operator should wear dark clothes when collecting work to avoid the interference of scattered light, and ensure the stability of the operator's vision and equipment. Pay attention to operation specifications at all times during spectral measurement.

[0058] 3) Sample sampling: In order to ensure the accuracy of each sampling area, the device probe height should be calculated through geometric relationship according to the 25-degree field of view of the sensor, and the instrument should be kept vertical above the target shrub to ensure that there is no shadow in the sampling area.

[0059] 4) Output sampling data: Set the instrument parameters to collect one spectral data, and the instrument automatically records 10 canopy spectral curves and automatically calculates the average value as the final spectral data, making the data more scientific.

[0060] S102: Refer to Figure 5 , the collection process of high-spectral image data in the region using an unmanned aerial vehicle (UAV) high-spectral platform is as follows:

[0061] This embodiment uses DJI Matrice 600 Pro six-rotor unmanned aerial vehicle (flight platform weight is 10 kg, maximum load is 5 kg) to carry Resonon Pika L airborne push-broom hyperspectral system (spectral range is 400 ~ 1000 nm, spectral resolution is 2.1 nm) to obtain hyperspectral image data. Before starting the flight task, prepare the work: use GoogleEarth Pro software to draw the vector file and flight vector file of the target area, and before the plane takes off, the target area range file needs to be transmitted into the onboard microcomputer, so that the camera can correctly identify the region boundary to collect spectral image. After checking the flight parameters, use A3 professional flight controller and iPad to execute the flight task (flight height 65 m, flight speed 3 m / s, lateral overlap rate 30%) according to the flight route. In addition, this experiment synchronously uses DJI wizard 4 unmanned aerial vehicle to obtain RGB orthographic image, which provides reference data for the geometric correction and orthographic correction of unmanned aerial vehicle high-spectral data (Resonon data). Considering the influence of the angle of the sun, all flight tasks are carried out in sunny and cloudless conditions between 10 am and 2 pm.

[0062] S103: Refer to Figure 6 , the collection process of each shrub geographic coordinate is as follows:

[0063] The embodiment adopts the polar point RTK (Real-Time Kinematic, RTK) of the South Surveying and Mapping to collect the geographic coordinate information of the target object. The system is composed of a reference station, a mobile station and a hand book. First, according to the target area range boundary, 8 points in a preset distance (for example, 20 m in the embodiment) within each boundary are taken as ground control points, to ensure that the control points are all in the image collected by the unmanned aerial vehicle, so as to facilitate the subsequent geometric correction processing of the hyperspectral image of the unmanned aerial vehicle. At the same time, a target cloth is also needed to be laid near the target middle area, which serves as a control point for the geometric correction of the spectral image data, and also provides correction data for the next stage of unmanned aerial vehicle image processing.

[0064] When deploying the equipment, an open area far away from high-voltage lines should be selected to avoid the obstruction of objects during the measurement process, while ensuring that the satellite signal strength of the site meets the standard (at least receiving 11-13 satellite signals at the same time), and then using the hand book to set the parameters of the two deployed equipment (fixed solution, built-in radio mode, air baud rate 9600, and keeping the same channel between the equipment).

[0065] In step S2, referring to Figure 7 The ground canopy hyperspectral reflectance data of each species collected are subjected to data dimension reduction processing of first derivative transformation and continuum removal transformation, specifically as follows:

[0066] (1) The ground canopy hyperspectral reflectance data of each species collected are subjected to first derivative transformation, which effectively amplifies the subtle features in the spectral curve through mathematical operation, enhances the synergistic ability between the characteristic bands, and has remarkable effect in suppressing background interference and highlighting the spectral response of key parameters.

[0067] The first derivative spectrum significantly enhances the slope change of the original spectral reflectance curve, and reveals the fluctuation rate of the reflectance curve. In the first derivative spectral curve, the wave bands corresponding to the peak and valley values of the original spectral curve (i.e. the wave bands with zero first derivative) are clearly visible, and the peak and valley features in the derivative spectrum become more prominent. This spectral processing amplifies the spectral difference between the ground vegetation and bare soil in the near-infrared band by several times, significantly weakening the interference of soil background on vegetation spectral data. In addition, through the first derivative spectral data, the related parameters (three-edge parameters) in the blue edge (490-530 nm), yellow edge (560-640 nm) and red edge (680-760 nm) regions can be extracted, and then the spectral differences between Tetraclinis articulata and other species are analyzed. The calculation formula of the first derivative spectrum is as follows:

[0068]

[0069] In the formula: is the wavelength value of the wave band i; is the wavelength of the spectral value; Δλ is the wavelength to difference.

[0070] The three-edge feature is a relevant feature variable based on the first derivative spectral position, wherein the "three edges" respectively refer to the red edge, yellow edge and blue edge of the first derivative spectrum, and the three-edge parameter refers to the maximum value of reflectivity corresponding to the three-edge position, the wavelength position, the integral area in the wave band range and the normalized ratio of the integral area.

[0071] (2) The collected ground canopy hyperspectral reflectance data of each species is subjected to continuum removal transformation. Continuum removal, also known as envelope removal, is a data processing method for enhancing spectral features. The continuum refers to a connecting line segment formed by connecting the maximum points in the local region of the original spectral reflectance curve, and the core constraint condition is that the outer angle of the connecting line segment must be greater than 180° at each connecting point. Therefore, the continuum of the spectral curve is like a "shell" around the original reflectance curve, that is, the envelope. Continuum removal transformation is to divide the reflectance value corresponding to each wavelength on the original spectral curve by the value on the continuum line at the wavelength. The spectrum after the processing transformation is the continuum-removed spectrum, and the specific calculation formula is as follows:

[0072]

[0073] In the formula: is the reflectance value after continuum removal; is the original reflectance value; is the envelope function value.

[0074] In step S3, different transformed spectral curves are analyzed, and classification features are extracted therefrom to construct the Tetraena mongolica classification feature variable. The Tetraena mongolica classification feature variable includes a vegetation index variable, a continuum removal parameter variable and a hyperspectral feature parameter variable, and the specific contents are as follows:

[0075] S301: After the first derivative transformation of the original spectral data, 19 relevant feature parameters are constructed in this embodiment for subsequent classification tasks. The relevant description and calculation method are specifically described in Table 1.

[0076] Table 1. Hyperspectral feature parameters and calculation methods

[0077]

[0078] S302: The reflectance data is normalized to 0~1 after continuum removal transformation of the original spectral data, wherein significant absorption valleys are presented at 407 nm, 500 nm and 685 nm. In this embodiment, 19 characteristic parameters are extracted from the continuum-removed spectrum for constructing the classification model of Tetraena boissieri, and the related description and calculation method are shown in Table 2.

[0079] Table 2. Continuum removal parameters and calculation methods

[0080]

[0081] S303: A vegetation index variable is constructed, and the specific process is as follows:

[0082] The vegetation index refers to the result obtained by linear or nonlinear combination of spectral data of each waveband according to the spectral reflection characteristics of shrubs. The vegetation index can effectively highlight the spectral characteristics of each species, and compared with a single waveband, it can reveal the unique classification characteristics of the target shrub. In this embodiment, 13 vegetation indexes are used for the extraction of Tetraena boissieri by analyzing the spectral characteristics of each species, and the specific process is shown in Table 3.

[0083] Table 3. Calculation formula of spectral transformation vegetation index

[0084]

[0085] Step S4 is to analyze the importance of classification of Tetraena boissieri for the three sets of characteristic variables, i.e., the vegetation index variable (VI), the continuum removal parameter variable (CR) and the hyperspectral characteristic parameter variable (GTC), in view of the significant redundancy phenomenon existing in the three sets of characteristic variables, so as to evaluate their contribution, and to carry out sensitivity analysis on the regularization coefficients of each parameter to screen the key characteristic variables which have significant influence on the stability of the classification model, and the specific process is as follows:

[0086] In order to make the classification effect of Tetraena boissieri more significant, LASSO analysis method is used to screen out variable characteristics with significant classification performance from the vegetation index variable (VI variable), the continuum removal parameter variable (CR variable) and the hyperspectral characteristic parameter variable (GTC variable), and the specific screening results are shown in Table 4.

[0087] LASSO is the abbreviation of Least Absolute Shrinkage and Selection Operator, which is a regularization method for realizing variable selection and coefficient shrinkage in the regression framework. It can automatically identify and remove redundant variables from the three groups of high-dimensional characteristics VI, CR and GTC, and construct a stability-importance coupling index to output a key characteristic set which has significant contribution to the generalization error of the classification model and stable coefficients, and the specific process is shown in Table 4.

[0088] Table 4. LASSO screening results of different characteristic variables

[0089]

[0090] In step S5, the vegetation index threshold classification method and the machine learning-based classification method are respectively used to construct classification models for extracting Tetraena boea in the target area, and the classification effects of the classification models are evaluated by combining the preset model classification effect evaluation indicators, and the best classification model is selected according to the evaluation results, including,

[0091] S501: Construct a classification performance evaluation indicator;

[0092] The classification performance evaluation indicator used in this embodiment includes overall classification accuracy (OA), Kappa coefficient, precision (Precision), recall (Recall) and F1 score. Among them, the overall classification accuracy represents the proportion of samples that the prediction model predicts correctly, and its limitation lies in that it cannot reflect the influence of FP and FN on the classification task, so it is usually used to describe the performance of the classification model superficially; the Kappa coefficient is an evaluation indicator for evaluating whether the predicted samples of the classification model are consistent with the true class labels, and its range is between 0 and 1, the larger the K value, the higher the classification accuracy; the precision is the proportion of samples whose actual labels are positive class samples in the samples predicted by the model as positive class samples, and its focus is to avoid predicting negative class samples as positive class; the recall is a measure of the recognition ability of the classification model to positive class samples, which ensures that as many positive class samples in the data as possible are identified; the F1 score is a comprehensive balance indicator of precision and recall, which takes into account the misjudgment and omission of samples in the classification process. The specific calculation formula is as follows:

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099] In the formula, TP and TN are true positives and true negatives, respectively, indicating the number of correctly identified samples of each class by the model; FP and FN are false positives and false negatives, respectively, indicating the number of misjudged samples of the model, which misjudges negative class samples as positive class or misjudges positive class samples as negative examples. is referred to as overall classification accuracy OA; is the sum of the product of true samples and predicted samples.

[0100] S502: Based on the threshold classification method of different transformed vegetation index, according to the unique spectral information of Tetraena mongolica in different transformed spectrum, the band combination and step-by-step elimination strategy is used to finally realize the accurate extraction of Tetraena mongolica. Including,

[0101] S5021: Based on the unique spectral information of Tetraena mongolica in different transformed spectrum, the band combination and step-by-step elimination strategy is used to eliminate plants other than Tetraena mongolica;

[0102] According to the analysis of Figure 7 the spectral curves of Tetraena mongolica, Nitraria sibirica, Cistanche deserticola, Kandahar pistachio, Reaumuria soongorica and Suaeda salsa in different transformed spectrum, it is found that the spectral curve of Cistanche deserticola presents a trend similar to bare soil, so it can be eliminated by using vegetation index NDVI CR ; secondly, according to the analysis of first derivative spectrum, it is found that at 770 nm and 928 nm, the spectral characteristics of Suaeda salsa are significantly opposite to those of other species, so the vegetation index FDVI DV1 is used to amplify the difference between the two bands, so that Suaeda salsa is separated from other species; from the original spectral curve, it can be concluded that the red valley characteristics of Tetraena mongolica are more significant than those of other species, and at 389 nm and 425 nm of the first derivative spectrum, the first derivative reflectivity of Tetraena mongolica is much larger than that of other species. Under the comprehensive consideration, the vegetation index SGVI and LDVI ABS_DV1 are used to amplify the spectral characteristic information of Tetraena mongolica, so that other species are eliminated.

[0103] As shown in Table 5, NDVI CR has the best elimination effect on Cistanche deserticola, which can separate all samples of Cistanche deserticola from Tetraena mongolica samples, further verifying the correctness of the spectral analysis of Cistanche deserticola. Secondly, this vegetation index also has a certain elimination effect on Nitraria sibirica and Suaeda salsa. From the classification results of FDVI DV1 , it is found that this vegetation index has good elimination effect on Suaeda salsa, and can also eliminate most of Cistanche deserticola, but similar to NDVI CR , the elimination effects of the above two vegetation indexes on Nitraria sibirica, Reaumuria soongorica and Kandahar pistachio are not significant. In addition, from the classification results of SGVI, it can be directly seen that this vegetation index shows good classification performance on the remaining species except Kandahar pistachio, among which the classification effect on Cistanche deserticola, Nitraria sibirica and Suaeda salsa is the best, and the classification effect on Reaumuria soongorica is also good. Finally, from the classification results of LDVI ABS_DV1The classification results of the vegetation index showed that it had certain classification ability for two species of Euphorbia macropodoides and Reaumuria soongorica, but the elimination effect for other species was poor.

[0104] Table 5. Univariate classification results of spectral transformation vegetation index

[0105]

[0106] S5022: Determine the optimal classification threshold of each vegetation index by the maximum inter-class variance method to develop the classification rule.

[0107] Referring to Figure 9 , the present embodiment uses the vegetation index combination of NDVI CR , FDVI DV1 , SGVI and LDVI ABS_DV1 in turn, gradually eliminates other species samples, and determines the optimal classification threshold of each vegetation index by the maximum inter-class variance method to develop the classification rule. Table 6 is the classification results of the spectral transformation vegetation index combination. From the classification results, the overall classification accuracy of Tetraena boea is 0.9167, and the Kappa coefficient is 0.7818. The classification performance and generalization ability of the method for the classification model of Tetraena boea are high. At the same time, for Tetraena boea, the Recall index is 0.8611, and the F1 score is 0.8378, which also shows that the model has high precision for the accurate extraction of Tetraena boea.

[0108] Table 6. Combination classification confusion matrix of spectral transformation vegetation index

[0109]

[0110] S503: Construct a Tetraena boea classification model based on machine learning algorithm;

[0111] After screening by the LASSO analysis method, three groups of classification characteristic variables are obtained, which are CR variable, VI variable and GTC variable. Using the three classification characteristic parameters as input variables, machine learning algorithms such as SVM, XGBoost and KNN are used to construct a Tetraena boea classification model. Finally, the best classification model is determined by evaluating the classification performance of each model.

[0112] Table 7 is the model classification results of three machine learning algorithms in different classification variables, from which it can be seen that in different machine learning models, the classification model constructed by vegetation index has the highest precision. Among them, in the three classification models based on SVM, except that the Kappa coefficient of the test set of SVM-CR (SVM classification model based on CR variable) model is low, the overall classification accuracy and Kappa coefficient of the other two models are good; among the models based on XGBoost, XGBoost-CR model also produces overfitting phenomenon, the Kappa coefficient of the test set is low, and the generalization ability of the model is weak; among the several classification models based on KNN, the overall classification accuracy is at a medium level, but the Kappa coefficient is low. Finally, by comprehensively comparing the classification accuracy of models of different algorithms, it can be concluded that the classification performance of SVM-VI, XGBoost-VI and KNN-VI models is good.

[0113] Table 7. Model classification results of machine learning algorithms in different classification variables

[0114]

[0115] Note: T represents the training set; S represents the test set; GTC represents the hyperspectral feature variable.

[0116] Table 8 is the classification results of the above three models on the test set, it can be found that SVM-VI performs well on the test set, only one poplar is misclassified as other species, so the recall rate of poplar is as high as 0.9167, and the F1 score is 0.88. In addition, the overall classification accuracy of this model reaches 0.9388, and the Kappa coefficient is 0.839, so the above data shows that the model performance of SVM-VI is superior, and has high prediction result reliability; the model accuracy of XGBoost-VI is at a medium level among the three models, compared with SVM-VI, the recall rate of poplar is only 0.7, and the F1 score is 0.7778. At the same time, from the Kappa coefficient of the model, it can be seen that the XGBoost-VI model has certain poplar classification performance; the model performance of KNN-VI is the worst, the misclassification rate of poplar is as high as 0.5, the recall rate is 0.5, the F1 score is 0.6316, and the Kappa coefficient of the model is only 0.5505, so the poplar extraction effect of this model is the worst.

[0117] Table 8. Confusion matrix of the best model of different algorithms on the test set

[0118]

[0119] S504: Compare the effects of different four-wood classification models, and select the best classification model to identify and extract four-wood;

[0120] Referring to Figure 10 The performance of the classification models constructed by the three machine learning algorithms was evaluated, and it was found that the model classification evaluation indicators of SVM-VI were higher than those of the other two classification models. Therefore, in the four-wood classification method based on machine learning, SVM-VI is the optimal classification model. By comparing the model accuracy of SVM-VI and the threshold classification method based on vegetation index, it can be seen that the Kappa coefficient and Recall of SVM-VI are higher, and the accuracy of four-wood extraction is more accurate. In summary, by comparing the classification effects of four-wood based on different classification methods of hyperspectral data, it is found that SVM-VI is the best classification model for four-wood.

[0121] In step S6, based on the best classification model of ground canopy spectral data, the four-wood in the target area is identified and classified by combining the unmanned aerial vehicle hyperspectral image data, and the spatial distribution data of four-wood is obtained, including,

[0122] The geographic coordinates of each shrub collected by RTK are used to obtain the unmanned aerial vehicle hyperspectral data by point extraction method for collaborative verification of air-ground spectral data. In addition, by using the unmanned aerial vehicle hyperspectral image data, combined with the best classification model of ground four-wood, the planar classification task of four-wood shrubs is realized, and the extraction effect of four-wood in the unmanned aerial vehicle classification image is evaluated. Finally, an automatic extraction method of four-wood based on hyperspectral data is constructed, which is as follows:

[0123] S601: Use sample point coordinates to extract unmanned aerial vehicle hyperspectral data, analyze sample spectral characteristics, and obtain four-wood classification effect based on the best ground classification model;

[0124] In this embodiment, the unmanned aerial vehicle hyperspectral data is preprocessed by methods such as radiometric calibration, geometric correction, image mosaicking and image filtering, and the geographic coordinates of each shrub sample point are obtained by the point extraction method in ArcGIS software. By removing outliers from the obtained hyperspectral data, 325 valid sample data are finally retained. By analyzing the different transformed spectral characteristics of the unmanned aerial vehicle hyperspectral data, it is found that the classification characteristics of four-wood are similar to those of ground objects. Therefore, the best classification variable (VI variable) of ground four-wood is used as the model input variable of unmanned aerial vehicle data. Table 9 is the result of LASSO screening of VI variables calculated based on unmanned aerial vehicle hyperspectral data. By stratified sampling method, 2 / 3 of the data are used as training set, and 1 / 3 of the data are used as test set, and four-wood classification model is constructed by SVM-VI.

[0125] Table 9. LASSO screening results of vegetation index variables

[0126]

[0127] Table 10 is the classification results of the SVM-VI model based on the UAV hyperspectral data. From the results, it is found that the overall classification accuracy of Tetraena mongolica using the SVM-VI model reaches 0.9115, and the Kappa coefficient is 0.814, which indicates that the model has high classification accuracy of Tetraena mongolica on hyperspectral image data, and has good classification reliability and model generalization ability. Secondly, the recall rate and F1 score of Tetraena mongolica both reach more than 0.85, which further illustrates the reliability of the model in extracting Tetraena mongolica from hyperspectral image data, and can effectively reduce the occurrence of "missed" and "misclassified" cases.

[0128] Table 10. SVM-VI model test set classification confusion matrix

[0129]

[0130] S602: Through the classification method of ground canopy spectral data, combined with UAV hyperspectral image data for classification, so as to quickly and accurately extract Tetraena mongolica automatically, specifically as follows:

[0131] Referring to Figure 11 , the present embodiment extracts Tetraena mongolica distributed in the region automatically through UAV hyperspectral image. First, the ENVI software is used to perform continuous uniform removal transformation processing on the hyperspectral image data. Second, the raster calculator of the ArcGIS software is used to calculate the vegetation index of different transformed spectra respectively, and the ENVI software is used to merge and process all the vegetation index raster data. Finally, the merged data is used as the input variable, and the SVM model is used to extract Tetraena mongolica distributed in the region. According to the classification results of the UAV image, most of the Tetraena mongolica in the sample area is accurately identified. At the same time, from the morphological feature analysis of the extraction results, it is found that the Tetraena mongolica patches present an approximate circular or elliptical structure, which is consistent with the actual growth morphological characteristics of Tetraena mongolica shrubs.

[0132] Table 11 is the superimposed analysis of the test set sample data, and the classification confusion matrix obtained after calculation. From the results, it is found that the extraction results of the UAV image are consistent with the actual distribution position of Tetraena mongolica, among which only 2 samples are not correctly extracted in the extraction results of 43 Tetraena mongolica verification samples, the recall rate is as high as 0.95, and the F1 score is 0.8723, which also verifies the accuracy and reliability of the method in the identification and extraction of Tetraena mongolica.

[0133] Table 11. Classification confusion matrix of test set sample points in UAV image

[0134]

[0135] In summary, by analyzing the different transformed spectral characteristics of the six shrubs, including Tetraena mongolica, Nitraria sibirica, Nitraria tangutorum, Nitraria sphaerocarpa, Reaumuria soongorica and Suaeda salsa, the following results were obtained: in the original reflectance spectrum, the sensitive bands were 645 nm, 682 nm, 750 nm, 765 nm and 941 nm; in the first derivative spectrum and the logarithmic derivative spectrum, the sensitive bands were 389 nm, 425 nm, 702 nm, 770 nm and 928 nm; in the continuum removal spectrum, the sensitive bands were 460 nm, 560 nm, 682 nm and 750 nm. Subsequently, three classification characteristic variables of Tetraena mongolica were constructed according to the above characteristics, which were CR variable, VI variable and GTC variable, and then the LASSO analysis method was used to screen the classification variables.

[0136] Two classification methods of Tetraena mongolica were used in this embodiment, which were the vegetation index threshold classification method based on different spectral transformations and the classification method based on machine learning. The former was to extract the sensitive bands according to the characteristic differences of different transformed spectra of Tetraena mongolica and other species, and then highlight the characteristic information of Tetraena mongolica by band combination, so as to use vegetation index combination to remove other species in turn. Specifically, the results of each vegetation index single variable classification were analyzed, and the vegetation index combinations of NDVI CR , FDVI DV1 , SGVI and LDVI ABS_DV1 were used in turn to remove other species samples, and the best classification threshold of each vegetation index was determined by the maximum inter-class variance method, and finally the classification rule was formulated. It was found from the classification results that the overall classification accuracy of Tetraena mongolica was 0.9167, the Kappa coefficient was 0.7818, the Recall was 0.8611 and the F1 score was 0.8378. The latter was the classification results of Tetraena mongolica obtained by taking CR variable, VI variable and GTC variable as machine learning input variables, which were as follows: among them, the classification models of SVM, XGBoost and KNN based on VI variable had the best performance, and after comparing the above three models, it was finally found that SVM-VI had the best extraction effect on Tetraena mongolica, with an overall classification accuracy of 0.9388, a Kappa coefficient of 0.839, a Recall of 0.9167 and an F1 of 0.88. By comparing the model accuracy of the two classification methods of Tetraena mongolica, the following conclusion was drawn: the best classification model of Tetraena mongolica was SVM-VI.

[0137] In addition, the unmanned aerial vehicle hyperspectral data is obtained by point extraction of RTK data of each shrub. Then, the spectral characteristics of reflectance spectrum and continuum-removed spectrum are analyzed respectively, and the results show that the reflectance spectrum characteristics of the unmanned aerial vehicle data are basically consistent with the spectral characteristics of the ground canopy reflectance curve, and there are some differences at the positions of blue edge and green peak. The analysis of the continuum-removed spectrum shows that each species has multiple obvious absorption valleys in the visible light range, and the absorption depth of Tetraena mongolica at 515 nm and 682 nm is significantly greater than that of other species. Therefore, the spectral analysis results show that the classification characteristics of Tetraena mongolica of the unmanned aerial vehicle hyperspectral image data are relatively the same as those of the ground object hyperspectral classification, so the best classification VI variable of Tetraena mongolica on the ground is used as the model input variable of the unmanned aerial vehicle data. It is found from the model classification results of SVM-VI that the overall classification accuracy of Tetraena mongolica reaches 0.9115, the Kappa coefficient is 0.814, and the Recall and F1 scores are all above 0.85. Secondly, through the analysis of the validation set sample points, it is found that the spatial distribution position of Tetraena mongolica extracted from the unmanned aerial vehicle image is relatively consistent with the actual position, the overall classification accuracy is 0.8938, the Kappa coefficient is 0.7826, the recall rate is as high as 0.95, and the F1 score is 0.8723, which further verifies that the method has high precision and reliability in the identification and extraction of Tetraena mongolica.

[0138] In the embodiment, the best classification model of Tetraena mongolica is explored through spectral dimension reduction, classification feature extraction, vegetation index threshold classification and machine learning algorithm. The classification of ground canopy spectral data is combined with the classification of unmanned aerial vehicle hyperspectral image data, so as to realize the rapid and accurate automatic extraction of Tetraena mongolica. The method provides key data support such as spatial distribution of Tetraena mongolica for feasibility evaluation of mining area and other artificial construction projects, and provides technical support for scientific protection decision of Tetraena mongolica. The distribution information of rare and protected plant Tetraena mongolica in the target area can be quickly obtained by using the technical method, and the accuracy of automatic extraction of Tetraena mongolica is ensured.

[0139] Embodiment 2

[0140] Referring to Figure 12 , the embodiment of the present application provides a kind of Tetraena mongolica automatic extraction system based on hyperspectral data, comprising,

[0141] Data acquisition platform is used to collect the ground canopy hyperspectral reflectance of each species and unmanned aerial vehicle hyperspectral image data in target area, and the geographic coordinates of each shrub are collected synchronously;

[0142] Hyperspectral data processing module is used to carry out first derivative transformation and continuum removal transformation processing to the ground canopy hyperspectral reflectance data;

[0143] The characteristic variable construction module is configured to analyze different transformed spectral curves and construct four-wood classification characteristic variables, including vegetation index variables, continuum removal parameter variables and hyperspectral characteristic parameter variables.

[0144] The characteristic variable screening module is configured to perform importance analysis and sensitivity analysis on the four-wood classification characteristic variables, and screen key characteristic variables.

[0145] The classification model construction module is configured to construct classification models for extracting four-wood in a target region by using a vegetation index threshold classification method and a machine learning-based classification method respectively, evaluate classification effects of the classification models in combination with preset evaluation indexes, and select an optimal classification model according to an evaluation result.

[0146] The regional four-wood identification module is configured to identify and classify four-wood in a target region based on the optimal classification model of ground canopy spectral data in combination with unmanned aerial vehicle hyperspectral image data, and obtain spatial distribution data of the four-wood.

[0147] The above description of the embodiments is only used to help understand the method of the present application and its core idea. It should be noted that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

[0148] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0149] It should be noted that some of the example embodiments are described as processes that are depicted as flow diagrams or flow charts. Although each can describe the operations as a sequential process, many of the operations can be performed in parallel, concurrently or simultaneously. In addition, the order of the operations can be re-arranged. A process can be terminated when its operations are completed, but could also be terminated or interrupted before, without completing, due to various failure scenarios or due to a user intervention. The processes shall be understood to be not limited by the order of that is described, unless such order is specifically shown. The processes can correspond to methods, functions, procedures, subroutines, subprograms, etc.

Claims

1. An automatic extraction method for *Tetraphyta tenuifolia* based on hyperspectral data, characterized in that, include, A data acquisition platform was built to collect hyperspectral reflectance of the ground canopy of each species and hyperspectral image data of UAVs in the target area, while simultaneously collecting the geographic coordinates of each shrub. The ground canopy hyperspectral reflectance data were subjected to first-order derivative transformation and continuum removal transformation. Different transformation spectral curves were analyzed and classification characteristic variables of *Tetraphyta glabra* were constructed. The classification characteristic variables of *Tetraphyta glabra* include vegetation index variables, continuum removal parameter variables, and hyperspectral characteristic parameter variables. Importance and sensitivity analyses were performed on the classification characteristic variables of the tetrahedron to screen key characteristic variables; Classification models were constructed using vegetation index threshold classification and machine learning-based classification methods to extract *Tetraphyta tenuifolia* from the target area. The classification performance of the models was evaluated using preset evaluation indicators, and the best classification model was selected based on the evaluation results. Based on the optimal classification model of ground canopy spectral data, combined with UAV hyperspectral image data, the tetrapods in the target area are identified and classified to obtain the spatial distribution data of tetrapods.

2. The automatic extraction method for tetrahedron based on hyperspectral data according to claim 1, characterized in that, The data acquisition platform includes using a ground-based hyperspectral instrument to collect hyperspectral reflectance data of the ground canopy; using a UAV hyperspectral platform to collect hyperspectral image data; and using the extreme point RTK of Southern Surveying and Mapping to collect the geographic coordinates of each shrub.

3. The automatic extraction method for tetrahedron based on hyperspectral data according to claim 1, characterized in that, The above process involves performing a first-order derivative transform and a continuum removal transform on the hyperspectral reflectance data of the ground canopy. The formula for the first-order derivative transform is as follows: In the formula: It is the wavelength value of band i; Wavelength The spectral value; Δλ is the wavelength. arrive Difference; The term "continuum removal parameter variable" refers to the analysis of the continuum removal spectrum to extract relevant spectral features. The continuum removal transformation formula is as follows: In the formula: The reflectance value after removing the continuum; This is the original reflectivity value; This represents the envelope function value.

4. The automatic extraction method for tetrahedron based on hyperspectral data according to claim 1, characterized in that, The vegetation index variable refers to the result obtained by linear and nonlinear combination calculations of each band of spectral data.

5. The automatic extraction method for tetrahedron based on hyperspectral data according to claim 1, characterized in that, The importance and sensitivity analyses of the classification characteristic variables of *Tetraphyta glabra* were performed to screen key characteristic variables. The LASSO analysis method was used to select variable characteristics with high classification performance from vegetation index variables, continuum parameter removal variables, and hyperspectral characteristic parameter variables.

6. The automatic extraction method for tetrahedron based on hyperspectral data according to claim 1, characterized in that, The step of using a vegetation index threshold classification method to construct a classification model for extracting *Tetraena mongolica* within the target area includes: Based on the unique spectral information of Tetracentron sinense in different transformed spectra, a strategy of band combination and stepwise elimination was used to remove plants other than Tetracentron sinense. The optimal classification threshold for each vegetation index was determined by the Otsu's method, and classification rules were formulated.

7. The automatic extraction method for tetrahedron based on hyperspectral data according to claim 1, characterized in that, The machine learning-based classification method constructs a classification model to extract *Tetraphyta tenuifolia* within the target area, including: Based on three sets of classification feature parameters—the vegetation index variable after feature screening, the continuum removed parameter variable, and the hyperspectral feature parameter variable—as input variables, a classification model for *Tetraphyta pulcherrima* is constructed using machine learning algorithms. These machine learning algorithms include Support Vector Machine, eXtreme Gradient Boosting, and K-Nearest Neighbors.

8. The automatic extraction method for tetrahedron based on hyperspectral data according to claim 1, characterized in that, The pre-defined evaluation metrics include overall classification accuracy (OA), Kappa coefficient, precision, recall, and F1 score, with the specific calculation formulas as follows: In the formula: TP and TN are true positives and true negatives, respectively, representing the number of samples correctly identified by the model; FP and FN are false positives and false negatives, respectively, representing the number of samples that the model misclassifies as positive or negative. This refers to the overall classification accuracy (OA). It is the sum of the products of the real samples and the predicted samples.

9. The automatic extraction method for tetrahedron based on hyperspectral data according to any one of claims 1-8, characterized in that, The optimal classification model based on ground canopy spectral data, combined with UAV hyperspectral imagery data, identifies and classifies *Tetraena mongolica* trees in the target area, obtaining spatial distribution data of *Tetraena mongolica* trees, including... The hyperspectral image data of the UAV is subjected to spectral transformation, vegetation indices of different transformed spectra are calculated, and all vegetation index raster data are merged. The merged raster data is used as input variables to construct a regional tetrapod classification model. The automatic extraction effect of tetrapods from UAV hyperspectral image data is evaluated by validation set sample points, and a tetrapod spatial distribution map is output.

10. The automatic extraction system for tetrahedron based on hyperspectral data according to claim 1, characterized in that, include, The data acquisition platform is used to collect the hyperspectral reflectance of the ground canopy of each species and the hyperspectral image data of UAVs in the target area, and simultaneously collect the geographic coordinates of each shrub. The hyperspectral data processing module is used to perform first-order derivative transformation and continuum removal transformation on the hyperspectral reflectance data of the ground canopy. The feature variable construction module is used to analyze different transform spectral curves and construct tetragonal tree classification feature variables. The tetragonal tree classification feature variables include vegetation index variables, continuum removal parameter variables, and hyperspectral feature parameter variables. The feature variable screening module is used to perform importance and sensitivity analysis on the classification feature variables of the tetrapod and screen key feature variables. The classification model building module is used to build classification models using vegetation index threshold classification method and machine learning-based classification method to extract tetrahedron in the target area, and to evaluate the classification effect of the classification model in combination with pre-set evaluation indicators, and select the best classification model based on the evaluation results. The regional tetrapanax tree identification module is used to identify and classify tetrapanax trees in the target area based on the best classification model of ground canopy spectral data and combined with UAV hyperspectral image data to obtain the spatial distribution data of tetrapanax trees.