Soil organic matter content monitoring model training method and device based on hyperspectral data

By using hyperspectral data and topographic factor data in soil organic matter content monitoring, the relevant spectral index and parameters are calculated, the characteristic data set is constructed and the monitoring model is trained, and the problem of high-precision monitoring of soil mass distribution in the existing technology is solved, and efficient and accurate monitoring of soil organic matter content is achieved.

CN119167096BActive Publication Date: 2025-05-23MINISTRY OF NATURAL RESOURCES LAND SATELLITE REMOTE SENSING APPL CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411669520.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-05-23
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high accuracy and rapid acquisition of soil mass distribution conditions, especially in complex terrain and landform areas, and traditional feature selection methods have problems of difficulty in balancing accuracy and efficiency in hyperspectral inversion.

Method used

By obtaining hyperspectral surface reflectivity data, topographic factor data and soil organic matter content, calculate the correlation between the two-band spectral index and spectral parameters and soil organic matter content, build a feature data set, and use the feature data set to train the soil organic matter content monitoring model to optimize the feature selection process to improve the model accuracy.

Benefits of technology

It realizes higher accuracy of soil organic matter content monitoring, improves monitoring efficiency, ensures the clarity and specificity of data acquisition, and improves the frequency and accuracy of soil organic matter content monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119167096B_ABST
    Figure CN119167096B_ABST
Patent Text Reader

Abstract

The present disclosure provides a soil organic matter content monitoring model training method and device based on hyperspectral data, which is applied to the field of soil organic matter content monitoring technology. The method includes calculating the correlation between the dual-band spectral index and spectral parameters obtained according to the hyperspectral surface reflectance data and the soil organic matter content; constructing a feature data set according to the spectral parameters and spectral index corresponding to the correlation whose absolute value is greater than a preset threshold, and terrain factor data; training the soil organic matter content monitoring model according to the feature data set, calculating the model accuracy and the accuracy improvement value; using the band and terrain factor type corresponding to the feature with an accuracy improvement value greater than 0 as the data acquisition condition when monitoring the soil organic matter content in the monitored area, and using the model trained with the feature data set corresponding to the feature with an accuracy improvement value greater than 0 as the final model. In this way, data acquisition conditions with high prediction accuracy and a soil organic matter content monitoring model can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of soil organic matter content monitoring, and in particular to a soil organic matter content monitoring model training method and device based on hyperspectral data. Background Art

[0002] Soil quality survey is an important business in the field of natural resource survey. Soil organic matter is an important indicator for evaluating soil quality. It is a decisive factor affecting soil fertility and crop yields, and is of great significance to the cycle of soil nutrients and sustainable agricultural development. The factors affecting soil quality are complex and closely related to factors such as topography, climate, parent material, vegetation and human activities. Traditional soil quality surveys are mainly based on field soil sampling and testing. This method can directly obtain reliable point data, but it is restricted by factors such as long field sampling cycles, large sample time spans, and high work implementation costs. It cannot support large-scale, high-frequency macro-dynamic monitoring. The traditional soil quality spatial mapping method represented by geostatistics has become the main means of soil quality mapping in the past because of its simplicity and significant interpolation effect. However, the geostatistical method does not take into account the relationship between soil quality and terrain factors, and it is difficult to achieve high-precision mapping of soil organic matter content in complex terrain and landforms.

[0003] The rapid development of satellite remote sensing technology has provided a stable data source for obtaining surface parameters. Compared with traditional methods, satellite remote sensing data has the advantages of being fast, economical, environmentally friendly, non-destructive, and repeatable, providing a new means for large-scale, high-precision, and high-frequency soil quality surveys and monitoring. As a spectrum fusion imaging technology, hyperspectral remote sensing obtains the geometric characteristics of the target by quickly acquiring continuous subdivided spectral information, and can quantitatively invert the spectral reflectance, radiation, and absorption characteristics of the target. In recent years, with the successive launch of the Gaofen-5, Zhuhai-1, and Ziyuan-1 02D satellites, the ability to obtain ground spectra with multi-space spectral resolution and monthly revisits in key areas has been formed, providing effective data guarantee for soil quality monitoring. Among them, Ziyuan-1 02D and Ziyuan-1 02E satellites can achieve a 2-day revisit observation of the ground at the fastest under networking conditions, greatly improving the observation efficiency of cultivated soil.

[0004] At present, the commonly used method for soil quality monitoring and evaluation is the spectral index method, which mainly uses the reflectance of two or more wavelengths for combined operations to highlight a certain characteristic or detailed information of the soil. Researchers have proposed soil spectral indices with different combinations to obtain soil quality distribution. For example, the soil organic matter content is estimated using the band reflectance ratio index after spectral transformation; according to the spectral absorption characteristics of the soil organic matter content, spectral indices such as difference index and normalized difference index are constructed to analyze the correlation between the index and the soil organic matter content. At the same time, researchers have also demonstrated the potential for quantitative inversion of soil organic matter content based on hyperspectral data. Most of them use models such as multivariate stepwise regression, partial least squares regression and BP neural network for inversion. For example, the soil organic matter content is estimated using sensitive spectral reflectance bands; based on soil spectral reflectance, a soil organic matter content classification model is established in combination with partial least squares regression method. However, these studies are all based on laboratory soil spectral reflectance data. Due to the differences between laboratory soil samples and field soil samples and the influence of observation scale, existing studies are difficult to directly apply to satellite remote sensing data. Therefore, it is necessary to study the soil organic matter content inversion model suitable for multi-source satellite collaborative observation to achieve high-precision and rapid acquisition of soil quality distribution. However, when studying the soil organic matter content inversion model suitable for multi-source satellite collaborative observation, it is particularly important to select features to determine the feature data set used for training the model. Common feature selection is the process of selecting a subset from a feature set and selecting the optimal subset using evaluation criteria. Subset generation is mainly completed through heuristic search, including sequential search, exhaustive search, and random search. The evaluation criteria have developed different algorithms based on actual needs and data characteristics. In hyperspectral soil inversion, algorithms such as variable importance projection, Pearson correlation coefficient, competitive adaptive weighted sampling, genetic algorithm and simulated annealing are commonly used feature selection methods, but there are considerable problems in the application of these more common feature selection algorithms in hyperspectral inversion: ① As more common feature selection techniques in various research fields, the above methods are not optimized for the characteristics of hyperspectral data. Some unsupervised algorithms place too much emphasis on the statistical analysis of the data itself, and the extracted features are usually difficult to guarantee the accuracy of inversion modeling; ② Currently commonly used feature selection methods usually have several random subset generation or evaluation processes, and there are certain problems with the stability of the methods. Under the same circumstances, there may be large differences in results, which interferes with the subsequent inversion modeling process; ③ Better feature selection results require more cumbersome calculation processes and consume a lot of computing power. Traditional methods usually find it difficult to strike a balance between accuracy and efficiency. Summary of the invention

[0005] The present invention provides a soil organic matter content monitoring model training method and device based on hyperspectral data.

[0006] According to a first aspect of the present disclosure, a soil organic matter content monitoring model training method based on hyperspectral data is provided. The method comprises:

[0007] Obtain the hyperspectral surface reflectance data of the sample points, the terrain factor data of the sample points, and the soil organic matter content of the sample points;

[0008] Calculating a plurality of dual-band spectral indices and a plurality of spectral parameters according to the hyperspectral surface reflectance data; and calculating the correlation between the soil organic matter content and the spectral parameters, and the soil organic matter content and the dual-band spectral index;

[0009] Constructing a characteristic data set according to the spectral parameters and spectral indices corresponding to the correlations whose absolute values ​​are greater than a preset threshold, and the terrain factor data;

[0010] Training a soil organic matter content monitoring model according to the characteristic data set, calculating the model accuracy, and calculating the accuracy improvement value according to the accuracy;

[0011] The bands and terrain factor types corresponding to the features with precision improvement values ​​greater than 0 are used as data acquisition conditions when monitoring the soil organic matter content in the monitored area; the soil organic matter content monitoring model used when monitoring the soil organic matter content in the monitored area is a model trained with the feature data set corresponding to the features with precision improvement values ​​greater than 0.

[0012] According to the above aspects and any possible implementation, an implementation is further provided, wherein the calculating of the spectral index and the spectral parameter according to the hyperspectral surface reflectance data comprises:

[0013] Smoothing the hyperspectral surface reflectance data;

[0014] Calculating spectral parameters according to the smoothed hyperspectral surface reflectance data; and performing spectral transformation on the smoothed hyperspectral surface reflectance data to obtain multiple bands;

[0015] Multiple dual-band spectral indices are calculated based on the hyperspectral surface reflectance data corresponding to each band.

[0016] According to the above aspects and any possible implementation, an implementation is further provided.

[0017] The calculation methods of spectral parameters include: averaging and slope;

[0018] The formula for calculating the spectral index based on the hyperspectral surface reflectance data corresponding to each band is: DI=p - q, RI = p / q,NDI=(pq) / ( p + q),DSI= ,

[0019] in, p , q is the hyperspectral surface reflectance data of any two bands, and p–q≠0 .

[0020] According to the above aspects and any possible implementation, an implementation is further provided.

[0021] The step of training the soil organic matter content monitoring model according to the characteristic data set and calculating the model accuracy includes:

[0022] Performing feature optimization calculation on the feature data set to obtain feature importance;

[0023] The features are sorted from high to low according to the feature importance, and the features are input into the soil organic matter content monitoring model in sequence according to the sorting order for training, and the accuracy of the model is calculated.

[0024] According to the above aspects and any possible implementation, an implementation is further provided.

[0025] The feature optimization calculation methods include: joint random frog RF, competitive self-organizing selection CARS, and variable importance factor VIP.

[0026] According to the above aspects and any possible implementation, an implementation is further provided.

[0027] The calculation methods for the accuracy of the soil organic matter content monitoring model include adjusting the coefficient of determination, root mean square error, and relative analytical error;

[0028] The accuracy improvement value S of the soil organic matter content monitoring model i The calculation formula is:

[0029] ,

[0030] in, , It indicates the maximum and minimum values ​​of the adjusted determination coefficient obtained when using the feature for inversion. represents the adjusted determination coefficient obtained when inversion is performed using the i-th feature, where i represents the order of the feature; , It represents the maximum and minimum values ​​of the root mean square error obtained when using the feature for inversion. represents the root mean square error obtained when inverting using the i-th feature; , It indicates the maximum and minimum values ​​of the relative analysis error obtained when using the feature for inversion. Represents the relative analysis error obtained when inverting using the i-th feature.

[0031] According to a second aspect of the present disclosure, a method for monitoring soil organic matter content based on hyperspectral data is provided. The method comprises:

[0032] Acquire high-spectral surface reflectance data of a preset band and terrain factor data of a preset type in the area to be monitored; the band and the type are respectively the band and terrain factor type in the data acquisition conditions obtained by the method of the first aspect above;

[0033] Calculating spectral index and spectral parameters according to the hyperspectral surface reflectance data;

[0034] The spectral index, spectral parameters and terrain factor data are input into a pre-trained soil organic matter content monitoring model, and the soil organic matter content of the monitored area is output; the pre-trained soil organic matter content monitoring model is a model trained with a feature data set corresponding to features with accuracy improvement values ​​greater than 0.

[0035] According to a third aspect of the present disclosure, a soil organic matter content monitoring model training device based on hyperspectral data is provided. The device comprises:

[0036] A data acquisition module is used to obtain the high-spectral surface reflectance data of the sample points, the terrain factor data of the sample points, and the soil organic matter content of the sample points;

[0037] A correlation calculation module, used to calculate a plurality of dual-band spectral indices and a plurality of spectral parameters according to the hyperspectral surface reflectance data; and calculate the correlation between the soil organic matter content and the spectral parameters, and the soil organic matter content and the dual-band spectral index;

[0038] A data set construction module, used to construct a feature data set according to the spectral parameters and spectral indices corresponding to the correlations whose absolute values ​​are greater than a preset threshold, and the terrain factor data;

[0039] An accuracy calculation module, used to train the soil organic matter content monitoring model according to the characteristic data set, calculate the model accuracy, and calculate the accuracy improvement value according to the accuracy;

[0040] The feature selection module is used to use the bands and terrain factor types corresponding to the features with precision improvement values ​​greater than 0 as data acquisition conditions when monitoring the soil organic matter content in the monitored area; the soil organic matter content monitoring model used when monitoring the soil organic matter content in the monitored area is a model trained with the feature data set corresponding to the features with precision improvement values ​​greater than 0.

[0041] According to a fourth aspect of the present disclosure, an electronic device is provided, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, the method described above is implemented.

[0042] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0043] The embodiment of the present disclosure provides a soil organic matter content monitoring model training method based on hyperspectral data, which calculates the correlation between the dual-band spectral index and spectral parameters and the soil organic matter content, and then constructs a feature data set with the spectral parameters and spectral index with high correlation, as well as terrain factor data, trains the soil organic matter content monitoring model according to the feature data set, calculates the model accuracy, and the accuracy improvement value; the model trained with the feature data set corresponding to the feature with an accuracy improvement value greater than 0 is used as the final model for soil organic matter content monitoring. In this way, higher-precision soil organic matter content monitoring can be achieved, and the band and terrain factor type corresponding to the feature with an accuracy improvement value greater than 0 are used as data acquisition conditions when monitoring the soil organic matter content in the monitored area, so that data acquisition during soil organic matter content monitoring is clearer and more specific, and the monitoring efficiency is better guaranteed while achieving higher monitoring accuracy.

[0044] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:

[0046] Figure 1 A flowchart of a soil organic matter content monitoring model training method based on hyperspectral data according to an embodiment of the present disclosure is shown;

[0047] Figure 2 shows a distribution diagram of sampling point locations within a study area according to an embodiment of the present disclosure;

[0048] Figure 3 A schematic diagram showing the correlation between the measured value and the predicted value of soil organic matter content according to an embodiment of the present disclosure is shown;

[0049] Figure 4 A result diagram of regional soil organic matter content distribution according to an embodiment of the present disclosure is shown;

[0050] Figure 5 A block diagram of a soil organic matter content monitoring model training device based on hyperspectral data according to an embodiment of the present disclosure is shown;

[0051] Figure 6 A schematic block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0053] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0054] In the present disclosure, firstly, the surface reflectance data is obtained by preprocessing the satellite-borne hyperspectral data, then the elevation data is calculated to obtain the basic terrain factors, and the spectral reflectance data and terrain factor data of the pixels are extracted based on the sampling points; on the basis of smoothing the reflectance data, spectral transformation processing and spectral parameter calculation are performed, and a dual-band index is constructed based on the spectral transformation reflectance data, and a feature data set is obtained by combining them; features are screened by performing image quality analysis and evaluation on the feature data set, and features are selected by using a multi-algorithm joint feature selection technology, and the optimal number of features is obtained by gradually increasing the accuracy of the feature comparison model, and the feature variables and number of features finally used are determined; finally, a soil organic matter content inversion model is obtained by comprehensive comparison of multiple machine learning models; and the soil organic matter content inversion model is trained using the optimal feature data set to obtain a soil organic matter content monitoring model.

[0055] Figure 1 The flowchart of the soil organic matter content monitoring model training method 100 based on hyperspectral data according to an embodiment of the present disclosure is shown. The method 100 includes:

[0056] Step 110, obtaining the hyperspectral surface reflectance data of the sample point, the terrain factor data of the sample point and the soil organic matter content of the sample point.

[0057] In some embodiments, synchronous ground sample collection is carried out according to the satellite transit time to obtain soil samples, and then the organic matter content of the soil samples is analyzed and measured to obtain the soil organic matter content of the sample points. For example, synchronous ground sample collection was carried out in a certain area, and the ground sampling time was April 2021 and April 2019. Synchronous observation data of 273 points were obtained, covering 10 typical soil types in black soil areas, such as black soil, chernozem, brown soil, and dark brown soil. Figure 2 As shown. It should be noted that the layout of sampling points needs to consider the principles of comprehensiveness, representativeness, objectivity, feasibility, and continuity, and select locations with obvious soil type characteristics, flat terrain, and good vegetation growth, away from towns, houses, roads, and ditches. For example, the "five-point sampling method" is used to collect soil samples at a depth of 0 to 20 cm at each sampling point, and the mixed soil sample of the five samples is used as the representative sample of the sampling point. The soil organic matter content of each sampling point is determined using the "potassium dichromate volumetric method".

[0058] In some embodiments, the satellite-borne hyperspectral data is preprocessed to obtain hyperspectral surface reflectance data. The hyperspectral surface reflectance data of the pixel where the sample point is located is extracted according to the sampling point position. That is, the hyperspectral surface reflectance data is obtained by preprocessing radiation correction, atmospheric correction, image quality improvement, orthorectification, and cultivated land area mask according to the product level; for example, according to the scope of the study area, GF5 and ZY1-02D ​​satellite L1A-level hyperspectral images are collected, and ENVI software is used for radiation calibration and FLAASH atmospheric correction. The image quality is improved by using the moment matching method, and then ENVI software is used for orthorectification. The cultivated land distribution vector data of the study area is used to realize the cultivated land area mask to obtain the cultivated land surface reflectance data. The hyperspectral surface reflectance data of the pixel where the sample point is located is extracted based on the sampling point point vector data.

[0059] In some embodiments, digital elevation model DEM data is collected, basic terrain factors are analyzed using ArcGIS / QGIS and other software, and the elevation, slope, aspect, and profile curvature data of the pixel where the sample point is located are extracted according to the sampling point location as the terrain factor data of the sample point; that is, the terrain factor is calculated based on the digital elevation model DEM data. The slope is the angle between the tangent plane of the sampling point and the horizontal ground, the aspect is the angle between the projection of the normal vector of the tangent plane of the sampling point on the horizontal plane and the due north direction through the point, and the profile curvature is a measure of the elevation change rate of the slope of the sampling point along the maximum slope drop direction.

[0060] Step 120, calculating a plurality of dual-band spectral indices and a plurality of spectral parameters according to the hyperspectral surface reflectance data; and calculating the correlation between the soil organic matter content and the spectral parameters, and the soil organic matter content and the dual-band spectral index.

[0061] In some embodiments, the calculation of spectral index and spectral parameters based on the hyperspectral surface reflectance data includes: smoothing the hyperspectral surface reflectance data; calculating spectral parameters based on the smoothed hyperspectral surface reflectance data; and spectrally transforming the smoothed hyperspectral surface reflectance data to obtain multiple bands; and calculating multiple dual-band spectral indexes based on the hyperspectral surface reflectance data corresponding to each band. Before the hyperspectral surface reflectance data is smoothed, the hyperspectral surface reflectance data can also be spectrally transformed to achieve the purpose of retaining the main information, reducing the amount of data, and enhancing or extracting useful information through function transformation in order to address the correlation and data redundancy of multispectral images. For example, the inverse, logarithm, square root, first-order differential, etc. can be calculated.

[0062] In some embodiments, the hyperspectral surface reflectance data is smoothed using the Savitzky-Golay filtering method, and spectral parameters such as mean and slope are calculated for the smoothed reflectance data, as shown in Table 1.

[0063] Table 1: Spectral parameter calculation formula table

[0064] ,

[0065] Continuation of Table 1:

[0066]

[0067] Of course, the formulas for calculating spectral parameters given in Table 1 can be manually added, deleted, and modified based on experience.

[0068] Then, the Pearson correlation coefficient was used to calculate the correlation between soil organic matter content and spectral parameters. The calculation formula is as follows, and the sensitive bands of soil organic matter content under different spectral transformations are analyzed:

[0069] ,in, , Respectively represent variables X ,variable Y The standard deviation of , Respectively represent variables X ,variable Y The mean of .

[0070] In some embodiments, according to the reflectance of any two bands, four dual-band spectral indices, namely, difference index, ratio index, normalized difference index, and difference square root index, are constructed by K-fold cross validation, and the calculation formula is as follows:

[0071] DI=p - q, RI = p / q,NDI=(pq) / ( p + q),DSI= ,

[0072] in, p , q is the hyperspectral surface reflectance data of any two bands, and p–q≠0 Then, the Pearson correlation coefficient was used to calculate the correlation between the dual-band spectral index and soil organic matter content.

[0073] Step 130: construct a feature data set according to the spectral parameters and spectral indices corresponding to the correlations whose absolute values ​​are greater than a preset threshold, and the terrain factor data.

[0074] In some embodiments, the various indexes are sorted in descending order according to the absolute value of the correlation (correlation coefficient). When the absolute value of the correlation coefficient is greater than 0.4, it is considered that there is a certain correlation between the dual-band spectral index / spectral parameter and the soil organic matter content, and the dual-band spectral index / spectral parameter is retained. The retained dual-band spectral index and spectral parameter are used as features, and then a band-by-band quality analysis is performed, that is, image strip noise detection is performed. Image strip noise detection compares all pixels in adjacent columns of the data. If the proportion of pixels in the current column that are larger or smaller than the pixel value of the next column is greater than a preset proportion, the current column is determined as a strip column. The formula is as follows:

[0075] or ,

[0076] in, Count is the number of pixels that meet the conditions, M is the number of lines in the current spectrum. threshold_sp1 and threshold_sp2 is the band judgment threshold, band j,i and band j,i+1 Represents two adjacent columns of pixels.

[0077] The relative radiation error evaluation includes problems such as inconsistency between CCD chips, bad lines, and abnormal bright spots. Considering the comprehensive image strip noise and relative radiation error problems, the criteria for image quality evaluation are divided as shown in Table 2. It is considered that when the value is 0 and 1, it indicates no quality problem or a minor quality problem; when the value is 2 and 3, it indicates a relatively serious quality problem, and the impact of the characteristic image on the application effect is relatively significant. In order to reduce the impact of quality problems on the results, the characteristics with detection values of 0 and 1 are used.

[0078] Table 2: Corresponding relationship table between image strip noise and values

[0079]

[0080] Then, combining the remaining features with the terrain factor data, a feature dataset is constructed.

[0081] Step 140, train the soil organic matter content monitoring model according to the feature dataset, calculate the model accuracy, and calculate the accuracy improvement value according to the accuracy.

[0082] In some embodiments, training the soil organic matter content monitoring model according to the feature dataset and calculating the model accuracy includes: performing feature optimization calculation on the feature dataset to obtain feature importance; sorting the features from high to low according to the feature importance, and sequentially inputting the features into the soil organic matter content monitoring model for training according to the sorting order, and calculating the accuracy of the model.

[0083] The feature optimization calculation methods include: combined random frog RF, competitive self-organization selection CARS, and variable importance factor VIP.

[0084] In some embodiments, the calculation method for the accuracy of the soil organic matter content monitoring model includes adjusted determination coefficient, root mean square error, and relative analysis error. According to the feature importance ranking, features are gradually added based on the integrated boosting tree LSBoost model for inversion, and the adjusted determination coefficient R 2 _adj, root mean square error RMSEP, and relative analysis error RPD multi-index comprehensive evaluation method are used to evaluate the model accuracy, and curves of multiple indexes with the addition of features are obtained. The index formulas are as follows:

[0085] ,

[0086] ,

[0087] ,

[0088] ,

[0089] Among them, Indicates the measured value, represents the mean of the measured values, represents the predicted value, n Indicates the number of samples, p Indicates the number of features.

[0090] According to R 2 _adj, RMSEP, and RPD values ​​are used to calculate the comprehensive improvement of accuracy of each feature under different feature selection methods. i When i=0, it means that the importance of this feature ranks first; when i>1, the calculation formula is as follows. i When >0, it is considered that the feature has a significant improvement in accuracy.

[0091] ,

[0092] in, , It indicates the maximum and minimum values ​​of the adjusted determination coefficient obtained when using the feature for inversion. represents the adjusted determination coefficient obtained when the i-th feature is used for inversion, and i represents the order of the feature; , It represents the maximum and minimum values ​​of the root mean square error obtained when using the feature for inversion. represents the root mean square error obtained when inverting using the i-th feature; , It indicates the maximum and minimum values ​​of the relative analysis error obtained when using the feature for inversion. Represents the relative analysis error obtained when inverting using the i-th feature.

[0093] Step 150, using the bands and terrain factor types corresponding to the features with precision improvement values ​​greater than 0 as data acquisition conditions for monitoring the soil organic matter content in the monitored area.

[0094] In some embodiments, the soil organic matter content monitoring model used when monitoring the soil organic matter content in the monitored area is a model trained using a feature data set corresponding to features with a precision improvement value greater than 0.

[0095] In some embodiments, the bands and terrain factor types corresponding to the features with accuracy improvement values ​​greater than 0 are used as data acquisition conditions when monitoring the soil organic matter content in the monitored area.

[0096] In some embodiments, the features with the above-mentioned accuracy improvement value greater than 0 can also be inverted based on a machine learning model (for example, an integrated boosting tree LSBoost model, a Gaussian process regression model), using an adjusted determination coefficient R 2_adj, root mean square error, and relative analysis error are used together to evaluate the prediction ability of the model, and a curve showing the change in accuracy with the number of features is obtained. When the coefficient reaches a peak and begins to decrease, the features before the peak are retained, and the features after the peak are removed, and the features are further screened. For example, 7 features are selected, as shown in Table 3, and used as data acquisition conditions for monitoring soil organic matter content in the monitored area.

[0097] Table 3: Characteristics and calculation formulas used as final data acquisition conditions

[0098]

[0099] In some embodiments, four machine learning models, namely, Gaussian process regression model, regression tree model, stepwise linear regression model, and integrated boosting tree model, are constructed according to the data set corresponding to the selected features, and the model parameters are optimized using the 10-fold cross validation method. Inversion is performed based on the machine learning models respectively, and the determination coefficient R is adjusted. 2 The model accuracy was compared by three indicators: _adj, root mean square error RMSEP, and relative analysis error RPD. Finally, the LSBoost model with the highest accuracy (model parameters: 30 ensemble learning cycles, 8 minimum leaf nodes, and 0.1 learning rate) was used as the soil organic matter content monitoring model. Table 4 shows the comparison results of different soil organic matter content inversion models. Figure 3 is the correlation between the measured and predicted values ​​of soil organic matter content, where R 2 =0.83, R 2 _adj=0.79, root mean square error RMSEP=5.15g / kg, relative analysis error RPD=2.32.

[0100] Table 4: Comparison results of different soil organic matter content inversion models

[0101]

[0102] Among them, R 2 The corresponding relationship between the numerical range of _adj and RPD and the five levels of model accuracy is as follows: Excellent model (R 2 _adj≥0.92;RPD≥2.4), good model (0.92>R 2 _adj≥0.81; 2.4>RPD≥2.0), approximate model (0.81>R 2 _adj≥0.64; 2.0>RPD≥1.8), with certain inversion capability (0.64>R 2 _adj≥0.49; 1.8>RPD≥1.4), no inversion capability (R 2 _adj<0.49;RPD<1.4).

[0103] Based on the above method 100, a soil organic matter content monitoring method based on hyperspectral data is provided in the present disclosure, including the following steps: obtaining hyperspectral surface reflectance data of a preset band and terrain factor data of a preset type in the area to be monitored; the band and the type are respectively the band and the type of terrain factor in the data acquisition conditions obtained in the above method; calculating spectral index and spectral parameters according to the hyperspectral surface reflectance data; inputting the spectral index, spectral parameters and the terrain factor data into a pre-trained soil organic matter content monitoring model, and outputting the soil organic matter content of the area to be monitored; the pre-trained soil organic matter content monitoring model is a model trained from a feature data set corresponding to features with accuracy improvement values ​​greater than 0.

[0104] In some embodiments, the trained integrated boosting tree LSBoost model is selected to detect the soil organic matter content in the monitored area, that is, the soil organic matter content monitoring model described above is applied to hyperspectral satellite images to realize the monitoring of regional soil organic matter content and perform remote sensing mapping, such as Figure 4 shown.

[0105] Based on this, the specific implementation process and methods of soil sample collection, satellite-borne hyperspectral data preprocessing, terrain factor extraction, soil organic matter spectral feature analysis, feature image quality evaluation, multi-algorithm joint feature selection for multi-index evaluation, soil organic matter content inversion model construction and remote sensing mapping are given. It has the advantages of strong timeliness, high accuracy, and spatial continuity. It can be widely used in multi-source hyperspectral satellite data. The features are easy to calculate and the quality is stable. The model is highly operable, and the scale, batch, and multi-satellite collaborative business monitoring capabilities of soil organic matter content are realized, which greatly increases the frequency of soil organic matter content monitoring and effectively improves the accuracy of soil organic matter content monitoring.

[0106] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0107] The above is an introduction to the method embodiment. The following is a further explanation of the scheme disclosed in the present invention through an apparatus embodiment.

[0108] Figure 5 FIG. 5 shows a block diagram of a soil organic matter content monitoring model training device 500 based on hyperspectral data according to an embodiment of the present disclosure. Figure 5 As shown, the device 500 includes:

[0109] The data acquisition module 510 is used to acquire the high-spectral surface reflectance data of the sample point, the terrain factor data of the sample point and the soil organic matter content of the sample point;

[0110] A correlation calculation module 520 is used to calculate a plurality of dual-band spectral indices and a plurality of spectral parameters according to the hyperspectral surface reflectance data; and calculate the correlation between the soil organic matter content and the spectral parameters, and the soil organic matter content and the dual-band spectral index;

[0111] A data set construction module 530 is used to construct a feature data set according to the spectral parameters and spectral indices corresponding to the correlations whose absolute values ​​are greater than a preset threshold, and the terrain factor data;

[0112] The precision calculation module 540 is used to train the soil organic matter content monitoring model according to the characteristic data set, calculate the model precision, and calculate the precision improvement value according to the precision;

[0113] The feature selection module 550 is used to use the bands and terrain factor types corresponding to the features with precision improvement values ​​greater than 0 as data acquisition conditions when monitoring the soil organic matter content in the monitored area; the soil organic matter content monitoring model used when monitoring the soil organic matter content in the monitored area is a model trained with the feature data set corresponding to the features with precision improvement values ​​greater than 0.

[0114] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0115] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0116] Figure 6 A schematic block diagram of an electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0117] The electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a ROM 602 or a computer program loaded from a storage unit 608 into a RAM 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An I / O interface 605 is also connected to the bus 604.

[0118] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0119] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as method 100. For example, in some embodiments, the method 100 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the method 100 in any other appropriate manner (e.g., by means of firmware).

[0120] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0122] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0124] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0125] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0126] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0127] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A soil organic matter content monitoring model training method based on hyperspectral data, characterized in that: include: Obtain the hyperspectral surface reflectance data of the sample points, the terrain factor data of the sample points, and the soil organic matter content of the sample points; Calculating a plurality of dual-band spectral indices and a plurality of spectral parameters according to the hyperspectral surface reflectance data; and calculating the correlation between the soil organic matter content and the spectral parameters, and the soil organic matter content and the dual-band spectral index; Constructing a characteristic data set according to the spectral parameters and spectral indices corresponding to the correlations whose absolute values ​​are greater than a preset threshold, and the terrain factor data; The soil organic matter content monitoring model is trained according to the characteristic data set, the model accuracy is calculated, and the accuracy improvement value is calculated according to the accuracy; the calculation method of the accuracy of the soil organic matter content monitoring model includes adjusting the determination coefficient, the root mean square error, and the relative analysis error; the accuracy improvement value S of the soil organic matter content monitoring model is calculated. i The calculation formula is: ,in, , It indicates the maximum and minimum values ​​of the adjusted determination coefficient obtained when using the feature for inversion. represents the adjusted determination coefficient obtained when inversion is performed using the i-th feature, where i represents the order of the feature; , It represents the maximum and minimum values ​​of the root mean square error obtained when using the feature for inversion. represents the root mean square error obtained when inverting using the i-th feature; , It indicates the maximum and minimum values ​​of the relative analysis error obtained when using the feature for inversion. represents the relative analysis error obtained when using the i-th feature for inversion; The bands and terrain factor types corresponding to the features with precision improvement values ​​greater than 0 are used as data acquisition conditions when monitoring the soil organic matter content in the monitored area; the soil organic matter content monitoring model used when monitoring the soil organic matter content in the monitored area is a model trained with the feature data set corresponding to the features with precision improvement values ​​greater than 0.

2. The method according to claim 1, characterized in that The calculating of spectral index and spectral parameter according to the hyperspectral surface reflectance data comprises: Smoothing the hyperspectral surface reflectance data; Calculating spectral parameters according to the smoothed hyperspectral surface reflectance data; and performing spectral transformation on the smoothed hyperspectral surface reflectance data to obtain multiple bands; Multiple dual-band spectral indices are calculated based on the hyperspectral surface reflectance data corresponding to each band.

3. The method according to claim 2, characterized in that The calculation methods of spectral parameters include: averaging and slope; The formula for calculating the spectral index based on the hyperspectral surface reflectance data corresponding to each band is: , in, p , q is the hyperspectral surface reflectance data of any two bands, and p–q≠0 .

4. The method according to claim 1, characterized in that: The step of training the soil organic matter content monitoring model according to the characteristic data set and calculating the model accuracy includes: Performing feature optimization calculation on the feature data set to obtain feature importance; The features are sorted from high to low according to the feature importance, and the features are input into the soil organic matter content monitoring model in sequence according to the sorting order for training, and the accuracy of the model is calculated.

5. The method according to claim 4, characterized in that The feature optimization calculation methods include: joint random frog RF, competitive self-organizing selection CARS, and variable importance factor VIP.

6. A soil organic matter content monitoring method based on hyperspectral data, characterized in that: include: Acquire high-spectral surface reflectance data of a preset band and terrain factor data of a preset type in the area to be monitored; the band and the type are respectively the band and terrain factor type in the data acquisition conditions obtained in any one of claims 1 to 5; Calculating spectral index and spectral parameters according to the hyperspectral surface reflectance data; The spectral index, spectral parameters and terrain factor data are input into a pre-trained soil organic matter content monitoring model, and the soil organic matter content of the monitored area is output; the pre-trained soil organic matter content monitoring model is a model trained with a feature data set corresponding to features with accuracy improvement values ​​greater than 0.

7. A soil organic matter content monitoring model training device based on hyperspectral data, characterized in that: include: A data acquisition module is used to obtain the high-spectral surface reflectance data of the sample points, the terrain factor data of the sample points, and the soil organic matter content of the sample points; A correlation calculation module, used to calculate a plurality of dual-band spectral indices and a plurality of spectral parameters according to the hyperspectral surface reflectance data; and calculate the correlation between the soil organic matter content and the spectral parameters, and the soil organic matter content and the dual-band spectral index; A data set construction module, used to construct a feature data set according to the spectral parameters and spectral indices corresponding to the correlations whose absolute values ​​are greater than a preset threshold, and the terrain factor data; The accuracy calculation module is used to train the soil organic matter content monitoring model according to the characteristic data set, calculate the model accuracy, and calculate the accuracy improvement value according to the accuracy; the calculation method of the accuracy of the soil organic matter content monitoring model includes adjusting the determination coefficient, the root mean square error, and the relative analysis error; the accuracy improvement value S of the soil organic matter content monitoring model i The calculation formula is: ,in, , It indicates the maximum and minimum values ​​of the adjusted determination coefficient obtained when using the feature for inversion. represents the adjusted determination coefficient obtained when inversion is performed using the i-th feature, where i represents the order of the feature; , It represents the maximum and minimum values ​​of the root mean square error obtained when using the feature for inversion. represents the root mean square error obtained when inverting using the i-th feature; , It indicates the maximum and minimum values ​​of the relative analysis error obtained when using the feature for inversion. represents the relative analysis error obtained when using the i-th feature for inversion; The feature selection module is used to use the bands and terrain factor types corresponding to the features with precision improvement values ​​greater than 0 as data acquisition conditions when monitoring the soil organic matter content in the monitored area; the soil organic matter content monitoring model used when monitoring the soil organic matter content in the monitored area is a model trained with the feature data set corresponding to the features with precision improvement values ​​greater than 0.

8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Soil heat flux prediction method based on multi-source satellite remote sensing data

    CN114563353A

  • Soil salinity inversion and salinization risk assessment method, device and equipment

    CN116757099A