Soil organic carbon prediction method and device based on data enhancement technology
The soil sample set is expanded through data augmentation technology, and combined with the random forest model, the problem of low accuracy of the soil organic carbon prediction model under small sample size is solved, achieving high-precision and stable rapid prediction of soil organic carbon.
Patent Information
- Application Number
- CN202510092612.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-30
AI Technical Summary
When using near-infrared spectroscopy technology to quickly predict soil properties, the prior art is limited by small sample size, resulting in low model accuracy and poor stability.
Using data augmentation technology, the rare and important sample sets are expanded through the combination of SmoteR method and Gaussian noise, and a high-precision soil organic carbon prediction model is obtained through random forest model training.
The accuracy and stability of the soil organic carbon prediction model are significantly improved, the problem of low model accuracy under small sample size is solved, and efficient rapid prediction of soil organic carbon is achieved.
Smart Images

Figure CN120064166A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of soil organic carbon prediction, and particularly relates to a method and device for predicting soil organic carbon based on data enhancement technology. Background Art
[0002] Soil organic carbon (SOC) is the largest carbon pool on land. The accurate and rapid acquisition of soil organic carbon is of great significance for improving soil fertility, ensuring food security, and monitoring climate change. Although traditional laboratory test and analysis methods can obtain high precision, the field soil sampling and laboratory chemical analysis have a long cycle, high cost, complex process, and poor real-time performance. Limited by the number of actual analysis samples, it is difficult to monitor soil organic carbon on a large scale. In addition, the acid-base waste generated by a large number of soil chemical tests and analyses can easily cause environmental pollution if not properly treated. The use of soil hyperspectral technology represented by visible-near-infrared spectroscopy technology to monitor soil property information has the advantages of being fast, simple, non-contact, environmentally friendly, and inexpensive. Scientific researchers in various countries have conducted a large number of studies and experiments on this, and the effects are remarkable.
[0003] Literature: Qi, M., Chen, S., Wei, Y., Zhou, H., Zhang, S., Wang, M., Zheng, J., Rossel, R. a. V., Chang, J., Shi, Z., & Luo, Z. (2024). Using visible-near infrared spectroscopy to estimate whole-profile soil organic carbon and its fractions. Soil & Environmental Health, 2(3), 100100. (Rossel, R. a. V., Shen, Z., Lopez, L. R., Behrens, T., Shi, Z., Wetterlind, J., Sudduth, K. A., Stenberg, B., Guerrero, C., Gholizadeh, A., Ben-Dor, E., St Luce, M., & Orellano, C. (2024). An imperative for soil spectroscopic modelling is to think global but fit local with transfer learning. Earth-Science Reviews, 254, 104797.) (Zhao, X., Wang, J., Koganti, T., & Triantafilis, J. (2024). Developing Vis–NIR libraries to predict cation exchange capacity (CEC) and pH in Australian sugarcane soil. Computers and Electronics in Agriculture, 221, 109004. It is disclosed that the number of soil samples can greatly affect the accuracy of the prediction model. However, a large amount of manpower and material resources are still required in the soil sample collection process. Therefore, the number of samples is one of the main limiting factors for the rapid prediction of soil properties using near-infrared spectroscopy. Prediction models established based on a small sample size often have problems such as low model accuracy and poor stability.
[0004] The emergence of data augmentation techniques provides a new idea for solving this bottleneck problem. The literature "Branco, P. O., Torgo, L., & Ribeiro, R. P. (2017). SMOGN: a Pre-processing Approach for Imbalanced Regression. Learning With Imbalanced Domains: Theory and Applications, 36–50." (Agrawal, A., Petersen, M. R. (2021). Detecting arsenic contamination using satellite imagery and machine learning. Toxics, 9(12), 333.) has applied and explored the use of data augmentation techniques to synthesize samples in fields such as image learning. However, the above literature has not been able to achieve high-precision and rapid prediction of small-sample soil properties by combining visible-near infrared spectroscopy with data augmentation techniques. Summary of the Invention
[0005] The present invention provides a method for predicting soil organic carbon based on data augmentation techniques, which can achieve high-precision and rapid prediction of small-sample soil properties.
[0006] The present invention provides a method for predicting soil organic carbon based on data augmentation techniques, including:
[0007] Obtain the spectral data of the target soil and the corresponding organic carbon content, and construct an original sample set based on the spectral data of the target soil;
[0008] Screen out extreme values from the organic carbon content sorted from high to low by the interquartile range method, calculate the phi coefficient between the organic carbon content and the extreme values, and use the spectral data corresponding to the organic carbon content with a phi coefficient greater than the correlation threshold as a rare and important sample set. Calculate the Manhattan distance between each sample in the rare and important sample set and the remaining samples, select K nearest neighbor samples from the remaining samples based on the obtained Manhattan distance, and use the samples with a Manhattan distance less than or equal to the set distance threshold among the K nearest neighbor samples as safe-distance samples. Each sample and its corresponding safe-distance sample are synthesized into new samples by the SmoteR method to expand the sample size. Use the samples with a Manhattan distance greater than the set distance threshold among the K nearest neighbor samples as unsafe-distance samples, introduce Gaussian noise to the unsafe-distance samples to synthesize new samples, and achieve data augmentation by synthesizing new samples in the original samples to obtain an expanded sample set. Divide the expanded sample set into a validation set and a training set;
[0009] Train a random forest model based on the training set to obtain a soil organic carbon prediction model.
[0010] Preferably, verify the soil organic carbon prediction model based on the validation sample set to obtain the prediction accuracy of the soil organic carbon prediction model. Adjust the prediction accuracy of the soil organic carbon prediction model by adjusting the correlation threshold and the K value. By comparing the obtained prediction accuracies, use the correlation threshold and the K value corresponding to the highest prediction accuracy as the final correlation threshold and K value, thereby obtaining the final soil organic carbon prediction model.
[0011] Preferably, use the coefficient of determination, the root mean square error, and the ratio of the difference between the third quartile and the first quartile to the standard prediction error as the evaluation parameters of the prediction accuracy.
[0012] Preferably, the method for obtaining the distance threshold includes sorting the obtained Manhattan distances from largest to smallest to obtain a distance sequence, and using half of the distance value corresponding to the median of the distance sequence as the distance threshold.
[0013] Preferably, each sample and the corresponding safe-distance sample are synthesized into new samples by the SmoteR method, including: using the method of linear interpolation to form new samples between each sample and the corresponding safe-distance sample.
[0014] Preferably, introduce Gaussian noise to the unsafe-distance samples to obtain new samples, including randomly adding a random perturbation term to the unsafe-distance samples, and the perturbation term follows a Gaussian distribution.
[0015] Preferably, screen out extreme values from the sorted organic carbon contents from high to low by the interquartile range method, including:
[0016] Obtain the first quartile Q1 and the third quartile Q3 from the sorted organic carbon contents from high to low, and calculate the difference between the third quartile and the first quartile to obtain the interquartile range IQR;
[0017] Points greater than Q3 + 1.5×IQR are extreme maxima, and points less than Q1 - 1.5×IQR are extreme minima. Construct extreme values through the extreme maxima and extreme minima.
[0018] Preferably, use the spectral data corresponding to the organic carbon contents with the phi coefficient less than or equal to the correlation threshold as a common and unimportant sample set, and randomly delete the samples in the common and unimportant sample set by the random undersampling method.
[0019] Preferably, remove the bands with excessive noise in the obtained spectral data of the target soil before constructing the original sample set.
[0020] The present invention also provides a soil organic carbon prediction device based on data augmentation technology, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the soil organic carbon prediction method based on data augmentation technology.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0022] By appropriately expanding the number of samples that have a high correlation with extreme values and are within the safe distance range, that is, the Manhattan distance in the K nearest neighbor samples is less than or equal to the set distance threshold, the present invention can accurately expand rare and important samples. While ensuring the quality of the training samples, the amount of training samples is also appropriate, which is conducive to efficiently training a random forest model to obtain a soil organic carbon prediction model that can accurately and quickly predict the soil organic carbon content. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic diagram of the soil organic carbon prediction method based on data augmentation technology proposed in a specific embodiment of the present invention;
[0024] Figure 2 It is a schematic diagram of the data augmentation algorithm SMOGN proposed in a specific embodiment of the present invention;
[0025] Figure 3 It is a comparison chart of soil organic carbon prediction results using the original sampling test data and data augmentation synthetic samples proposed in a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The present invention will be further described below with reference to specific embodiments.
[0027] In order to achieve high-precision and rapid prediction of small-sample soil properties, specific embodiments of the present invention expand rare and important samples, so as to efficiently train a soil organic carbon prediction model that can quickly and accurately predict the soil organic carbon content.
[0028] Specific embodiments of the present invention provide a soil organic carbon prediction method based on data augmentation technology, as Figure 1 shown:
[0029] (1) Acquisition of soil chemical properties and soil spectral data: The soil organic carbon content is measured using the international standard (ISO 10694.1995). The soil spectrum is obtained by a FOSS XDS spectrometer equipped with a high-intensity contact probe with a built-in light source, and the initial spectral data with a wavelength range of 400 - 2500 nm.
[0030] (2) Spectral data preprocessing: Remove the bands with excessive noise in the initial spectral data. The remaining spectral data is the 500 - 2500 nm band in the initial spectrum, obtaining the original sample set and the corresponding organic carbon content.
[0031] (3) As Figure 2 shown, perform data augmentation on the original sample set to obtain an expanded sample set, and divide the expanded sample set into a validation set and a training set:
[0032] Sort the organic carbon content obtained in step (2) from high to low, obtain the first quartile Q1 and the third quartile Q3 from the sorted organic carbon content from high to low, and calculate the difference between the third quartile and the first quartile to obtain the interquartile range IQR; the points greater than Q3 + 1.5×IQR are extreme high values, and the points less than Q1 - 1.5×IQR are extreme low values. Construct extreme values through the extreme high values and extreme low values, where Q1 is the value ranked 25% of the total, and Q3 is the value ranked 75% of the total.
[0033] In a specific embodiment of the present invention, calculate the phi coefficient between the organic carbon content and the extreme values. The spectral data corresponding to the organic carbon content with a phi coefficient greater than the correlation threshold is used as a rare and important sample set, and the spectral data corresponding to the organic carbon content with a phi coefficient less than or equal to the correlation threshold is used as a common and unimportant sample set.
[0034] The larger the phi coefficient provided in the specific embodiment of the present invention indicates that the organic carbon content is more correlated with the extreme values. Since the range of the phi coefficient is between 0 and 1, the size of the correlation threshold is taken between 0 and 1. By setting the correlation threshold, the present invention can flexibly condition the rare and important sample size based on the prediction accuracy.
[0035] In a specific embodiment of the present invention, randomly delete and discard samples in the common and unimportant sample set through the random undersampling method to obtain the remaining samples. The proportion of randomly deleting and discarding samples is random, ranging from 0 to 1, which is beneficial to reducing the computational overhead while ensuring a good training effect.
[0036] In a specific embodiment of the present invention, an oversampling method is used for the rare and important sample set, that is, sample expansion is performed through the SmoteR method or the method of introducing Gaussian noise. The specific steps are as follows:
[0037] In a specific embodiment of the present invention, the Manhattan distance between each sample in a rare and important sample set and the remaining samples is calculated by the KNN algorithm. Based on the obtained Manhattan distance, K nearest neighbor samples are selected from the remaining samples. Samples with a Manhattan distance less than or equal to a set distance threshold among the K nearest neighbor samples are used as safe distance samples. Each sample and its corresponding safe distance sample are synthesized into new samples by the SmoteR method to expand the sample size. Samples with a Manhattan distance greater than the set distance threshold among the K nearest neighbor samples are used as unsafe distance samples, and Gaussian noise is introduced into the unsafe distance samples to synthesize new samples. Data augmentation is achieved by synthesizing new samples in the original samples to obtain an expanded sample set. The expanded sample set is divided into a validation set and a training set, where the range of the K value is 0 - the number of rare and important samples.
[0038] Data with a high correlation with extreme values indicates that it is rare and has a small quantity. During the actual prediction process, the small quantity leads to insufficient representativeness of the samples, causing the problem of data imbalance, affecting the generalization and fairness of the model, and there are samples with insufficient representativeness that affect the prediction accuracy of the model. Therefore, it is necessary to expand the samples with a high correlation with extreme values. In a specific embodiment of the present invention, the SmoteR linear interpolation is used to expand the number of samples with a high correlation with extreme values and within the safe distance range. The main consideration is that since linear interpolation constructs a relationship between points: the value of the synthesized sample point = the value of the original sample point + z × (the value of the nearest neighbor sample point - the value of the original sample point), when the distance between points is too far, the relationship between the values of the two points is not high, and if linear interpolation is performed, there is a higher risk, and the obtained value may lack representativeness. Therefore, a more conservative method is used in a specific embodiment of the present invention. By setting a distance threshold, this method performs interpolation between points with a relatively close distance, and uses Gaussian noise when the distance is far, allowing an increase in generalization ability in rare cases where the decision boundary can be expanded.
[0039] In a specific embodiment, the method for obtaining the distance threshold in a specific embodiment of the present invention includes sorting the obtained Manhattan distances from largest to smallest to obtain a distance sequence, and taking half of the distance corresponding to the median of the distance sequence as the distance threshold. Since the median represents the central tendency in the data set, in order to adapt to the distribution of different data sets, balance the risk and diversity of generating synthetic samples, reduce the influence of outliers on the distance metric, and make the algorithm more robust, this embodiment takes half of the distance corresponding to the median of the distance sequence as the distance threshold.
[0040] In a specific embodiment of the present invention, taking half of the distance corresponding to the median of the distance sequence as the distance threshold can more strictly limit the distance between samples, thereby reducing the risk of using samples with a relatively large distance for synthesis during the interpolation process. This is because samples with a relatively large distance may have significant differences in the feature space, and directly interpolating may generate synthetic samples that do not conform to the data distribution law, thereby affecting the generalization ability of the model. The median may lead to over-interpolation. If the median is used as the threshold, more samples may be considered "safe" and thus used for interpolation. Although this can increase the number of synthetic samples, it also increases the probability of generating unrealistic samples, especially in the case of uneven data distribution. By using half of the median as the threshold, it can be ensured that interpolation is only performed when the distance between samples is relatively close. The synthetic samples generated in this way are closer to the local structure of the original data and can better maintain the authenticity and diversity of the data. A shorter distance threshold can effectively reduce the influence of noise samples on the synthetic samples. In the dataset, noise samples may cause the synthetic samples to deviate from the true data distribution, and using half of the median can better filter out these noise samples. The distribution characteristics of different datasets are different, and using half of the median as the threshold can better adapt to various data distributions. In some datasets, the distance between samples may be relatively large, and using the median as the threshold may result in too many samples being used for interpolation, while half of the median can more flexibly adjust the interpolation range to adapt to different data distributions. In the presence of outliers or uneven data distribution, half of the median can provide a more robust threshold selection. The median itself is insensitive to outliers, and its half further reduces the influence of outliers on the interpolation process, improving the robustness of the algorithm. In the experimental verification of the SMOGN algorithm, using half of the median as the distance threshold showed good performance on multiple datasets. The experimental results show that this choice can effectively improve the generalization ability and prediction accuracy of the model while reducing the risk of overfitting.
[0041] Specifically, each sample and the corresponding safe-distance sample are synthesized into new samples through the SmoteR method, including: using the method of linear interpolation to form new samples between each sample and the corresponding safe-distance sample. Specifically for linear interpolation, for each nearest neighbor sample, a random number z (between 0 and 1) is calculated, and then according to the formula: synthetic sample point value = original sample point value + z × (nearest neighbor sample point value - original sample point value). It can be understood that the sample point values provided in this embodiment include organic carbon values and corresponding spectral values.
[0042] In a specific embodiment, the specific embodiment of the present invention introduces Gaussian noise into the unsafe distance samples to obtain new samples, including randomly adding a random perturbation term to the unsafe distance samples, and the perturbation term follows a Gaussian distribution, that is, the synthetic sample point value = the original sample point value + the random perturbation term. The above operation can simulate the noise in real-world data.
[0043] (5) Establish a soil organic carbon prediction model based on data augmentation technology: Use the data obtained by the processing in step (4) for data augmentation, and randomly divide the augmented sample set into a training set and a validation set according to a ratio of 2:1. Use the spectral data of the training set as the input of the random forest model, and use the soil organic carbon content of the soil samples as the output of the model to obtain a soil organic carbon prediction model based on data augmentation, and use the soil organic carbon prediction model to predict the organic carbon content of the soil to be measured.
[0044] The specific embodiment of the present invention also includes evaluating the accuracy of the prediction results of the soil organic carbon prediction model using the validation set: Use three model evaluation parameters to evaluate the prediction model, and compare the prediction results based on the original sampling test data and the prediction results based on data augmentation.
[0045] The specific embodiment of the present invention verifies the prediction accuracy of the soil organic carbon prediction model based on the validation sample set to obtain the prediction accuracy of the soil organic carbon prediction model. Adjust the prediction accuracy of the soil organic carbon prediction model by adjusting the correlation threshold and the K value. By comparing the obtained prediction accuracies, use the correlation threshold and the K value corresponding to the highest prediction accuracy as the final correlation threshold and K value, so as to obtain the final soil organic carbon prediction model.
[0046] Compared with using the original sampling test data for soil organic carbon modeling, the data augmentation synthetic samples obtained by using this method have greatly improved the model prediction accuracy and stability, solved the problem of low accuracy and stability of the prediction model with a small sample size, and provided a method with high detection accuracy and strong stability for soil organic carbon prediction.
[0047] In an embodiment 1, this embodiment provides a soil organic carbon prediction method based on data augmentation synthetic samples. In this embodiment, a total of 92 soil surface layer samples (0-20 cm) are collected for research, including:
[0048] (1) Acquisition of soil chemical properties and soil spectral data: Measure soil organic carbon using the international standard (iso ISO 10694.1995). Soil spectra are obtained by a FOSS XDS type spectrometer equipped with a high-intensity contact probe with a built-in light source, and its wavelength range is 400-2500 nm.
[0049] (2) Spectral data preprocessing: Remove the bands with excessive noise in the initial spectrum. The remaining spectral data is the 500 - 2500 nm band in the initial spectrum, obtaining the original sample set and the corresponding organic carbon content.
[0050] (3) Data augmentation: Based on the samples of the spectral data retained in step (2), according to the correlation between soil organic carbons, by setting the threshold to 0.4, the samples are divided into a rare and important sample set and a common and unimportant sample set. Among them, the samples with a correlation lower than the threshold are used to construct the common and unimportant sample set, and the rest are the rare and important sample set. For the common and unimportant sample set, the random undersampling method is used for processing, and some samples are randomly discarded. For the rare and important sample set, the oversampling method is used for processing.
[0051] According to the distance between the selected sample and its K = 5 nearest neighbor samples, oversampling will use SmoteR or introduce a Gaussian noise strategy to generate synthetic samples. The main idea is: if the selected nearest neighbor sample is "safe", that is, a sample within the safe distance, SmoteR is used to synthesize new samples; if the nearest neighbor sample is not within the safe range, that is, an unsafe distance sample, it cannot be used for interpolation. At this time, a new sample will be generated by introducing Gaussian noise. The threshold for determining "safe" or "unsafe" distance depends on the distance between the sample and all the remaining samples in the partition, and is set to half of the median of this distance.
[0052] (4) Establish a soil organic carbon prediction model based on data augmentation: According to the data enhanced in step (3), the data is randomly divided into a training set and a validation set in a ratio of 2:1. Use the random forest algorithm to establish a prediction model of visible - near - infrared spectroscopy based on data augmentation.
[0053] (5) Evaluate the accuracy of the model prediction results: Select the coefficient of determination (R 2 ), root mean square error (RMSE), and the ratio of the difference between the third quartile and the first quartile to the standard prediction error (RPIQ) as evaluation parameters, and compare the prediction results based on the original sampling test data and the prediction results based on the data - augmented synthetic samples.
[0054] In Comparative Example 1, the difference between this Comparative Example 1 and the above - mentioned Example 1 is that after obtaining the original sample set and the corresponding organic carbon content in step (2) of Example 1, data augmentation is not performed, and the following step (3) is directly carried out.
[0055] (3) Establish a soil organic carbon prediction model based on the original sampling test data. According to the spectral data retained in step (2), divide the data into a training set and a validation set randomly at a ratio of 2:1. Use the random forest algorithm to establish a visible-near-infrared spectral soil organic carbon prediction model based on the original sampling test data.
[0056] (4) Evaluate the accuracy of the model prediction: Select the coefficient of determination (R 2 ), root mean square error (RMSE), and the ratio of the difference between the third quartile and the first quartile to the standard prediction error (RPIQ) as evaluation parameters, and compare the prediction results based on the original sampling test data and the prediction results based on the data augmentation synthetic samples.
[0057] Compare the prediction results based on the original sampling test data and the soil organic carbon prediction results based on the data augmentation synthetic samples. The specific results are shown in Table 1, Figure 3 as follows.
[0058] Table 1 Comparison of soil organic carbon prediction accuracies based on the original sampling test data and data augmentation synthetic samples
[0059]
[0060] As can be seen from Table 1, the model performance of predicting soil organic carbon using the random forest model established with the original sampling test data is low. The R 2 is 0.32, the RMSE is 16.14 g / kg, and the RPIQ is 1.20, so it cannot predict soil organic carbon well. The model prediction performance of predicting soil organic carbon using the random forest model established with the data augmentation synthetic samples is greatly improved compared with the model established using the original sampling test data. The R 2 of the model after data augmentation reaches 0.86. Compared with the R 2 of 0.32 of the model constructed using the original sampling test data, the R 2 is increased by 168.75%. The RMSE is 7.32 g / kg. Compared with the RMSE of 16.14 g / kg using the original sampling test data, the RMSE is reduced by 54.65%. The RPIQ is 2.20. Compared with the RPIQ of 1.20 using the original sampling test data, the RPIQ is increased by 83.33%, which can predict soil organic carbon better and significantly improve the model accuracy and stability of predicting soil organic carbon using visible-near-infrared spectroscopy.
[0061] Figure 3 In, the abscissa represents the measured values of each soil organic carbon, and the ordinate represents the predicted values of soil organic carbon. Figure 3a shows the accuracy of the soil organic carbon prediction model established based on the original sampling test data. Figure 3 b shows the accuracy of the soil organic carbon prediction model representing the one established based on the data augmentation synthetic samples. The black dashed line is the 1:1 line, and the red solid line is the fitting line. Figure 3 It can be seen that the distance of the predicted value of the model established based on the data augmentation synthetic samples from the 1:1 line is less than that of the predicted value of the model established based on the original sampling test data, indicating that the prediction accuracy of the model based on the data augmentation synthetic samples is better than that of the model based on the original sampling test data.
[0062] In the specific implementation of the present invention, a model of soil organic carbon based on visible-near infrared spectroscopy technology combined with the original sampling test data and data augmentation synthetic samples was established, and at the same time, the prediction results of the models of soil organic carbon based on visible-near infrared spectroscopy technology combined with the original sampling test data and combined with data augmentation synthetic samples were compared. It was found that the soil organic carbon prediction model based on visible-near infrared spectroscopy technology and data augmentation technology obtained by the present invention has excellent prediction performance, significantly improving the accuracy and stability of rapid prediction of soil organic carbon in the case of small sample sizes.
Claims
1. A soil organic carbon prediction method based on data enhancement technology, characterized in that: include: Obtaining the spectral data and the corresponding organic carbon content of the target soil, and constructing an original sample set based on the spectral data of the target soil; The extreme values are screened out from the organic carbon content sorted from high to low by the interquartile range method, the phi coefficient of the organic carbon content and the extreme value is calculated, the spectral data corresponding to the organic carbon content with the phi coefficient greater than the correlation threshold is taken as the rare and important sample set, the Manhattan distance between each sample and the remaining samples in the rare and important sample set is calculated, K nearest neighbor samples are selected from the remaining samples based on the obtained Manhattan distance, the samples with Manhattan distance less than or equal to the set distance threshold in the K nearest neighbor samples are taken as safe distance samples, each sample and the corresponding safe distance sample are synthesized into a new sample by the SmoteR method to expand the sample size, the samples with Manhattan distance greater than the set distance threshold in the K nearest neighbor samples are taken as unsafe distance samples, Gaussian noise is introduced into the unsafe distance samples to synthesize new samples, data enhancement is achieved by synthesizing new samples in the original samples to obtain an expanded sample set, and the expanded sample set is divided into a validation set and a training set; The random forest model was trained based on the training set to obtain the soil organic carbon prediction model.
2. The soil organic carbon prediction method based on data enhancement technology according to claim 1 is characterized in that: The soil organic carbon prediction model was verified based on the validation sample set to obtain the prediction accuracy of the soil organic carbon prediction model. The prediction accuracy of the soil organic carbon prediction model was adjusted by adjusting the correlation threshold and K value. The prediction accuracies obtained were compared and the correlation threshold and K value corresponding to the highest prediction accuracy were used as the final correlation threshold and K value to obtain the final soil organic carbon prediction model.
3. The soil organic carbon prediction method based on data enhancement technology according to claim 2 is characterized in that: The coefficient of determination, root mean square error and the ratio of the difference between the third quartile and the first quartile to the standard prediction error were used as evaluation parameters for prediction accuracy.
4. The soil organic carbon prediction method based on data enhancement technology according to claim 1 or 2, characterized in that: The method for obtaining the distance threshold includes sorting the obtained Manhattan distances from large to small to obtain a distance sequence, and taking half of the distance value corresponding to the median of the distance sequence as the distance threshold.
5. The soil organic carbon prediction method based on data enhancement technology according to claim 1 is characterized in that: Each sample and the corresponding safety distance sample are synthesized into a new sample through the SmoteR method, including: using a linear interpolation method to form a new sample between each sample and the corresponding safety distance sample.
6. The soil organic carbon prediction method based on data enhancement technology according to claim 1 is characterized in that: Introducing Gaussian noise into the unsafe distance sample to obtain a new sample includes randomly adding a random disturbance term into the unsafe distance sample, where the disturbance term obeys Gaussian distribution.
7. The soil organic carbon prediction method based on data enhancement technology according to claim 1 is characterized in that: The extreme values were screened out from the organic carbon content sorted from high to low by the interquartile range method, including: The quartile Q1 and the lower quartile Q3 were obtained from the organic carbon content sorted from high to low, and the interquartile range Q3 was obtained by calculating the difference between the lower quartile and the quartile; The point greater than Q3+1.5×IQR is regarded as the maximum value, and the point less than Q1-1.5×IQR is regarded as the minimum value, and the extreme value is constructed by the maximum value and the minimum value.
8. The soil organic carbon prediction method based on data enhancement technology according to claim 1 is characterized in that: The spectral data corresponding to the organic carbon content with phi coefficient less than or equal to the correlation threshold are taken as the common and unimportant sample set, and the samples in the common and unimportant sample set are randomly deleted by random undersampling method.
9. The soil organic carbon prediction method based on data enhancement technology according to claim 1, characterized in that: Before constructing the original sample set, the bands with excessive noise in the spectral data of the target soil are removed.
10. A soil organic carbon prediction device based on data enhancement technology, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the soil organic carbon prediction method based on data enhancement technology as described in any one of claims 1-9.
Citation Information
Cited By
Soil sample screening method based on instance transfer learning
CN121144847A