Method for predicting heavy metals in farmland soil in arid region based on Vis-NIR and pXRF spectrum fusion
By using the fusion method of Vis-NIR and pXRF spectrum and a random forest algorithm in the arid areas, the problems of low concentration heavy metal detection accuracy and high cost in the arid areas in the prior art are solved, and high-precision, fast and low-cost heavy metal prediction are achieved.
Patent Information
- Application Number
- CN202510354508.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems such as limited accuracy, many interference factors and high detection costs in the detection of low-concentration heavy metals in arid areas, and it is difficult to meet the needs of large-scale and rapid monitoring.
The fusion method based on Vis-NIR and pXRF spectrum is adopted, and the spectral quality is optimized through 15 pretreatment methods, and the prediction model is constructed in combination with the random forest machine learning algorithm to achieve efficient prediction of soil heavy metal content.
It realizes high-precision, fast and low-cost soil heavy metal prediction, which is especially suitable for accurate monitoring of low-concentration heavy metals in arid agricultural soils, improving the sensitivity and stability of detection.
Smart Images

Figure CN120148677A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of spectral data processing and soil detection, and in particular to a method for predicting heavy metals in farmland soil based on Vis-NIR and pXRF spectrum fusion, which is suitable for the prediction and ecological risk assessment of low-concentration heavy metals in arid areas. Background Art
[0002] Heavy metal pollution in soil has become a major threat to global agricultural ecological security, especially in arid areas. Due to the fragility of the soil environment, the accumulation of heavy metals in the soil and its long-term impact are difficult to predict. Traditional heavy metal detection methods mainly rely on laboratory analysis, such as inductively coupled plasma mass spectrometry (ICP-MS) and atomic absorption spectroscopy (AAS). Although these methods have high detection accuracy, sample preparation is complex, detection costs are high, and it is difficult to meet the actual needs of large-scale and rapid monitoring.
[0003] In recent years, proximal sensing technology has attracted extensive attention in the field of soil pollution monitoring due to its advantages of non-destructive, rapid and on-site detection. Among them, Vis-NIR and pXRF spectroscopy are two important means of predicting soil heavy metals. Vis-NIR spectroscopy can achieve rapid and non-destructive detection based on the absorption and scattering characteristics of soil to light, and can indirectly reflect the influence of soil organic matter and mineral composition on heavy metal content, which is suitable for large-scale monitoring; while pXRF spectroscopy can directly measure the concentration of heavy metals by detecting the characteristic X-rays of soil elements, without the need for complex sample pretreatment, and has strong on-site detection capabilities. Although both have their own advantages, visible-near infrared spectroscopy is easily affected by soil organic matter and mineral composition, making it difficult to directly detect heavy metals, while pXRF spectroscopy has limited sensitivity in low-concentration heavy metal detection and is greatly affected by the sample matrix effect. A single sensor has problems such as limited accuracy and many interference factors when detecting low-concentration heavy metals, which affects the reliability of monitoring. Therefore, how to fully integrate different spectral information, give full play to their respective advantages, and improve the sensitivity and stability of heavy metal detection has become a technical problem that needs to be solved in this field.
[0004] Therefore, there is an urgent need for an efficient soil heavy metal prediction method that integrates Vis-NIR and pXRF spectra to achieve more accurate and rapid large-scale farmland soil heavy metal monitoring and provide scientific support for agricultural ecological security in arid areas. Summary of the invention
[0005] Aiming at the problems existing in the background technology, the purpose of the present invention is to provide a rapid prediction method for heavy metals in soil based on the spectral fusion of Vis-NIR and pXRF, which is particularly suitable for the precise monitoring of low-concentration heavy metals in arid agricultural soil. The present invention provides a precise prediction method for heavy metals in soil based on the spectral fusion of Vis-NIR and pXRF. First, 15 preprocessing methods are respectively performed on the Vis-NIR and pXRF spectra to optimize the spectral quality, remove noise, and enhance key spectral features. Subsequently, the preprocessed spectral data are fused to make full use of the advantages of organic matter and mineral information of the Vis-NIR spectrum and the direct element detection ability of the pXRF spectrum to achieve efficient prediction of the heavy metal content in soil. In the modeling process, the Random Forest machine learning algorithm is used to construct a prediction model for the fused spectrum, and the optimal preprocessing combination for different heavy metal elements is systematically screened to ensure the stability and high accuracy of the model. In addition, for the selected optimal preprocessing scheme, key spectral bands are further extracted to optimize the model input variables, reduce data redundancy, and improve the prediction performance to meet the needs of sustainable agricultural development and ecological environment protection.
[0006] To achieve the above object, the present invention adopts the following technical solutions: A rapid prediction method for heavy metals in soil based on the spectral fusion of Vis-NIR and pXRF, including:
[0007] Step 1: Obtain the data of n soil samples in the arid area farmland, and perform heavy metal detection on the soil sample data to obtain the content information of heavy metals in the soil;
[0008] Step 2: Obtain the spectral data of n soil samples and perform spectral measurement to obtain the original Vis-NIR and pXRF spectral data;
[0009] Step 3: Perform 15 preprocessings on the original Vis-NIR and pXRF spectral data of the soil respectively;
[0010] Step 4: Directly fuse the 15 kinds of preprocessed Vis-NIR spectral data and 15 kinds of pXRF spectral data to obtain 225 kinds of directly fused spectral data sets;
[0011] Step 5: Based on the fused 225 kinds of spectral data and the measured data of heavy metals in the soil, construct a random forest model;
[0012] Step 6: Evaluate the model results and select the optimal spectral preprocessing combination for each heavy metal;
[0013] Step 7: Based on the optimal spectral preprocessing combination, extract the importance index of the spectral data to determine the key spectral features.
[0014] Further, a plurality of sampling points are arranged in the sampling area according to the uniform sampling method, and the longitude and latitude coordinates of the sampling points are located by using a handheld GPS;
[0015] After collecting the soil samples, they are air-dried at room temperature of 25°C in the laboratory, ground after removing gravel and plant and animal residues, and sieved;
[0016] Preferably, the soil samples are divided into two parts. One part is sieved through a 100-mesh (0.15 mm) nylon sieve to remove larger particles for measuring soil properties. The other part is sieved through a 200-mesh (0.074 mm) nylon sieve to remove larger particles for measuring Vis-NIR and pXRF spectral curves;
[0017] Preferably, laboratory chemical analysis methods are used for soil heavy metals. Graphite furnace atomic absorption spectrometry is used for lead and cadmium; atomic fluorescence spectrometry is used for the determination of arsenic; the determination of copper is by flame atomic absorption spectrophotometry;
[0018] Heavy metal data of n soil samples will be obtained.
[0019] Further, Vis-NIR spectral measurement is carried out under controllable laboratory conditions using a Vis-NIR spectral radiometer. The measurement is carried out in a dark room. The instrument is preheated for 5 minutes, and a white board is used for reflectance correction. During measurement, the air-dried and sieved soil samples are placed in a petri dish with a diameter of 5 cm, and spectra are collected at n different points for each sample. During the spectral measurement process, white board correction is carried out multiple times to reduce the influence of noise. The obtained spectral data is imported into software for baseline correction, and the average value of n scan data is calculated. Since the noise is relatively large at both ends of the spectrum, cropping and resampling are carried out, and finally n spectral bands are obtained; The pXRF spectrometer measures the soil. Each sample is measured n times, each time not less than 30 seconds, and calibration is carried out once every 10 samples measured. Finally, the average value of n measurements is taken as the pXRF spectral data of the sample. To remove the low-energy band, the energy range is narrowed. Subsequently, to reduce redundancy and improve the modeling efficiency, the spectral data is resampled, and finally n spectral bands are obtained;
[0020] The curves of Vis-NIR and pXRF spectra of n soil samples are smoothed to obtain the original Vis-NIR and pXRF spectral curves.
[0021] Further, for the original Vis-NIR and pXRF spectral data, scattering correction and fractional-order differential preprocessing are carried out respectively;
[0022] Preferably, the preprocessing of scatter correction includes baseline correction (SUB) to eliminate spectral shifts caused by external factors; absorbance conversion (ABS) to convert reflectance to absorbance to reduce the influence of physical properties such as particle size on the spectrum; maximum reflectance correction (CMR) to achieve effective comparison between spectra through spectral maximum normalization; continuous normalization (CONR) to enhance the visualization effect of spectral data and improve the modeling accuracy; multiple scatter correction (MSC) to reduce the scattering effect in spectral data and improve the modeling stability; standard normal variate transformation (SNV) to standardize the data into a normal distribution with a mean of 0 and a standard deviation of 1 to reduce the influence of particle size effect on the spectrum;
[0023] Preferably, the baseline effects in spectral signals, such as vertical drift and tilt, can be corrected by differential operations. Among them, the first-order differential (FD) can eliminate baseline shift, and the second-order differential (SD) can remove both baseline shift and slope. The fractional-order differential (FOD) extends the traditional integer-order differential to any order and has significant advantages in signal detection and feature extraction;
[0024] Furthermore, the direct cascade method is used to fuse the preprocessed Vis-NIR and pXRF spectra;
[0025] 225 kinds of spectral data after fusion are obtained.
[0026] Furthermore, random forest can discover complex non-linear relationships, thus better extracting important frequency bands from soil spectra. The number of trees in the random forest model is set to 1000. To control overfitting, improve accuracy and robustness, 5 nodes are selected, and 1 / 3 of the samples are selected for feature selection. Modeling is carried out for 225 kinds of fused spectral data and soil heavy metal data;
[0027] Furthermore, ten-fold cross-validation is performed on the data to obtain the optimal fusion spectral model for each soil heavy metal;
[0028] Preferably, the indicators for evaluating a model include: coefficient of determination (R2), Lin's concordance correlation coefficient (LCCC), and ratio of interquartile range (RPIQ).
[0029] Furthermore, important spectral bands are extracted from the optimal fusion spectral model for each soil heavy metal. The variable importance in random forest modeling is normalized to the range of 0 - 1 for easy comparison;
[0030] Furthermore, after the machine learning model is established, it can be used to predict the heavy metal content in unknown soil samples. The steps include:
[0031] Obtain the VIS-NIR spectral data of the unknown sample;
[0032] Further, determine whether the sample matches the previously established machine learning model;
[0033] Further, if it matches, input the spectral data into the model to obtain the quantitative result of soil heavy metals in the unknown sample; if it does not match, only a reference result can be provided;
[0034] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:
[0035] Based on the fusion of Vis-NIR and pXRF spectra, combined with a variety of spectral preprocessing methods and the random forest machine learning algorithm, the present invention realizes high-precision, fast, and low-cost soil heavy metal prediction, and is particularly suitable for the precise monitoring of low-concentration heavy metals in arid agricultural soils. Through the fusion of 225 kinds of spectral data, the spectral characteristics are optimized, the model stability is enhanced, and the key spectral bands are extracted through variable importance analysis to reduce redundancy and improve the prediction performance. In addition, ten-fold cross-validation is adopted to ensure the high accuracy of the model, and it can be applied to the prediction of heavy metal content in unknown soil samples, providing an efficient and reliable technical means for sustainable agricultural development and ecological environment protection. Description of the Drawings
[0036] Figure 1 It is a schematic flow chart of the method for predicting heavy metals in farmland soil in arid areas based on the fusion of Vis-NIR and pXRF spectra provided by the present invention.
[0037] Figure 2 It is the original curve graph of soil Vis-NIR spectrum and pXRF spectrum provided by the embodiment of the present invention.
[0038] Figure 3 It is the soil Vis-NIR spectrum curve graph after 15 kinds of preprocessing provided by the embodiment of the present invention.
[0039] Figure 4 It is the soil pXRF spectrum curve graph after 15 kinds of preprocessing provided by the embodiment of the present invention.
[0040] Figure 5 It is the scatter plot of the predicted value and the measured value of the best preprocessing combination for soil heavy metal Cd based on the random forest machine learning method provided by the embodiment of the present invention.
[0041] Figure 6 It is the scatter plot of the predicted value and the measured value of the best preprocessing combination for soil heavy metal Pb based on the random forest machine learning method provided by the embodiment of the present invention.
[0042] Figure 7Scatter plot of the predicted values and measured values of the optimal pretreatment combination for soil heavy metal As based on the random forest machine learning method provided by the embodiments of the present invention.
[0043] Figure 8 Scatter plot of the predicted values and measured values of the optimal pretreatment combination for soil heavy metal Cu based on the random forest machine learning method provided by the embodiments of the present invention.
[0044] Figure 9 Comparison chart of important bands before and after pretreatment of four heavy metals provided by the embodiments of the present invention. Detailed implementation manners
[0045] Hereinafter, specific embodiments of the present invention will be described in detail with reference to the accompanying drawings. According to these detailed descriptions, those skilled in the art can clearly understand the present invention and can implement the present invention. Without departing from the principle of the present invention, the features in different embodiments can be combined to obtain new implementation manners, or some features in certain embodiments can be replaced to obtain other preferred implementation manners. The embodiments of the present invention provide a method for high-precision prediction of heavy metals in arid area farmland soil based on Vis-NIR and pXRF spectral fusion. Figure 1 is the overall flowchart of the method. As Figure 1 shown, the method includes the following steps.
[0046] Step 1: Obtain arid area farmland soil sample data. In the arid area farmland area, select multiple sampling points according to the uniform sampling method or specific sampling strategy. Use a handheld GPS to record the longitude and latitude coordinates of the sampling points to ensure the spatial representativeness of the samples. The collected soil samples need to be air-dried at room temperature of 25°C, ground after removing gravel and animal and plant residues, and sieved using nylon sieves of different specifications (100 mesh and 200 mesh) to ensure the particle size for different experimental requirements.
[0047] Step 2: Detect heavy metals in the soil sample data to obtain the content information of four heavy metals, cadmium, lead, arsenic, and copper; and obtain the spectral data of the soil samples through spectral measurement to obtain the original Vis-NIR and pXRF spectral data, as Figure 2 .
[0048] Step 3: To improve the quality of spectral data, 15 preprocessing methods were applied to Vis-NIR and pXRF spectra respectively, including: scattering correction (baseline correction, absorbance conversion, maximum reflectance correction, continuum normalization, multiplicative scatter correction, standard normal variate transformation, etc.); fractional order derivative processing, and the Grünwald-Letnikov (G-L) method was used to calculate the fractional order derivatives of different orders to enhance characteristic information and remove noise. After preprocessing, the effects of different methods on the spectral signal were compared to obtain the optimal spectral preprocessing effect, and the visualization comparison results were plotted, as shown in Figures 3 and 4.
[0049] Step 4: The direct cascade method was used to combine the 15 preprocessed Vis-NIR spectral data and 15 pXRF spectral data, generating a total of 225 fused spectral datasets. This data fusion method fully utilized the advantages of Vis-NIR for soil organic matter and mineral information, as well as the ability of pXRF to directly detect heavy metal elements, ensuring the comprehensiveness of spectral data and minimizing information redundancy.
[0050] Step 5: Using the 225 fused spectral data and the measured soil heavy metal data, a random forest algorithm was used to build a model. 1000 decision trees were set to improve the model stability; a 1 / 3 variable selection method for samples was adopted to reduce the model complexity and prevent overfitting; 10-fold cross-validation was performed to evaluate the model performance and ensure the reliability of the prediction results.
[0051] Step 6: Model evaluation and selection of the optimal preprocessing scheme Different fused spectral models were evaluated through indicators such as the coefficient of determination, Lin's concordance correlation coefficient, and interquartile range ratio, and the optimal spectral preprocessing combination for each heavy metal was screened out, and then the accuracy comparison results were obtained.
[0052] Step 7: Based on the selected optimal preprocessing combination, the most important spectral variables were extracted from the fused spectral data. In the random forest model, using the variable importance analysis method, the variable weights were normalized to between 0 and 1, and the key spectral characteristic bands were screened out to reduce data redundancy and improve the model generalization ability, such as Figure 9 .
[0053] In this example, based on the prediction method of heavy metals in arid area farmland soil based on the fusion of Vis-NIR and pXRF spectra, the comparison diagrams of the predicted content and the actual content for the quantitative analysis of heavy metals cadmium, lead, arsenic, and copper in soil are shown in Figure 5 , 6 , 7, and 8. As shown in Figure 5 , 6, as can be seen from 7 and 8, the predicted Lin's concordance coefficients of the four soil heavy metals are all above 0.6, and the interquartile range ratios are all above 2.5, indicating that the method has high accuracy for the quantitative analysis of soil heavy metal content in farmland in arid areas.
[0054] As described above, the above are only specific embodiments of the present invention. Any feature disclosed in this specification, unless specifically described, can be replaced by other equivalent or alternative features with similar purposes; all the disclosed features, or all the steps in all the methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.
Claims
1. A method for predicting heavy metals in farmland soil in arid areas based on Vis-NIR and pXRF spectrum fusion, characterized in that: The following steps are involved: Step 1: Collect n samples of farmland soil in arid areas and measure the heavy metal content to obtain measured data of soil heavy metals; Step 2: Perform Vis-NIR spectrum and pXRF spectrum measurements on n soil samples to obtain original spectrum data; Step 3: 15 spectral preprocessing methods were applied to the raw Vis-NIR and pXRF spectral data to optimize the signal quality; Step 4: Directly fuse the Vis-NIR and pXRF spectral data after 15 preprocessing processes to construct 225 fused spectral data sets; Step 5: Based on the fused spectral data and the measured data of soil heavy metals, a random forest model is used for modeling; Step 6: Evaluate the model performance and select the optimal spectral preprocessing combination for each heavy metal; Step 7: Based on the optimal spectral preprocessing combination, extract key spectral features to improve the accuracy and stability of heavy metal prediction.
2. The method according to claim 1, characterized in that The 15 spectral preprocessing methods used in step 3 include: baseline correction (SUB), absorbance conversion (ABS), maximum reflectance correction (CMR), continuous normalization (CONR), multiple scattering correction (MSC), standard normal variate transformation (SNV), 0th derivative, 0.25th derivative, 0.5th derivative, 0.75th derivative, 1st derivative, 1.25th derivative, 1.5th derivative, 1.75th derivative, and 2nd derivative.
3. The method according to claim 1, characterized in that In step 4, a direct cascade fusion strategy is used for the 15 preprocessed Vis-NIR and pXRF spectral data to form a multi-dimensional spectral dataset.
4. The method according to claim 1, characterized in that In step 5, the random forest algorithm is used to model the fused spectral data and soil heavy metal content data to improve the robustness and generalization ability of the prediction.
5. The method according to claim 1, characterized in that In step 7, based on the optimal spectral preprocessing combination, feature importance analysis is used to extract key spectral features to optimize the recognition ability of heavy metal elements and improve the accuracy of the detection model.