A multi-feature and integrated optimization method for predicting soil salinity conductivity
By combining the time-domain and frequency-domain characteristics and correlation coefficients of radar echoes, an ensemble learning algorithm is used to construct a regression model, which solves the problem of insufficient accuracy in measuring the electrical conductivity of saline-alkali land in existing technologies. This achieves high-precision and efficient prediction of soil electrical conductivity, and is suitable for complex geological environments and large-scale applications.
Patent Information
- Application Number
- CN202511031768.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing ground-penetrating radar-based methods for measuring soil conductivity rely on single-feature inputs and fail to effectively integrate multi-scale features and depth information. This results in insufficient prediction accuracy of the model in complex geological environments and an inability to effectively handle data imbalance, thus limiting the high-precision measurement of conductivity in saline-alkali land.
A multi-feature and ensemble optimization approach is adopted, combining the time-domain and frequency-domain features and correlation coefficients of radar echoes. A regression model is constructed through an ensemble learning algorithm, a preferred machine learning model is selected, weight coefficients are determined in different conductivity ranges, features are optimized using the MRMR algorithm, and a partitioned weighted ensemble learning strategy is used for prediction.
It significantly improves the accuracy of electrical conductivity prediction, increases the coefficient of determination to over 0.90, enhances the model's adaptability and generalization ability, reduces prediction bias, and is suitable for large-scale saline-alkali land surveys and agricultural soil monitoring, thereby improving measurement efficiency and result reliability.
Smart Images

Figure CN120524463B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of saline-alkali soil conductivity prediction, and particularly relates to a multi-feature and integrated optimization saline-alkali soil conductivity prediction method. BACKGROUND
[0002] In the field of saline-alkali soil research and management, accurate measurement of soil conductivity is crucial for assessing the degree of soil salinization and developing scientific and reasonable improvement strategies. Changes in soil conductivity reflect the ion concentration, water content, and physical and chemical properties of the soil. Therefore, obtaining high-precision conductivity data not only helps to optimize farmland water management, but also provides a scientific basis for saline-alkali soil management and ecological restoration.
[0003] With the development of technology, various soil conductivity measurement techniques have been proposed and applied, including laboratory chemical analysis, electromagnetic induction (EMI), and inversion methods based on ground penetrating radar (GPR). Among them, GPR has the advantages of non-contact measurement, rapid detection, and large-area coverage, and has become an important direction in recent years for soil conductivity research.
[0004] Existing GPR-based soil conductivity measurement methods have single feature input and rely only on echo correlation coefficients as prediction variables, resulting in limited model expression ability and difficulty in accurately depicting the spatial distribution and variation of soil conductivity. Due to limited feature selection, multi-scale features and depth information cannot be effectively integrated, making it difficult to adapt to changes in complex geological environments.
[0005] Most existing studies use polynomial fitting or traditional linear regression models. These methods are computationally efficient for low-dimensional data, but they are prone to overfitting or lack of generalization ability when dealing with high-dimensional, nonlinear features. In particular, in cases where the saline-alkali soil conductivity is high and the geological conditions are complex, the model error is large and the prediction stability is poor. Some studies use a single machine learning model for conductivity inversion, which fails to fully utilize the advantages of ensemble learning and cannot balance the adaptability of different data features, resulting in limited prediction accuracy. In addition, some methods do not adequately consider the data imbalance problem during data processing, resulting in prediction accuracy deviations in high or low conductivity areas, limiting their application in large-scale, high-precision soil conductivity measurement. SUMMARY
[0006] The purpose of the present invention is to overcome the defects of the existing technology and provide a multi-feature and integrated optimization method for predicting the conductivity of saline-alkali land. According to the characteristics of saline-alkali land, the method utilizes the echo correlation coefficient and combines the time domain and frequency domain characteristics of the echo. Through an integrated learning algorithm and fusing multiple features to construct a regression model, the correlation between GPR signals and soil conductivity is characterized from multiple angles, which can improve the prediction accuracy of conductivity.
[0007] The technical solution provided by the present invention is:
[0008] A multi-feature and integrated optimization method for predicting electrical conductivity of saline-alkali land includes the following steps:
[0009] Step 1: obtaining radar echo characteristic parameters of multiple sampling locations in the saline-alkali land and the conductivity corresponding to each radar echo characteristic parameter to form a sample set;
[0010] The radar echo characteristic parameters include: radar echo correlation coefficient characteristics, radar echo spectrum characteristics, radar echo time domain characteristics, amplitude difference between adjacent frequency points of radar echo signal, difference between adjacent positive peaks of radar echo signal and difference between adjacent negative peaks of radar echo signal;
[0011] Step 2: Using the sample set to train different machine learning models, and based on the training results, screening out multiple preferred machine learning models as conductivity prediction models;
[0012] Step 3: determining a weight coefficient of each conductivity prediction model in each conductivity interval according to the accuracy of the prediction results of the conductivity prediction model for samples in different conductivity intervals;
[0013] Step 4: Detecting the sampling position to be measured in the saline-alkali land by radar, obtaining radar echo characteristic parameters of the sampling position to be measured, and inputting the radar echo characteristic parameters of the sampling position to be measured into multiple conductivity prediction models to obtain multiple conductivity prediction values; and determining the conductivity interval to which the sampling point to be measured belongs based on the multiple conductivity prediction values;
[0014] Step 5: Calculate the conductivity prediction value of the sampling location to be measured:
[0015] ;
[0016] in, For the The weight coefficient of the conductivity prediction model in the conductivity interval to which the sampling position to be measured belongs, For the The conductivity prediction model predicts the conductivity value of the sampling location to be measured. U is the number of conductivity prediction models.
[0017] Preferably, the radar echo correlation coefficient feature is characterized by:
[0018] ;
[0019] wherein, and are normalization coefficients, is a standard radar wavelet, represents the radar echo at any sampling position in the saline-alkali land; represents time, represents the length of the selected time window.
[0020] Preferably, in the step one, the method for obtaining the radar echo frequency spectrum feature is characterized by:
[0021] calculating the sampling frequency step:
[0022] ;
[0023] performing Fourier transform on the sampling sequence to obtain the sampling frequency at each sampling position:
[0024] ;
[0025] determining the frequency range of the sampling sequence , and taking the integer characteristic frequency within the frequency range as the radar echo frequency spectrum feature;
[0026] wherein, , ;
[0027] wherein, is the index of the frequency component, is the base number of the natural logarithm, is the imaginary unit, and are the indices of the lower limit frequency and the upper limit frequency of the working frequency band, respectively; represents the sampling sequence, represents the number of the sampling sequence; and represent the lower limit and the upper limit of the working frequency band, respectively.
[0028] Preferably, the radar echo time domain feature is the time sequence of multiple peaks of the radar echo signal.
[0029] Preferably, before the step two, it further comprises:
[0030] The MRMR algorithm is used to screen out key feature frequencies and key wave peaks from the radar echo frequency spectrum features and the radar echo time domain features respectively, and the values of the key feature frequencies and the peak values of the key wave peaks are fused respectively to obtain fused frequency spectrum features and fused time domain features.
[0031] The radar echo correlation coefficient features, the fused frequency spectrum features, the fused time domain features, the amplitude difference between adjacent frequency points of the radar echo signal, the difference between adjacent positive peaks of the radar echo signal, and the difference between adjacent negative peaks of the radar echo signal are used as the radar echo feature parameters in the sample set.
[0032] Preferably, the calculation formula of the fused frequency spectrum features is:
[0033] ;
[0034] wherein, ;
[0035] wherein, represents the fused frequency spectrum features, represents the fusion coefficient of the key feature frequency , represents the actual value of the key feature frequency , represents the proportion of the actual value of the key feature frequency , represents the proportion of the MRMR score value of the key feature frequency ; is the number of the key feature frequencies.
[0036] Preferably, the calculation formula of the fused time domain features is:
[0037] ;
[0038] wherein, ;
[0039] wherein, represents the fused time domain features, represents the fusion coefficient of the key peak value , represents the peak value of the key wave peak , represents the proportion of the peak value of the key wave peak , represents the proportion of the MRMR score value of the key wave peak ; is the number of the key wave peaks.
[0040] Preferably, the preferred machine learning model comprises: an ensemble of boosted trees model, an ensemble of bagged trees model, a Gaussian process regression model and a linear support vector machine model.
[0041] Preferably, the calculation formula of the weight coefficient of each conductivity prediction model in each conductivity interval is:
[0042] ;
[0043] wherein, U is the number of conductivity prediction models, is the decision coefficient of the i-th conductivity prediction model in the corresponding conductivity interval. i
[0044] Preferably, in the step three, the conductivity is divided into three intervals of low conductivity, medium conductivity and high conductivity;
[0045] wherein, the conductivity interval corresponding to the low conductivity is: [0, 750] uS / cm, the conductivity interval corresponding to the medium conductivity is (750, 2000] uS / cm, and the conductivity interval corresponding to the high conductivity is: (2000, 7000] uS / cm.
[0046] The beneficial effects of the present application are:
[0047] (1) The prediction accuracy is significantly improved: the present application combines the time domain, frequency domain and echo correlation coefficient features, and optimizes the machine learning model, so that the decision coefficient (R 2 ) of the conductivity prediction is improved to 0.90 or more, which is improved by more than 15% compared with the traditional single feature input method.
[0048] (2) The model has stronger adaptability and better generalization ability: through testing and verification in different saline-alkali soil environments (low humidity and high humidity areas), it is proved that the method of the present application has better adaptability; in the low conductivity area and the high conductivity area, the prediction accuracy is better than that of the existing method, so that it is applicable to a wider range of application scenarios.
[0049] (3) Optimizing data distribution and improving the universality of the model: the present application adopts a partition weighted ensemble learning strategy in the data processing process, which reduces the prediction deviation of the traditional model in the high conductivity or low conductivity area, so that the conductivity prediction remains balanced in the whole range, and the reliability of the result is improved; in addition, the MRMR algorithm is used to effectively screen the key features, reduce the redundant information, and make the prediction result more stable.
[0050] (4) Improve measurement efficiency, suitable for large-scale applications: The application combines GPR fast data acquisition and machine learning online estimation, which improves the measurement efficiency of soil conductivity by more than 3 times, significantly reduces the data processing time compared with traditional methods, and is suitable for large-scale saline-alkali land investigation, agricultural soil monitoring, irrigation optimization and ecological environment assessment, providing scientific basis for precision agriculture and land improvement.
[0051] (5) Promote saline-alkali land management and agricultural modernization development: High-precision soil conductivity data can provide precise fertilization, irrigation and soil improvement decision support for agricultural managers, help to improve agricultural production efficiency, and promote the sustainable development of saline-alkali land improvement and ecological environment. BRIEF DESCRIPTION OF DRAWINGS
[0052] Figure 1 The discrete sequence diagram of the echo of each sampling position extracted in time domain described in the application.
[0053] Figure 2 The predicted response result diagram obtained after training by using the integrated bagging tree model in the embodiment of the application.
[0054] Figure 3 The predicted response result diagram obtained after training by using the integrated bagging tree model in the embodiment of the application.
[0055] Figure 4 The predicted response result diagram obtained after training by using the integrated bagging tree model in the embodiment of the application.
[0056] Figure 5 The predicted response result diagram obtained after training by using the integrated bagging tree model in the embodiment of the application.
[0057] Figure 6 The predicted response result diagram obtained after training by using the integrated bagging tree model in the embodiment of the application.
[0058] Figure 7 The predicted response result diagram obtained after training by using the integrated bagging tree model in the embodiment of the application.
[0059] Figure 8 The predicted response result diagram obtained after training by using the integrated bagging tree model in the embodiment of the application.
[0060] Figure 9 The predicted response result diagram obtained after training by using the integrated bagging tree model in the embodiment of the application.
[0061] Figure 10 The predicted response result diagram obtained after training by using the integrated bagging tree model in the embodiment of the application.
[0062] Figure 11The prediction effect diagram of the partition adaptive weighted ensemble learning method of the application. DETAILED DESCRIPTION
[0063] The application will be further described in detail below with reference to the accompanying drawings so as to enable those skilled in the art to implement the application according to the description and the drawings.
[0064] The application provides a saline-alkali soil conductivity prediction method with multiple features and integrated optimization, and the specific implementation process is as follows.
[0065] 1. Data acquisition
[0066] In this data acquisition, a separated transmitting-receiving pulse GPR system is used to measure the saline-alkali soil, and the system works in the near-field region. In order to ensure sufficient signal coupling, the antenna height is set to be about 5 cm away from the ground, and the center frequency of the transmitted pulse signal is 100 MHz. At the same time, a WET sensor is introduced to obtain the reference data of soil conductivity.
[0067] 2. Feature extraction
[0068] Based on the propagation and response characteristics of electromagnetic waves in high-conductivity saline-alkali soil, three types of input parameters are selected, namely echo time domain (dimension expansion) features, frequency domain (dimension expansion) features and echo correlation features. The time domain (dimension expansion) features of the echo mainly reflect the instantaneous change of the radar echo signal in time, and the peak amplitude and position can reveal the reflection intensity and depth information of the underground medium interface; the frequency domain (dimension expansion) features depict the energy distribution of the radar signal at different frequency components, which can reflect the absorption and attenuation characteristics of the soil to different frequency signals, thereby indirectly reflecting the conductivity; and the echo correlation coefficient is used to measure the similarity between the measured echo and the standard waveform, which is a comprehensive reflection of the features in the time domain and the frequency domain of the echo, and reflects the trend of the overall shape of the signal changing with the conductivity in a new dimension. The fusion of the three types of features as model input can comprehensively capture the nonlinear relationship between the radar signal and the soil conductivity from multiple dimensions, and enhance the perception ability of the model to complex soil environment. At the same time, in the soil with weak dielectric constant, using the three types of features to describe the collected echo signal helps to retain the details of the original signal. Thus, the prediction accuracy, generalization ability and anti-interference performance can be effectively improved. The following will be described in detail:
[0069] (1) Echo correlation coefficient feature
[0070] The radar echo correlation coefficient is an important indicator for measuring the relationship between the change of GPR echo signal and soil conductivity. The calculation method is as follows:
[0071] Let the extracted standard radar wavelet be Let the radar echo at any sampling position in the saline-alkali soil be , the radar echo correlation coefficient is defined as follows:
[0072] ;
[0073] wherein, and are normalization coefficients, is the correlation coefficient determined by the GPR echo and , denotes a time variable, the argument of the integral function and , is the upper limit of integration, representing the length of the selected time window. As , the expression tends to the long-term correlation of the two signals and , that is, reflects the degree of similarity of the entire time range. This feature can effectively characterize the similarity of the radar echo and the standard wavelet, and thus reflect the changes of soil conductivity.
[0074] (2) Echo spectrum feature (single frequency amplitude)
[0075] Since the spectral characteristics of the radar echo signal are significantly affected by soil conductivity, the specific frequency component is extracted and used as a feature input. Its calculation method is as follows:
[0076] Let the number of sampling points of the radar echo at each sampling position be N , and the sampling interval time be , then the sampling frequency can be defined as:
[0077] ;
[0078] Then the discrete Fourier transform (DFT) frequency step is:
[0079]
[0080] The DFT is performed on the sampling sequence, and let the sampling sequence be , then the DFT formula can be expressed as:
[0081] ;
[0082] wherein, is the index of the frequency component, is the base of the natural logarithm, is the imaginary unit.
[0083] Determine the working frequency band of the GPR system, which is defined as from to Then, there are
[0084]
[0085]
[0086] where, and are the lower and upper limits of the GPR working frequency band, respectively, and are the indexes of the lower and upper limits of the working frequency band, respectively, since the working frequency band of GPR is to , the obtained after using the discrete Fourier transform, the value of is selected in the range of k , the minimum is , and the maximum is , so the frequency range of the sampling sequence is ultimately . Select an integer characteristic frequency in this range as the model input.
[0087] (3) Echo time-domain features (peak information)
[0088] In the time domain, the peak information of the echo signal can reflect the energy distribution of the GPR wave and has important physical significance. The method of extracting the peak value is as follows:
[0089] Let the ground penetrating radar (GPR) echo signal be a discrete sequence:
[0090]
[0091] where, is the signal amplitude of the m th sampling point, N is the total number of sampling points.
[0092] Assume that V peak values (including positive and negative peaks) need to be extracted from each echo signal, let be the preset peak amplitude threshold, the positive peak is greater than the threshold, and the negative peak is less than the threshold, define the peak set:
[0093]
[0094] where, represents the amplitude of the th peak, represents the index corresponding to the th peak, and needs to satisfy the time sequence constraint of .
[0095] Set initial index , i.e. starting from the first data point, the iteration steps to find the th peak value are as follows:
[0096] 1. Search for the next peak value
[0097] If searching for a positive peak (local maximum), find the minimum such that:
[0098] If searching for a negative peak (local minimum), find the minimum such that: ; where represents the th sampling point, i.e. the next point adjacent to ;
[0099] 2. Record the peak value:
[0100] ;
[0101] At this time, is the positive (or negative) peak value extracted at .
[0102] 3. Update the search starting point to ensure that the subsequent peak index is increasing in time:
[0103]
[0104] After V iterations, the peak value set of each echo signal P , Figure 1 is obtained, where is the discrete sequence diagram of the echo at each sampling position after time domain extraction, the horizontal axis represents time (ns), and the vertical axis represents amplitude. Each peak value (positive and negative) in the diagram is represented by .
[0105] In summary, various effective features of radar echo in the time-frequency domain are deeply excavated, and the correlation information between radar echo and soil conductivity is captured from different angles, providing rich and representative input information for machine learning models.
[0106] (4) Feature expansion
[0107] A differential enhancement feature construction method is adopted to expand the input data dimension and enhance the sensitivity of the model to the microscopic changes of soil conductivity.
[0108] Specifically, in the frequency domain feature, the single frequency amplitude on the frequency domain reflects the response ability of the soil electrical structure to the radar signal at the corresponding frequency. In order to further depict the response gradient change between different frequency components, the present application calculates the amplitude difference between adjacent frequency points (such as: ), thereby reflecting the frequency domain slope characteristics of the radar echo signal in the frequency spectrum change process. This difference can represent the response difference characteristics of the dielectric parameters or conductive ions in the soil to different frequencies of electromagnetic waves.
[0109] In the time domain feature, the extracted multiple wave peaks reflect the peak response of the radar signal after reflection and scattering with the underground structure on the time axis. These wave peaks correspond to the interface reflection characteristics of different depths or different physical structures in the soil. In order to depict the gradient change of the local structure in the time domain, the present application selects multiple representative wave peak pairs to perform difference construction, such as the first wave peak value minus the third wave peak value, the second wave peak value minus the fourth wave peak value, and the like, thereby constructing the morphological change rate characteristics in the time domain. These differences not only reflect the energy attenuation trend of the radar wave in the underground medium during propagation, but also effectively represent the coherence characteristics of the near-field echo in the saline-alkali soil surface detection.
[0110] By introducing the above difference features, the present application effectively expands the expression ability of the feature space, so that the potential but not explicitly modeled structural information in the original features can be extracted, thereby improving the resolution and fitting accuracy of the machine learning model for predicting soil conductivity.
[0111] The time domain feature of the echo mainly reflects the instantaneous change of the radar echo signal in time, and the peak amplitude and position can reveal the reflection intensity and depth information of the underground medium interface; the frequency domain feature depicts the energy distribution of the radar signal at different frequency components, which can reflect the absorption and attenuation characteristics of the soil to different frequency signals, thereby indirectly reflecting the conductivity; and the echo correlation coefficient feature is used to measure the similarity between the measured echo and the standard waveform, which is a comprehensive reflection of the signal overall form in the new dimension with the change trend of the conductivity.
[0112] Fusing the three types of features as model inputs can comprehensively capture the nonlinear relationship between the radar signal and the soil conductivity from multiple dimensions, and enhance the perception ability of the model to complex soil environment. At the same time, in the soil with weak dielectric constant, using the three types of features to depict the collected echo signal helps to retain the details of the original signal. Thus, the prediction accuracy, generalization ability and anti-interference performance can be effectively improved.
[0113] In summary, the present application deeply excavates the multiple effective features of the radar echo in the time-frequency domain, captures the correlation information between the radar echo and the soil conductivity from different angles, and provides rich and representative input information for the machine learning model.
[0114] 3. Model selection and optimization
[0115] The present application constructs a soil conductivity prediction model by a machine learning method, mainly including data preprocessing, model pre-training and comparison, optimal model selection, and prediction of conductivity, etc. At the same time, in order to further improve the prediction accuracy, and give full play to the advantages of each model in different data subsets or different soil conductivity intervals, an ensemble learning method is proposed. The basic idea is: in the training stage, the input features are predicted by multiple regression models, and the performance is evaluated by cross-validation method, etc. On this basis, for different soil conductivity intervals (such as low, medium and high conductivity regions), the advantages of each model in these intervals are utilized, and the prediction results of each model are fused by partition weighted integration method, so as to realize the improvement of the overall prediction performance.
[0116] (1) Data preprocessing: first, the processed feature variables are divided into training set, validation set and test set according to a certain proportion, and all feature data are normalized to ensure the stability of model training. The best combination of time domain features (wave peak) and frequency domain features (single frequency amplitude) is set, that is, from the extracted wave echo V M X
[0117] (2) Model pre-training and comparison: using the selected combination of time domain features (wave peak) and frequency domain features (single frequency amplitude), train multiple models including the following regression models, such as polynomial regression, support vector machine regression, Gaussian process regression, random forest regression, ensemble bagging tree, etc., and calculate the R² value of each model under different feature combinations. The S fold cross-validation method is used to prevent model overfitting, and ensure the robustness and generalization ability of the model.
[0118] (3) Feature optimization and fusion: based on MRMR (minimum redundancy maximum relevance) algorithm and pre-training results, the MRMR importance score of each time domain and frequency domain feature is obtained. In order to further improve the model's ability to identify key features of soil conductivity, a weighted calculation strategy of feature amplitude ratio and MRMR importance score ratio fusion is adopted to construct the representative time domain and frequency domain input features after fusion.
[0119] Specifically, in terms of time domain features, the multiple key peak values extracted from each ground penetrating radar echo signal are normalized, and the peak amplitude is divided by the sum of the peak amplitudes to obtain the actual response proportion of each peak. This proportion reflects the energy proportion of each peak in the entire radar echo and represents the intensity of the reflection of underground electrical changes. At the same time, the importance score corresponding to each peak obtained by the MRMR (minimum redundancy maximum relevance) algorithm represents the statistical contribution of each peak to the prediction of the target variable (soil conductivity) in the entire sample. After normalizing these scores, the global statistical importance proportion of each peak is obtained, which can be understood as the discriminant ability weight of the peak to the conductivity change in the entire data set. The average of the above two percentages is used to construct the fusion weighting coefficient of each peak. Then, the corresponding peaks in each echo signal are weighted and summed using the fusion coefficient to obtain a time domain fusion feature value, which is used as one of the final input features of the model. This value takes into account the local response intensity (physical strength) of the single signal and the discriminant ability of the global feature (statistical weight), and has high representativeness.
[0120] In the frequency domain feature aspect, the same method is used to calculate the relative single frequency amplitude proportion of the selected multiple single frequency amplitudes, and then fuse it with the corresponding MRMR importance proportion to obtain the weighting coefficient of each frequency point, thereby constructing the fusion feature value in the frequency domain.
[0121] The specific calculation method of the fusion feature value is as follows:
[0122] Suppose after pre-training, the key peaks extracted from the ground penetrating radar echo signal are , that is, the key peaks are extracted S , their actual values are represented by , and the MRMR score values are represented by , then the proportion of the actual value of each key peak is represented by , i =1,…, S , which is calculated by the following formula:
[0123]
[0124] The proportion of the MRMR score value of each key peak is represented by , i =1,…, S , which is calculated by the following formula:
[0125]
[0126] Then, the average of the two proportions is taken to obtain the fusion coefficient
[0127]
[0128] Thus, the fusion coefficients of different key peaks are obtained , and then the actual values of the key peaks corresponding to the sampling echoes of all sampling positions are multiplied by the corresponding fusion coefficients respectively to perform weighted summation, and the sum value is the time domain fusion feature corresponding to all sampling echoes
[0129] .
[0130] Similarly, assuming that after pre-training, the key single-frequency amplitudes extracted from the ground penetrating radar echo signal are , that is, the key frequencies are extracted Q , their actual values are represented by , and the MRMR score values are represented by , then the proportion of the actual value of each key frequency is represented by , =1,…, Q , which is calculated by the following formula:
[0131]
[0132] The proportion of the MRMR score value of each key frequency is represented by , =1,…, Q , which is calculated by the following formula:
[0133]
[0134] Next, the two proportions are averaged to obtain the fusion coefficient
[0135]
[0136] Thus, the fusion coefficients of different key single-frequency amplitudes are obtained , and then the actual values of the key single-frequency amplitudes corresponding to the sampling echoes of all sampling positions are multiplied by the corresponding fusion coefficients respectively to perform weighted summation, and the sum value is the frequency domain fusion feature corresponding to all sampling echoes :
[0137] .
[0138] Through the feature fusion strategy, the generalization performance and anti-interference ability of the model are further improved, and it is suitable for more complex or subtle soil environments with more significant differences.
[0139] (4) Select the optimal model: the optimized fusion feature data and the difference feature obtained by feature expansion are combined with the echo correlation coefficient feature, and then input into different machine learning regression models for training, and multiple rounds of training and verification are performed, the prediction accuracy of different models is compared, and the model with the best prediction performance is selected.
[0140] (5) Ensemble learning: in order to further improve the prediction accuracy and overall robustness of the model in different soil conductivity intervals, the application introduces an adaptive weighting mechanism based on model performance indicators in the ensemble learning strategy, which dynamically evaluates the prediction ability of each regression model in different conductivity ranges, realizes intelligent regulation and precision improvement of the fusion model.
[0141] Specifically, based on the knowledge of geophysics and agricultural ecology, it is known that the soil conductivity of saline-alkali land has significant non-uniformity in actual distribution, which can be divided into low [0, T 1 ] medium ( T 1 , T 2 ] high ( T 2 , T 3 ] conductivity three intervals. The application first divides the overall training set according to the above intervals, and trains and verifies a plurality of candidate regression models (such as ensemble boosting tree, ensemble bagging tree, Gaussian process regression, linear support vector machine regression, etc.) in each sub-interval, records the prediction results and performance indicators, especially the coefficient of determination (R²) value.
[0142] At this time, in order to ensure that the integrated weighting strategy calls the correct set of weighting coefficients (i.e. uses low interval, medium interval or high interval weighting strategy), a robust judgment mechanism needs to be developed. The mechanism is as follows:
[0143] Let the average value of the prediction values of the U selected models be: =1,…,U
[0144]
[0145] The following judgment criteria are defined:
[0146] If , it is judged that the sampling point belongs to the low conductivity interval, if , it is judged that the sampling point belongs to the medium conductivity interval, and if , it is judged that the sampling point belongs to the low conductivity interval.
[0147] Subsequently, the present invention uses the coefficient of determination R² of the model as the basis for performance weighting, and obtains the relative weight coefficient by normalizing the R² value of each model within the conductivity range. The specific calculation method is:
[0148] ;
[0149] in, For the The weighted coefficients of the model, For the The determination coefficient of each model in the corresponding conductivity range is, U is the number of weighted models. In the actual prediction stage, for any new test sample, the system first inputs it into each trained model to obtain a number of initial prediction values. Based on the average of this set of prediction values, it determines whether the sample prediction result belongs to the low, medium, or high concentration range. Then, the weight coefficient table within the corresponding range is automatically called, and the output results of each model are integrated in a weighted average manner to finally obtain a comprehensive prediction value:
[0150] ;
[0151] in, For the The predicted value of the model for this sample is is the final weighted fusion result.
[0152] This method not only adjusts the fusion strategy according to the historical performance of the model in each interval, but also introduces incremental samples during the system operation and updates the performance evaluation results of each interval at irregular intervals, thereby dynamically refreshing the weighting coefficients and realizing continuous adaptive optimization of the integrated model.
[0153] (6) Predicting electrical conductivity: Finally, the input feature data after optimization and integration processing is input into the model with the best performance or after integration fusion. The model outputs the predicted value of soil electrical conductivity based on the relationship between the learned features and soil electrical conductivity, thereby achieving a high-precision estimation of the electrical conductivity of saline-alkali land.
[0154] The implementation process and technical effects of the present invention are further described below in conjunction with specific embodiments to further demonstrate the effectiveness and superiority of the present invention.
[0155] Example
[0156] The representative saline-alkali region, the Yellow River big bend area (40.755572ºN, 107.736707ºE), is selected as the experimental site. A 175-meter-long measuring line is selected in the area, and the soil conductivity in the area presents a gradient change, which is typical. 16 areas are selected along the measuring line for random sampling, and soil samples are collected at each sampling location to determine the water content and ion concentration, which will constitute the basic data set for the inversion of the conductivity.
[0157] In the feature extraction aspect, the time-frequency domain features are extracted from the echo data collected by GPR, and a feature data set is established. First, the GPR echo correlation coefficient is calculated, which is used to represent the similarity of the radar signal and is one of the input variables of the machine learning model. Thirteen different single frequency amplitudes are extracted by DFT (Discrete Fourier Transform), which can be used to analyze the influence of soil conductivity on different frequency components. Considering that the coupling between the echo and the saline-alkali soil is more prominent in the early signal, the first 8 peaks of the echo P m ( m =1,…,8) are selected as the echo time domain features. These peak information contains rich soil property information and is of great value to the estimation of soil conductivity.
[0158] Next, the features are expanded. In the time domain, the 8 peaks collected are calculated according to , and 6 new time domain expanded features are obtained. In the frequency domain, the 13 single frequency amplitude values collected are subtracted between adjacent values, and 12 new frequency domain expanded features are obtained.
[0159] In the model pre-training, all feature variables are pre-processed, and the data set is divided for training, validation and testing. Then, the selected 8 peaks and 13 single frequency amplitude values are combined with the echo correlation coefficient for model training, and S fold cross-validation is used to prevent overfitting, and the prediction results are compared. Then, based on the MRMR algorithm, the prediction features with the highest correlation with soil conductivity and the lowest redundancy between each other are selected. Based on the prediction accuracy of different models, the first, second, third, fourth and fifth wave peaks in the time domain are selected as the key wave peaks, and the single frequency amplitudes selected in the frequency spectrum are 72 MHz, 136 MHz, 200 MHz, 216 MHz and 296 MHz as the key frequencies, and their MRMR importance scores are obtained. Then, the actual value proportion (percentage) of the 5 time domain features and frequency domain features of the standard radar wavelet is calculated.
[0160] After the above steps, at this time, two percentages of 5 time domain characteristics and 5 frequency domain characteristics of the standard radar wave are obtained, and then the two percentages are averaged to obtain the final feature fusion percentage, that is, the feature fusion coefficient. The specific data is shown in Table 1-2. Then, the five groups of peak values and the five groups of amplitude values of each echo signal are weighted by the corresponding feature fusion coefficients, and finally the fused time domain features and fused frequency domain features are obtained.
[0161] Table 1 Key Peak Fusion Coefficient
[0162]
[0163] Table 2 Key Feature Frequency Fusion Coefficient
[0164]
[0165] Finally, the echo correlation coefficient feature, 6 groups of time domain extended dimension features and 12 groups of frequency domain extended dimension features, fused time domain features and fused frequency domain features are put into all regression models for retraining to perform final optimization of the model.
[0166] In the model training and optimization, first, all feature variables are preprocessed, and data sets are divided for training, verification and testing. Then the above features are input into the integrated bagging tree model, the integrated boosting tree model, the Gaussian process regression model, the linear support vector machine model, the medium regression tree model and the linear regression model for training. S-fold cross-validation is used to prevent overfitting, the prediction results are compared, and the R 2 ≥0.85 prediction model is selected as the best regression model.
[0167] The four conductivity prediction models selected in this embodiment are: the integrated boosting tree model (R 2 =0.88), the integrated bagging tree model (R 2 =0.87), the Gaussian process regression model (R 2 =0.85), and the linear support vector machine (SVM) model (R 2 =0.85). The prediction accuracy of these four models for soil conductivity is very high, among which the integrated boosting tree model shows the highest prediction accuracy. Figures 2 to 5 The response graphs of the processing results of the relationship between the echo correlation coefficient, the extended dimension and the fused time domain features and frequency domain features of each sampling position and the soil conductivity of the four different regression models are respectively, Figures 6 to 9 The scatter plots of the real response and the predicted response are respectively.
[0168] Figure 2After the optimization and selection of features, the predicted response result graph obtained by using the integrated boosting tree model for training is shown in FIG. 6, in which the horizontal axis represents the sampling position of each, i.e. the test point, and the vertical axis represents the value of the conductivity. There are two curves in the figure, which are the real conductivity curve and the predicted conductivity curve respectively. From the figure, we can see the prediction of the model, and the R 2 = 0.88 is obtained by using the integrated boosting tree model.
[0169] Figure 3 After the optimization and selection of features, the predicted response result graph obtained by using the integrated bagging tree model for training is shown in FIG. 7, in which the horizontal axis represents the sampling position of each, i.e. the test point, and the vertical axis represents the value of the conductivity. There are two curves in the figure, which are the real conductivity curve and the predicted conductivity curve respectively. From the figure, we can see the prediction of the model, and the R 2 = 0.87 is obtained by using the integrated bagging tree model.
[0170] Figure 4 After the optimization and selection of features, the predicted response result graph obtained by using the Gaussian process regression model for training is shown in FIG. 8, in which the horizontal axis represents the sampling position of each, i.e. the test point, and the vertical axis represents the value of the conductivity. There are two curves in the figure, which are the real conductivity curve and the predicted conductivity curve respectively. From the figure, we can see the prediction of the model, and the R 2 = 0.85 is obtained by using the Gaussian process regression model.
[0171] Figure 5 After the optimization and selection of features, the predicted response result graph obtained by using the linear support vector machine (SVM) model for training is shown in FIG. 9, in which the horizontal axis represents the sampling position of each, i.e. the test point, and the vertical axis represents the value of the conductivity. There are two curves in the figure, which are the real conductivity curve and the predicted conductivity curve respectively. From the figure, we can see the prediction of the model, and the R 2 = 0.85 is obtained by using the linear support vector machine (SVM) model.
[0172] Figure 6 The predicted-real response scatter plot is shown in FIG. 10, in which the perfect prediction line represents a straight line drawn when the predicted response and the real response are equal, i.e. perfect prediction. From the figure, we can see that each point is generally distributed near the perfect prediction line, indicating that the prediction effect of the model is good.
[0173] Figure 7To make the predicted value (predicted response) after the ensemble bagging tree model prediction as the ordinate, and the true value (true response) as the abscissa, a prediction-true response scatter plot is formed, where the perfect prediction line represents a straight line drawn when the predicted response and the true response are equal, i.e. perfect prediction. As can be seen from the figure, each point is generally distributed near the perfect prediction line, indicating that the prediction effect of the model is also good.
[0174] Figure 8 To make the predicted value (predicted response) after the ensemble bagging tree model prediction as the ordinate, and the true value (true response) as the abscissa, a prediction-true response scatter plot is formed, where the perfect prediction line represents a straight line drawn when the predicted response and the true response are equal, i.e. perfect prediction. As can be seen from the figure, each point is generally distributed near the perfect prediction line, indicating that the prediction effect of the model is also good.
[0175] Figure 9 To make the predicted value (predicted response) after the ensemble bagging tree model prediction as the ordinate, and the true value (true response) as the abscissa, a prediction-true response scatter plot is formed, where the perfect prediction line represents a straight line drawn when the predicted response and the true response are equal, i.e. perfect prediction. As can be seen from the figure, each point is generally distributed near the perfect prediction line, indicating that the prediction effect of the model is also good.
[0176] As can be seen from the figure, among the four models, the ensemble boosting tree model shows the highest prediction accuracy, and its prediction effect is better than that of other models. At the same time, it can be seen from the prediction results that the echo signal components of different sampling points can better reflect the change rule of soil conductivity. In the low conductivity area, the four models all show stable and accurate prediction ability. This is mainly due to the fact that in the low conductivity area, the reflection of the echo signal is strong, so that the correlation between its components and the soil conductivity is high. However, in the high conductivity area, due to the enhancement of signal scattering and attenuation effect, the correlation between echo signal and soil conductivity decreases, resulting in certain influence on the prediction performance of the model.
[0177] In order to further improve the prediction accuracy of soil conductivity in saline-alkali soil, we introduce the ensemble learning (Ensemble Learning) strategy on the basis of the optimized machine learning model. The ensemble learning method can effectively combine the prediction ability of multiple models, so that the advantages of different models in different soil conductivity ranges can be fully played, thereby improving the overall prediction performance.
[0178] Based on geophysical and agricultural ecological knowledge, different types of soil and their surrounding environment will have an impact on soil conductivity. Therefore, a single model is difficult to maintain stable prediction performance in the entire soil conductivity distribution range. After analyzing the collected soil sample data, it is found that the conductivity value is distributed in the range of [0, 7000], and can be roughly divided into low conductivity ([0, 750] uS / cm), medium conductivity ((750, 2000] uS / cm), and high conductivity ((2000, 7000] uS / cm) according to experience.
[0179] In order to take advantage of different models in different conductivity prediction intervals, a partition weighted integration strategy is adopted, that is, in different soil conductivity ranges, the best performing models are weighted and integrated to ensure the optimal prediction accuracy in each range.
[0180] First, the training set is divided into three parts according to low, medium, and high conductivity, and four models are used for prediction. The predicted values are averaged to determine the conductivity interval to which the predicted values belong; then, according to the R 2 value of the four models in the conductivity interval, the formula
[0181]
[0182] is used to determine the weight coefficient of different models in the corresponding conductivity interval. Since the data in the three conductivity ranges is predicted by four models, at this time U = 4. The weight coefficients of the models are shown in Table 3:
[0183] Table 3 Weight coefficients of each model
[0184]
[0185] The above is the weighted proportion of each interval. It can also be seen that the four models can better predict the soil data in the low conductivity range, and the performance is comparable. In the medium conductivity range, the linear SVM, ensemble bagging tree, and ensemble boosting tree models perform well. In the high conductivity range, the ensemble bagging tree and ensemble boosting tree models perform best. This result can be used as experience to help predict the conductivity of unknown soil areas during actual measurement.
[0186] Finally, the prediction results of each model are integrated by weighted averaging to improve the overall prediction accuracy. The experimental results show that the final determination coefficient (R²) of this scheme reaches 0.90, which is significantly improved compared to single models, as shown in Figures 10-11 .
[0187] Figure 10The prediction result response graph obtained by using the partition adaptive weighted ensemble learning method for training is shown in FIG. 6, wherein the horizontal axis represents the sampling position of each, that is, the test point, and the vertical axis represents the value of the conductivity. There are two curves in the figure, which are the real conductivity curve and the predicted conductivity curve respectively. The experimental results show that the final determination coefficient (R2) of the scheme reaches 0.90, which is obviously improved compared with a single model.
[0188] Figure 11 The prediction-real response scatter diagram is constructed by using the prediction value (predicted response) as the vertical coordinate and the real value (real response) as the horizontal coordinate after training by using the partition adaptive weighted ensemble learning method, wherein the perfect prediction line represents a straight line drawn when the predicted response and the real response are equal, that is, when the perfect prediction is drawn. As can be seen from the figure, each point is well distributed in a very small range near the perfect prediction line, which further illustrates that the prediction accuracy of the ensemble learning is greatly improved compared with the single model.
[0189] The above is the training process. In actual testing, first, data acquisition and feature extraction are performed by using GPR, and then four models are used for prediction. After taking the average of the prediction values, the determination of the adaptive weighting coefficient and the next prediction are started, and the final result is obtained. At the same time, in the prediction measurement stage, not only can the prediction results of the model be used to adaptively select different weighting coefficients according to the determination coefficient to perform fusion prediction, but also the historical training data can be used for dynamic optimization. In this way, not only can the performance difference of the model in different conductivity intervals be accurately adapted, but also the system can be automatically updated without manual setting of the weight, and the robustness and stability of the model to the distribution change of the data are better, so as to improve the accuracy and consistency of the soil conductivity prediction and reduce the model bias.
[0190] In order to verify the superiority of the results, a comparative analysis is made with the existing curve polynomial fitting model and the robust linear regression model (only using the echo correlation coefficient as the prediction feature). Table 4 lists the best correlation coefficient values of the two traditional models and the four machine learning models and the ensemble learning method adopted in the application. As can be seen from the table, the method adopted in the application is obviously better than the traditional polynomial fitting and robust linear regression model, and especially after combining the time domain and frequency domain features of the echo, the value of R2 is significantly improved.
[0191] Table 4 R of different models 2 value
[0192]
[0193] Through the integrated learning method, different models can play the best prediction ability in the respective skilled conductivity range, thereby improving the overall accuracy.
[0194] The application fuses the correlation features of ground penetrating radar (GPR) echoes and the extension and fusion of echo time-frequency domain features.
[0195] The application adopts multiple machine learning models (such as integrated bagging tree, linear support vector machine regression, Gaussian process regression, etc.) for training, and calculates R² values of each model for comparison.
[0196] In summary, the application breaks through the limitations of traditional ground penetrating radar inversion technology and provides an innovative solution for efficient, accurate and stable soil conductivity measurement.
[0197] Although the embodiments of the application have been disclosed as above, they are not limited to the application listed in the specification and the embodiments, and can be fully applied to various fields suitable for the application, and other modifications can be easily realized by those skilled in the art, therefore the application is not limited to specific details and the figures shown and described herein, without departing from the general concept defined by the claims and the equivalent scope.
Claims
1. A method for saline-alkali soil conductivity prediction with multi-features and integrated optimization, characterized in that, The method comprises the following steps: Step one, obtaining radar echo characteristic parameters of multiple sampling positions in saline-alkali soil and the conductivity corresponding to each radar echo characteristic parameter to form a sample set; The radar echo characteristic parameters include radar echo correlation coefficient characteristics, radar echo spectrum characteristics, radar echo time domain characteristics, amplitude difference between adjacent frequency points of radar echo signals, difference between adjacent positive peaks of radar echo signals, and difference between adjacent negative peaks of radar echo signals. Step two, training different machine learning models using the sample set, and selecting multiple optimal machine learning models as conductivity prediction models according to the training results; Step three, determining the weight coefficients of each conductivity prediction model in each conductivity interval according to the accuracy of the prediction results of the samples in different conductivity intervals; Step four, detecting the to-be-detected sampling position in the saline-alkali soil by using a radar, obtaining radar echo characteristic parameters of the to-be-detected sampling position, inputting the radar echo characteristic parameters of the to-be-detected sampling position into the multiple conductivity prediction models to obtain multiple conductivity prediction values, and determining the conductivity interval to which the to-be-detected sampling position belongs according to the multiple conductivity prediction values; Step five, calculating the conductivity prediction value of the to-be-detected sampling position: ; wherein, is the number of conductivity prediction models, is the weight coefficient of the conductivity prediction model in the conductivity interval to which the sampling position belongs, is the conductivity prediction value of the conductivity prediction model, is the number of conductivity prediction models, U is the number of conductivity prediction models.
2. The method of claim 1, wherein, The radar echo correlation coefficient characteristics are: ; wherein and is a normalization coefficient, is a standard radar wavelet, denotes the radar return at any sampling location in a saline soil; denotes time, denotes the length of the selected time window.
3. The method of claim 2, wherein, In step one, the method for obtaining the radar echo spectrum characteristics is: Calculate the sampling frequency step size: ; Perform Fourier transform on the sampling sequence to obtain the sampling frequency of each sampling position. ; determining a frequency range of the sampling sequence ] and integer characteristic frequencies within the frequency range as radar return spectral features; wherein , ; wherein is the index of the frequency component, is the base of the natural logarithm, is the imaginary unit, and are the indices of the lower and upper limit frequencies of the operating frequency band, respectively; denotes the sampling sequence, denotes the number of sampling sequences; and denote the lower and upper limit of the operating frequency band, respectively.
4. The method of claim 1-3, wherein, The radar echo time domain characteristics are the time sequence of multiple peak values of the radar echo signal.
5. The method of claim 4, wherein, Before step two, the following steps are further included: Filter key feature frequencies and key wave peaks from the radar echo spectrum characteristics and the radar echo time domain characteristics based on the MRMR algorithm, fuse the values of the key feature frequencies and the peak values of the key wave peaks to obtain fused spectrum characteristics and fused time domain characteristics, and use the radar echo correlation coefficient characteristics, the fused spectrum characteristics, the fused time domain characteristics, the amplitude difference between adjacent frequency points of the radar echo signals, the difference between adjacent positive peaks of the radar echo signals, and the difference between adjacent negative peaks of the radar echo signals as the radar echo characteristic parameters in the sample set. The calculation formula of the fused spectrum characteristics is:
6. The method of claim 5, wherein, The calculation formula of the fused time domain characteristics is: ; wherein ; In the formula, indicates the fusion spectrum feature, indicates the key feature frequency fusion coefficient, indicates the key feature frequency actual value, indicates the key feature frequency actual value proportion, indicates the key feature frequency MRMR score value proportion; is the number of key feature frequencies.
7. The method of claim 6, wherein, The optimal machine learning models include an integrated boosting tree model, an integrated bagging tree model, a Gaussian process regression model, and a linear support vector machine model. ; wherein ; In the formula, indicates the fusion time domain feature, indicates the fusion coefficient of the key peak value , indicates the peak value of the key peak , indicates the peak value ratio of the key peak , indicates the MRMR score value ratio of the key peak , is the number of key peaks.
8. The method of claim 7, wherein, The calculation formula of the weight coefficients of each conductivity prediction model in each conductivity interval is:
9. The method of claim 8, wherein, In step three, the conductivity is divided into three intervals of low conductivity, medium conductivity, and high conductivity. ; wherein, U is the number of conductivity prediction models, is the coefficient of determination of the i th conductivity prediction model in the corresponding conductivity interval.
10. The method of claim 9, wherein, The conductivity interval corresponding to the low conductivity is [0, 750] uS / cm, the conductivity interval corresponding to the medium conductivity is (750, 2000] uS / cm, and the conductivity interval corresponding to the high conductivity is (2000, 7000] uS / cm.
Citation Information
Patent Citations
GIL three-pillar grounding electrode copper / graphite system performance optimization method and system
CN119252386A
Optimization sampling method of electric power digital power meter, power meter and equipment
CN119691636A