Soil property spatial prediction method fusing multi-source data and its spatial autocorrelation
By fusing the soil property predictions from linear and nonlinear models using a multi-scale geographic normalized weighted fusion model, the problem of unconsidered spatial autocorrelation in soil property prediction is solved, thus improving prediction accuracy and environmental adaptability.
Patent Information
- Application Number
- CN202310939608.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-28
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2043-07-28
AI Technical Summary
Existing technologies fail to effectively consider the spatial autocorrelation of soil properties in soil property prediction, resulting in insufficient prediction accuracy, especially poor adaptability in complex environments.
A multi-scale geographic normalized weighted fusion model is adopted, which combines linear and nonlinear models. The predicted values of soil attribute data from the linear model and the nonlinear model are fused through the multi-scale geographic normalized weighted fusion model. Considering the spatial autocorrelation of soil attributes, the range, optimal normalization index and bandwidth of the multi-scale geographic normalized weighted fusion model are used to calculate the weights, and finally the predicted values of soil attribute data are obtained.
It improves the accuracy of soil property prediction and its adaptability in complex environments. The fusion prediction results are superior to conventional linear and nonlinear methods, and maintain good prediction performance.
Smart Images

Figure CN117056858B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a soil attribute spatial prediction method by fusing multi-source data and spatial autocorrelation thereof, and belongs to the technical field of digital soil mapping and data fusion. Background Art
[0002] Current research on predicting basic soil properties, such as soil organic carbon, often relies on constructing various prediction models to characterize the linear or nonlinear relationships between multiple environmental factors and the predicted soil properties. Commonly used linear models include multiple linear regression (MLR) and geographically weighted regression (GWR). GWR is an extension of the ordinary linear regression model. It adheres to the spatial autocorrelation principle of the first law of geography and innovatively applies the concept of local regression. By incorporating the spatial location of the data into the local regression equation to calculate the weight of the data, it can estimate the parameters of the dependent and explanatory variables at each location while taking into account the spatial weights of neighboring points. Due to the spatial autocorrelation of soil properties, GWR outperforms the ordinary linear regression model based on the least squares method. A series of models have been derived from GWR as its core concept. Among them, the multiscale geographically weighted regression (MGWR) method is an extension of GWR. It is not restricted to the same spatial scale for all modeling processes and is more adaptable to large-scale and environmentally complex regions. Compared to the standard GWR, which uses a single bandwidth for all relationships between environmental variables and the predicted variable in the model, MGWR accounts for variations in the spatial scale of relationships by assigning a separate bandwidth to each relationship in the model. Furthermore, unlike the weighted least squares algorithm used in the standard GWR model, MGWR uses a backward fitting algorithm. Regarding nonlinear models, commonly used machine learning models include random forests (RF), support vector machines (SVM), and artificial neural networks (ANN). Because the relationships between soil properties and environmental variables are complex and inherently nonlinear, nonlinear machine learning methods generally outperform linear methods in training accuracy. Among these nonlinear methods, artificial neural networks are widely used due to their powerful adaptive nonlinear fitting capabilities.
[0003] Although nonlinear methods such as artificial neural networks have shown greater accuracy than conventional linear methods in predicting basic soil properties such as soil organic carbon, they still have drawbacks. One of these is their lack of consideration of spatial autocorrelation. This means that when predicting soil properties, they only consider the relationship between the soil property and the environmental variables at that location, while ignoring the influence of neighboring environmental variables. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a soil attribute spatial prediction method that integrates multi-source data and their spatial autocorrelation.
[0005] To solve the above technical problems, the present invention provides a soil attribute spatial prediction method that integrates multi-source data and their spatial autocorrelation, comprising:
[0006] Collect environmental variables of the target area;
[0007] Inputting the environmental variables of the target area into a pre-trained linear relationship model of soil attribute data and environmental variables and a pre-trained non-linear relationship model of soil attribute data and environmental variables, respectively, to obtain a linear model soil attribute data prediction value and a non-linear model soil attribute data prediction value;
[0008] The predicted values of soil attribute data of the linear model and the predicted values of soil attribute data of the nonlinear model are fused using a multi-scale geographic normalization weighted fusion model to obtain the final predicted values of soil attribute data of the target area.
[0009] Furthermore, the linear relationship model between soil attribute data and environmental variables adopts an ANN model.
[0010] Furthermore, the nonlinear relationship model between soil attribute data and environmental variables adopts the MGWR model.
[0011] Furthermore, the fusion process of the multi-scale geographic normalization weighted fusion model is as follows:
[0012] Obtaining the range of the multi-scale geographic normalized weighted fusion model and the target grid to be predicted;
[0013] Taking the target grid to be predicted as the center, all grids are divided into four quadrants on the plane, and the linear model soil property data prediction value and the nonlinear model soil property data prediction value of each grid layer within the range are obtained; all grids are raster data composed of multiple grids in the target area generated by map software;
[0014] Obtaining predetermined optimal normalized index and optimal bandwidth corresponding to the linear model and optimal normalized index and optimal bandwidth corresponding to the linear nonlinear model;
[0015] The spatial weights corresponding to different quadrants of the target grid at each layer are calculated based on the optimal bandwidth of the corresponding linear model. The spatial weights corresponding to different quadrants of the target grid at each layer are summed and averaged to obtain the average spatial weight of each quadrant. The average spatial weight of each quadrant is then normalized based on the optimal normalization index of the corresponding linear model to obtain a normalized spatial weight matrix. The average value of the soil attribute data in the same quadrant of all grid layers is calculated based on the predicted values of the linear model soil attribute data of all grid layers to obtain the independent variable matrix of the corresponding linear model. The final value corresponding to the linear model is obtained by multiplying the normalized spatial weight matrix and the independent variable matrix of the corresponding linear model and summing them.
[0016] The spatial weights corresponding to different quadrants of the target grid at each layer are calculated based on the optimal bandwidth of the corresponding nonlinear model. The spatial weights corresponding to different quadrants of the target grid at each layer are summed and averaged to obtain the average spatial weight of each quadrant. The average spatial weight of each quadrant is then normalized based on the optimal normalization index of the corresponding nonlinear model to obtain a normalized spatial weight matrix. The average value of the soil attribute data in the same quadrant of all grid layers is calculated based on the predicted values of the nonlinear model soil attribute data of all grid layers to obtain the independent variable matrix of the corresponding nonlinear model. The final value of the nonlinear model is obtained by multiplying the normalized spatial weight matrix and the independent variable matrix of the corresponding nonlinear model and summing them.
[0017] The final value corresponding to the linear model and the final value corresponding to the nonlinear model are added together to obtain the final predicted value of the soil property data of the target grid.
[0018] Furthermore, the process of obtaining the optimal normalized index and optimal bandwidth of the corresponding linear model and the optimal normalized index and optimal bandwidth of the corresponding linear nonlinear model specifically includes:
[0019] Randomly generate several sets of normalized exponents and bandwidths corresponding to linear models and normalized exponents and bandwidths corresponding to nonlinear models within a preset range, wherein the sum of the normalized exponents of the linear model and the normalized exponents of the nonlinear model is 1;
[0020] Obtain the measured values of soil property data of any grid in the target area;
[0021] The grid is used as the target grid, and each set of normalized exponents and bandwidths corresponding to the linear model and the normalized exponents and bandwidths corresponding to the nonlinear model are used as the normalized exponents and bandwidths used in the multi-scale geographic normalized weighted fusion model. The multi-scale geographic normalized weighted fusion model is used to calculate the final predicted values of the soil attribute data corresponding to the grid.
[0022] The mean square error is calculated between the final predicted value of the soil attribute data corresponding to each grid and the measured value of the soil attribute data of the grid. The normalized index and bandwidth of the corresponding linear model and the normalized index and bandwidth of the corresponding nonlinear model corresponding to the final predicted value of the soil attribute data with the smallest mean square error are the optimal normalized index and optimal bandwidth of the corresponding linear model and the optimal normalized index and optimal bandwidth of the corresponding linear nonlinear model.
[0023] Furthermore, the expression of the multi-scale geographic normalization weighted fusion model is:
[0024]
[0025] Among them, y i ,X ij , are the predicted value corresponding to the i-th grid and the independent variable matrix and spatial weight matrix of the j-th fused variable obtained by searching within the range of the i-th grid; h is the number of fused variables; p j ,bw j are the normalized index and bandwidth corresponding to the jth fused variable; Represents normalizing the spatial weight matrix of the jth fused variable to (0, p j ), so that The sum of all spatial weight values in is p j .
[0026] A soil property spatial prediction device integrating multi-source data and their spatial autocorrelation, comprising:
[0027] The collection module is used to collect environmental variables of the target area;
[0028] A prediction module is used to input the environmental variables of the target area into a pre-trained linear relationship model of soil attribute data and environmental variables and a pre-trained non-linear relationship model of soil attribute data and environmental variables, respectively, to obtain a linear model soil attribute data prediction value and a non-linear model soil attribute data prediction value;
[0029] The fusion module is used to fuse the linear model soil attribute data prediction value and the nonlinear model soil attribute data prediction value using a multi-scale geographic normalization weighted fusion model to obtain the final prediction value of the soil attribute data of the target area.
[0030] A computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described above.
[0031] A computer device comprising:
[0032] One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods described.
[0033] The beneficial effects achieved by this invention include significantly improved accuracy of the fused prediction results compared to the original fused data and superiority over conventional linear and nonlinear fusion methods. Furthermore, because the fusion process considers the spatial autocorrelation of soil properties, this method exhibits the strongest adaptability to the overall environment of the prediction area, maintaining good prediction performance in a variety of complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic flow diagram of the present invention;
[0035] Figure 2 It is a schematic diagram of a model of the present invention;
[0036] Figure 3 is a flow chart of an implementation method;
[0037] Figure 4 This is a schematic diagram of accuracy verification comparison;
[0038] Figure 5 This is a comparative schematic diagram of the spatial prediction results of soil organic carbon storage obtained using different methods. DETAILED DESCRIPTION
[0039] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0040] Example 1, as Figure 1 and 3 As shown, this embodiment introduces a soil attribute spatial prediction method that integrates multi-source data and their spatial autocorrelation, including:
[0041] Collect environmental variables of the target area;
[0042] Inputting the environmental variables of the target area into a pre-trained linear relationship model of soil attribute data and environmental variables and a pre-trained non-linear relationship model of soil attribute data and environmental variables, respectively, to obtain a linear model soil attribute data prediction value and a non-linear model soil attribute data prediction value;
[0043] The predicted values of soil attribute data of the linear model and the predicted values of soil attribute data of the nonlinear model are fused using a multi-scale geographic normalization weighted fusion model to obtain the final predicted values of soil attribute data of the target area.
[0044] The fusion process of the multi-scale geographic normalization weighted fusion model is as follows:
[0045] Obtaining the range of the multi-scale geographic normalized weighted fusion model and the target grid to be predicted;
[0046] Taking the target grid to be predicted as the center, all grids are divided into four quadrants on the plane, and the linear model soil property data prediction values and nonlinear model soil property data prediction values of each grid layer within the range are obtained; all the grids are raster data composed of multiple grids in the target area generated by ArcMap.
[0047] The calculation of the range is implemented on ArcMap, that is, the values of all the grids predicted by the two fused models (in this case, the ANN model and the MGWR model) in the study area are calculated. The calculation of the range is actually based on the semivariogram. When the distance between the sampling points increases, the semivariogram reaches a relatively stable constant from the initial value, and the constant value is called the base value; and when the value of the semivariogram reaches the base value from the initial value, the interval distance of the sampling points is called the range. The minimum value of the range in the ANN model and the MGWR model is taken as the calculation range of the fusion model, that is, the range obtained by the present invention.
[0048] The existing training method is used to obtain the trained ANN model and MGWR model, which is not the content of the invention and will not be described here.
[0049] like Figure 2 As shown, obtaining the predetermined optimal normalized index and optimal bandwidth corresponding to the linear model and the optimal normalized index and optimal bandwidth corresponding to the linear nonlinear model;
[0050] The spatial weights corresponding to different quadrants of the target grid at each layer are calculated based on the optimal bandwidth of the corresponding linear model. The spatial weights corresponding to different quadrants of the target grid at each layer are summed and averaged to obtain the average spatial weight of each quadrant. The average spatial weight of each quadrant is then normalized based on the optimal normalization index of the corresponding linear model to obtain a normalized spatial weight matrix. The average value of the soil attribute data in the same quadrant of all grid layers is calculated based on the predicted values of the linear model soil attribute data of all grid layers to obtain the independent variable matrix of the corresponding linear model. The final value corresponding to the linear model is obtained by multiplying the normalized spatial weight matrix and the independent variable matrix of the corresponding linear model and summing them.
[0051] The spatial weights corresponding to different quadrants of the target grid at each layer are calculated based on the optimal bandwidth of the corresponding nonlinear model. The spatial weights corresponding to different quadrants of the target grid at each layer are summed and averaged to obtain the average spatial weight of each quadrant. The average spatial weight of each quadrant is then normalized based on the optimal normalization index of the corresponding nonlinear model to obtain a normalized spatial weight matrix. The average value of the soil attribute data in the same quadrant of all grid layers is calculated based on the predicted values of the nonlinear model soil attribute data of all grid layers to obtain the independent variable matrix of the corresponding nonlinear model. The final value of the nonlinear model is obtained by multiplying the normalized spatial weight matrix and the independent variable matrix of the corresponding nonlinear model and summing them.
[0052] The final value corresponding to the linear model and the final value corresponding to the nonlinear model are added together to obtain the final predicted value of the soil property data of the target grid.
[0053] The process of obtaining the optimal normalized index and optimal bandwidth of the corresponding linear model and the optimal normalized index and optimal bandwidth of the corresponding linear nonlinear model specifically includes:
[0054] Randomly generate several sets of normalized exponents and bandwidths corresponding to linear models and normalized exponents and bandwidths corresponding to nonlinear models within a preset range, wherein the sum of the normalized exponents of the linear model and the normalized exponents of the nonlinear model is 1;
[0055] Obtain the measured values of soil property data of any grid in the target area;
[0056] The grid is used as the target grid, and each set of normalized exponents and bandwidths corresponding to the linear model and the normalized exponents and bandwidths corresponding to the nonlinear model are used as the normalized exponents and bandwidths used in the multi-scale geographic normalized weighted fusion model. The multi-scale geographic normalized weighted fusion model is used to calculate the final predicted values of the soil attribute data corresponding to the grid.
[0057] The mean square error is calculated between the final predicted value of the soil attribute data corresponding to each grid and the measured value of the soil attribute data of the grid. The normalized index and bandwidth of the corresponding linear model and the normalized index and bandwidth of the corresponding nonlinear model corresponding to the final predicted value of the soil attribute data with the smallest mean square error are the optimal normalized index and optimal bandwidth of the corresponding linear model and the optimal normalized index and optimal bandwidth of the corresponding linear nonlinear model.
[0058] The method of the present invention assigns different spatial weights to fused variables at different locations based on the distance between the fused variable and the grid where the target fused data is located. The range of the fused variable under consideration is determined by the minimum range calculated by the semivariogram function (within the range, the smaller the distance between the sampling points, the greater the similarity. When the distance exceeds the range, the spatial autocorrelation is considered to be negligible). In addition, considering that different variables have different spatial scales, we combine the idea of MGWR to give different fused variables different bandwidths during fusion. Finally, after normalizing the corresponding spatial weights of each variable based on the specified normalization index, the prediction result is obtained by weighted aggregation. It can be expressed by the following formula:
[0059]
[0060] Among them, y i ,X ij , are the predicted value corresponding to the i-th grid and the independent variable matrix and spatial weight matrix of the j-th fused variable obtained by searching within the range of the i-th grid; h is the number of fused variables; p j ,bw j are the normalized index and bandwidth corresponding to the jth fused variable; Represents normalizing the spatial weight matrix of the jth fused variable to (0, p j ), which makes The sum of all spatial weight values in is p j .also, Each spatial weight in the spatial weight matrix is calculated based on the Gaussian kernel function, as shown below:
[0061] w=exp(-D 2 / bw 2 )
[0062] Where D is the distance to the corresponding grid.
[0063] In this embodiment, taking soil organic carbon storage as an example, based on the measured organic carbon storage of 952 soil profiles (of which 762 or 80% profile data are used for training, and 190 or 20% profile data are used for verification), and 9 environmental variables (precipitation, temperature, multi-scale ridge index, green light band, red light band, near infrared band and enhanced vegetation index), the ANN model and MGWR model are constructed and trained using the training set data respectively, and then the conventional linear fusion method multivariate linear regression MLR_F, the nonlinear fusion method artificial neural network ANN_F and our MGNW_F are used to fuse the ANN and MGWR prediction results, and the three models are trained separately using the training set data. The prediction results of ANN, MGWR, MLR_F, ANN_F and MGNW_F are respectively verified for accuracy using the verification set data, and the results are as follows: Figure 4 As shown, it can be seen that MGNW_F has the lowest mean square error (MSE) and mean absolute error (MAE) among the five models, and its correlation coefficient is also the highest among the five models.
[0064] like Figure 5 Figure 2 shows the prediction results of ANN, MGWR, MLR_F, ANN_F, and MGNW_F, where (a) is the ANN model, (b) is the MGWR model, (c) is the MLR_F, (d) is the ANN_F, and (e) is the MGNW_F. It can be seen that the prediction results of the fusion model well combine the characteristics of the ANN model and the MGWR model, achieving a good fusion effect.
[0065] Example 2, based on the same inventive concept as Example 1, introduces a soil property spatial prediction device that integrates multi-source data and their spatial autocorrelation, including:
[0066] The collection module is used to collect environmental variables of the target area;
[0067] A prediction module is used to input the environmental variables of the target area into a pre-trained linear relationship model of soil attribute data and environmental variables and a pre-trained non-linear relationship model of soil attribute data and environmental variables, respectively, to obtain a linear model soil attribute data prediction value and a non-linear model soil attribute data prediction value;
[0068] The fusion module is used to fuse the linear model soil attribute data prediction value and the nonlinear model soil attribute data prediction value using a multi-scale geographic normalization weighted fusion model to obtain the final prediction value of the soil attribute data of the target area.
[0069] Example 3, based on the same inventive concept as other examples, this example introduces a computer-readable storage medium storing one or more programs, characterized in that the one or more programs include instructions that, when executed by a computing device, enable the computing device to execute any of the methods described.
[0070] Example 4, based on the same inventive concept as the other examples, this example introduces a computer device, including:
[0071] One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods described.
[0072] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0073] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0074] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0075] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0076] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for spatial prediction of soil properties by integrating multi-source data and their spatial autocorrelation, characterized in that: include: Collect environmental variables of the target area; Inputting the environmental variables of the target area into a pre-trained linear relationship model of soil attribute data and environmental variables and a pre-trained non-linear relationship model of soil attribute data and environmental variables, respectively, to obtain a linear model soil attribute data prediction value and a non-linear model soil attribute data prediction value; The linear model soil property data prediction value and the nonlinear model soil property data prediction value are fused using a multi-scale geographic normalization weighted fusion model to obtain a final prediction value of the soil property data of the target area; The fusion process of the multi-scale geographic normalization weighted fusion model is as follows: Obtaining the range of the multi-scale geographic normalized weighted fusion model and the target grid to be predicted; Taking the target grid to be predicted as the center, all grids are divided into four quadrants on the plane, and the linear model soil property data prediction value and the nonlinear model soil property data prediction value of each grid layer within the range are obtained; all grids are raster data composed of multiple grids in the target area generated by map software; Obtaining predetermined optimal normalized index and optimal bandwidth corresponding to the linear model and optimal normalized index and optimal bandwidth corresponding to the linear nonlinear model; The spatial weights corresponding to different quadrants of the target grid at each layer are calculated based on the optimal bandwidth of the corresponding linear model. The spatial weights corresponding to different quadrants of the target grid at each layer are summed and averaged to obtain the average spatial weight of each quadrant. The average spatial weight of each quadrant is then normalized based on the optimal normalization index of the corresponding linear model to obtain a normalized spatial weight matrix. The average value of the soil attribute data in the same quadrant of all grid layers is calculated based on the predicted values of the linear model soil attribute data of all grid layers to obtain the independent variable matrix of the corresponding linear model. The final value corresponding to the linear model is obtained by multiplying the normalized spatial weight matrix and the independent variable matrix of the corresponding linear model and summing them. The spatial weights corresponding to different quadrants of the target grid at each layer are calculated based on the optimal bandwidth of the corresponding nonlinear model. The spatial weights corresponding to different quadrants of the target grid at each layer are summed and averaged to obtain the average spatial weight of each quadrant. The average spatial weight of each quadrant is then normalized based on the optimal normalization index of the corresponding nonlinear model to obtain a normalized spatial weight matrix. The average value of the soil attribute data in the same quadrant of all grid layers is calculated based on the predicted values of the nonlinear model soil attribute data of all grid layers to obtain the independent variable matrix of the corresponding nonlinear model. The final value of the nonlinear model is obtained by multiplying the normalized spatial weight matrix and the independent variable matrix of the corresponding nonlinear model and summing them. The final value corresponding to the linear model and the final value corresponding to the nonlinear model are added together to obtain the final predicted value of the soil property data of the target grid.
2. The soil attribute spatial prediction method based on the fusion of multi-source data and their spatial autocorrelation according to claim 1 is characterized in that: The linear relationship model between soil attribute data and environmental variables adopts an ANN model.
3. The soil attribute spatial prediction method based on the fusion of multi-source data and spatial autocorrelation according to claim 1 is characterized in that: The nonlinear relationship model between soil attribute data and environmental variables adopts the MGWR model.
4. The soil attribute spatial prediction method based on integration of multi-source data and spatial autocorrelation according to claim 1 is characterized in that: The process of obtaining the optimal normalized index and optimal bandwidth of the corresponding linear model and the optimal normalized index and optimal bandwidth of the corresponding linear nonlinear model specifically includes: Randomly generate several sets of normalized exponents and bandwidths corresponding to linear models and normalized exponents and bandwidths corresponding to nonlinear models within a preset range, wherein the sum of the normalized exponents of the linear model and the normalized exponents of the nonlinear model is 1; Obtain the measured values of soil property data of any grid in the target area; The grid is used as the target grid, and each set of normalized exponents and bandwidths corresponding to the linear model and the normalized exponents and bandwidths corresponding to the nonlinear model are used as the normalized exponents and bandwidths used in the multi-scale geographic normalized weighted fusion model. The multi-scale geographic normalized weighted fusion model is used to calculate the final predicted values of the soil attribute data corresponding to the grid. The mean square error is calculated between the final predicted value of the soil attribute data corresponding to each grid and the measured value of the soil attribute data of the grid. The normalized index and bandwidth of the corresponding linear model and the normalized index and bandwidth of the corresponding nonlinear model corresponding to the final predicted value of the soil attribute data with the smallest mean square error are the optimal normalized index and optimal bandwidth of the corresponding linear model and the optimal normalized index and optimal bandwidth of the corresponding linear nonlinear model.
5. The soil attribute spatial prediction method by integrating multi-source data and spatial autocorrelation according to claim 1 is characterized in that: The expression of the multi-scale geographic normalization weighted fusion model is: ; in, y i 、 X ij 、 They correspond to i The predicted value of each grid and the corresponding i The first grid is searched within the range j The independent variable matrix and spatial weight matrix of the fused variables; h is the number of fused variables; 、 They correspond to j Normalized exponent and bandwidth of the fused variables; Representatives will j The spatial weight matrix of the fused variables is normalized to , making The sum of all spatial weight values in is .
6. A soil property spatial prediction device that integrates multi-source data and their spatial autocorrelation, characterized in that: include: The collection module is used to collect environmental variables of the target area; A prediction module is used to input the environmental variables of the target area into a pre-trained linear relationship model of soil attribute data and environmental variables and a pre-trained non-linear relationship model of soil attribute data and environmental variables, respectively, to obtain a linear model soil attribute data prediction value and a non-linear model soil attribute data prediction value; A fusion module is used to fuse the soil attribute data prediction value of the linear model and the soil attribute data prediction value of the nonlinear model using a multi-scale geographic normalization weighted fusion model to obtain a final prediction value of the soil attribute data of the target area; The fusion process of the multi-scale geographic normalization weighted fusion model is as follows: Obtaining the range of the multi-scale geographic normalized weighted fusion model and the target grid to be predicted; Taking the target grid to be predicted as the center, all grids are divided into four quadrants on the plane, and the linear model soil property data prediction value and the nonlinear model soil property data prediction value of each grid layer within the range are obtained; all grids are raster data composed of multiple grids in the target area generated by map software; Obtaining predetermined optimal normalized index and optimal bandwidth corresponding to the linear model and optimal normalized index and optimal bandwidth corresponding to the linear nonlinear model; The spatial weights corresponding to different quadrants of the target grid at each layer are calculated based on the optimal bandwidth of the corresponding linear model. The spatial weights corresponding to different quadrants of the target grid at each layer are summed and averaged to obtain the average spatial weight of each quadrant. The average spatial weight of each quadrant is then normalized based on the optimal normalization index of the corresponding linear model to obtain a normalized spatial weight matrix. The average value of the soil attribute data in the same quadrant of all grid layers is calculated based on the predicted values of the linear model soil attribute data of all grid layers to obtain the independent variable matrix of the corresponding linear model. The final value corresponding to the linear model is obtained by multiplying the normalized spatial weight matrix and the independent variable matrix of the corresponding linear model and summing them. The spatial weights corresponding to different quadrants of the target grid at each layer are calculated based on the optimal bandwidth of the corresponding nonlinear model. The spatial weights corresponding to different quadrants of the target grid at each layer are summed and averaged to obtain the average spatial weight of each quadrant. The average spatial weight of each quadrant is then normalized based on the optimal normalization index of the corresponding nonlinear model to obtain a normalized spatial weight matrix. The average value of the soil attribute data in the same quadrant of all grid layers is calculated based on the predicted values of the nonlinear model soil attribute data of all grid layers to obtain the independent variable matrix of the corresponding nonlinear model. The final value of the nonlinear model is obtained by multiplying the normalized spatial weight matrix and the independent variable matrix of the corresponding nonlinear model and summing them. The final value corresponding to the linear model and the final value corresponding to the nonlinear model are added together to obtain the final predicted value of the soil property data of the target grid.
7. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any one of the methods according to claims 1 to 5 .
8. A computer device, characterized in that: include, One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for executing any one of the methods according to claims 1 to 5.
Citation Information
Patent Citations
Soil nutrient prediction and comprehensive evaluation method based on machine learning algorithm
CN109374860A
Target soil property content prediction method based on soil transfer function
WO2021184550A1