A geographically weighted regression model based on nonlinear kernel mapping and its modeling method and application

Through the geographically weighted regression model based on nonlinear kernel mapping, the problem of difficulty in handling non-normal distribution and complex spatial structure in existing technologies is solved, and high-precision identification of geochemical anomalies and efficient exploration of deep mineral resources are achieved.

CN120277638BActive Publication Date: 2025-09-09CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510775991.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-09
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

Existing geochemical data processing methods are unable to effectively handle non-normal distribution and complex spatial structures, resulting in limited accuracy in mineral resource exploration.

Method used

A geographically weighted regression model based on nonlinear kernel mapping is adopted. By integrating the nonlinear modeling of kernel functions and the spatial locality analysis of geographically weighted regression, it dynamically adapts to the boundaries and scale differences of geological units and improves data utilization and recognition rate.

Benefits of technology

It significantly improves the accuracy and stability of geochemical anomaly identification, especially the ability to identify weak anomaly areas, and provides efficient and accurate technical means for deep mineral resource exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277638B_ABST
    Figure CN120277638B_ABST
Patent Text Reader

Abstract

The present invention discloses a geographically weighted regression model based on nonlinear kernel mapping, its modeling method and application. The model first obtains geochemical sample data containing spatial positions, constructs a standardized sample set, then constructs an initial geographically weighted regression model based on the sample set, and solves the linear optimal solution of its parameter vector, then reconstructs local weights based on the sample set, and performs adaptive bandwidth optimization to generate an adaptive spatial weight matrix for geographically weighted regression, and then obtains a kernelized geographically weighted regression model through a nonlinear kernel function. The model is based on a joint nonlinear system of kernel mapping and spatial weights, and achieves the technical purpose of local linearization of nonlinear problems in geographic models. When this model is used in geochemical anomaly identification, it can significantly improve the accuracy and stability of anomaly detection, especially it can more accurately identify weak anomaly areas, providing efficient and accurate technical means for deep mineral resource exploration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a geographically weighted regression model, in particular to a geographically weighted regression model based on nonlinear kernel mapping and a modeling method and application thereof, belonging to the technical field of geochemical exploration. Background Art

[0002] Shallow mineral resources on Earth are becoming increasingly depleted, and the exploration of deep mineral resources is gaining more and more attention. However, element concentration detection is easily affected by complex geological conditions, and efficient mineral resource exploration has become an industry challenge. Over the past few decades, scholars have proposed many methods for identifying exploration geochemical anomalies, which are mainly divided into two categories: those based on data frequency distribution and those based on data frequency-spatial distribution. Methods based on data frequency distribution include classical statistics, multivariate statistics (such as principal component analysis), and data exploration analysis methods. This type of method assumes that the data follows a normal distribution or approximates a normal distribution through transformation. However, geochemical data often exhibits significant non-normal distribution characteristics (such as skewness and multimodal distribution) due to the superposition of multiple geological effects, resulting in deviations in statistical parameter estimation and limited anomaly identification accuracy. Methods based on data frequency-spatial distribution are represented by geostatistical models (such as Kriging interpolation) and fractal / multi-fractal models. Although they can take into account the spatial correlation and local characteristics of the data, they have two limitations: first, they cannot adequately describe the multi-scale characteristics of the data (such as the scale differences in element enrichment in different mineralization stages) and complex spatial structures (such as the discontinuous effects of faults and lithologic boundaries); second, they are essentially still based on linear assumptions or simple nonlinear transformations, making it difficult to adapt to the complex nonlinear synergistic mechanisms between elements in the mineralization process (such as the coupled enrichment of trace elements and major elements, and the nonlinear correlation of isotope fractionation).

[0003] To address the nonlinear characteristics of geochemical data, various local anomaly identification methods have emerged, such as local singularity analysis and local domain statistics. These methods focus on local anomalies using sliding windows or spatial weight matrices. However, these methods generally suffer from the following problems: 1) they rely on manually set window sizes or weight functions, which lacks adaptability to spatial nonstationarity; 2) they fail to effectively integrate nonlinear feature modeling, making it easy to miss key anomaly signals when dealing with nonlinear mapping relationships between element concentrations. Therefore, there is an urgent need for an adaptive geographic model with high accuracy and high recognition rate. Summary of the Invention

[0004] To address the challenges of the prior art, the first objective of the present invention is to provide a geographically weighted regression model based on nonlinear kernel mapping. This model, based on a joint nonlinear system of kernel mapping and spatial weights, achieves the technical goal of locally linearizing nonlinear problems in geographic models. For geochemical data characterized by multi-element nonlinear correlations and spatial heterogeneity, it effectively leverages the local spatial characteristics of the data with nonlinear covariates, thereby improving the model's data utilization, reducing the loss of nonlinear signals, and significantly enhancing the model's accuracy and recognition rate.

[0005] A second objective of the present invention is to provide a geographically weighted regression (GWR) modeling method based on nonlinear kernel mapping. This method integrates nonlinear kernel modeling with spatial localization analysis using GWR to accurately identify weak geochemical anomalies under complex geological conditions. This model implicitly maps the raw characteristics of geochemical elements into a high-dimensional space, effectively capturing the complex nonlinear synergistic mechanisms between elements while avoiding the direct display and calculation of high-dimensional features. This overcomes the limitations of traditional GWR linear assumptions and improves the accuracy of complex data characterization, such as non-normal distributions. Furthermore, based on the GWR local weighting mechanism and adaptive bandwidth optimization, it dynamically adapts to differences in geological unit boundaries and scales, thereby enhancing the ability to characterize spatial heterogeneity.

[0006] A third objective of the present invention is to provide an application of a geographically weighted regression model based on nonlinear kernel mapping for geochemical anomaly identification. Based on the excellent performance of this model, its application in identifying geochemical anomalies can significantly improve the accuracy and stability of anomaly detection, particularly by more accurately identifying weakly anomalous areas, providing a highly efficient and accurate technical approach for deep mineral resource exploration.

[0007] In order to achieve the above technical objectives, the present invention provides a geographically weighted regression modeling method based on nonlinear kernel mapping, which is characterized by comprising:

[0008] Step S1: Acquire geochemical sample data including spatial locations and construct a standardized sample set;

[0009] Step S2: construct an initial geographically weighted regression model based on the standardized sample set, and solve the linear optimal solution of its parameter vector;

[0010] Step S3: Based on the standardized sample set, the local weights are calculated using a Gaussian weight function, and adaptive bandwidth optimization is performed to generate an adaptive spatial weight matrix for geographically weighted regression;

[0011] Step S4: Based on the initial geographically weighted regression model, a nonlinear kernel function is used to convert the original features in the sample set into a high-order space through implicit mapping to obtain a kernelized geographically weighted regression model.

[0012] As a preferred solution, the geochemical sample data includes: multi-dimensional element feature vectors, spatial coordinates and sample labels.

[0013] As a preferred solution, the standardization process is: normalizing the sample data to eliminate dimensional differences.

[0014] As a preferred solution, the geochemical sample data is , its sample feature matrix is:

[0015] Formula 1: ;

[0016] In formula 1, For the Sample No. element feature vector, is the total number of samples, is the feature dimension.

[0017] As a preferred solution, the process of solving the linear optimal solution of the eigenvector in the initial geographically weighted regression model is: obtaining the objective function of its target position by estimating the regression coefficient of spatial variation through local weighted least squares, and then converting it into the matrix form of the loss function and performing differentiation to obtain the result.

[0018] As a preferred solution, the mathematical expression of the objective function is:

[0019] Formula 2: ;

[0020] The matrix form of the loss function is:

[0021] Formula 3: ;

[0022] The process of finding the optimal solution is:

[0023] Formula 4: ;

[0024] Formula 5: ;

[0025] In formulas 2 to 5, The parameter vector to be estimated, is a matrix composed of the geochemical element characteristics of all sample points, The geochemical characteristic vector reconstructed for the target, is the spatial weight.

[0026] As a preferred solution, the process of generating the adaptive spatial weight matrix is ​​as follows: based on the sample spatial position , constructing the spatial weight matrix of geographically weighted regression using Gaussian weights , the process is:

[0027] Formula 6: ;

[0028] Formula 7: ;

[0029] In Equation 6 and Equation 7, The target sample and The Euclidean distance of samples, For bandwidth.

[0030] As a preferred solution, the bandwidth The adaptive optimization process is as follows: determine the initial bandwidth through cross-validation and Akaike information criterion, select the bandwidth interval for iteration, and calculate the AUC value corresponding to the bandwidth in the iterative process. When the AUC value is ≥90%, the method is obtained.

[0031] As a preferred solution, the process of obtaining the kernelized geographically weighted regression model is as follows:

[0032] Step S4-1: Map the original features to the high-dimensional space through implicit mapping, and define the high-dimensional space loss function. The calculation process is:

[0033] Formula 8: ( ), ;

[0034] Formula 9: ;

[0035] Substituting Equation 9 into Equation 3, we get:

[0036] Formula 10: ;

[0037] Formula 11: ;

[0038] Step S4-2: Construct a kernel matrix based on the Gaussian kernel function and the implicit mapping described in step S4-1. , the process is:

[0039] Formula 12: ;

[0040] Formula 13: ;

[0041] Equation 14: ;

[0042] Equation 15: ;

[0043] Step S4-3: According to the reproducing kernel Hilbert space theory, the nonlinear Convert it into a linear combination of high-dimensional features, and then substitute it into the objective function to convert it into a kernel-only matrix And the kernelized objective function of the regularization term, the process is:

[0044] Equation 16: ;

[0045] Equation 17: ;

[0046] In formulas 8 to 17, Represents a nonlinear mapping relationship, is a high-dimensional feature matrix, is the high-dimensional space parameter vector,

[0047] is the coefficient to be optimized, which is , is the regularization parameter.

[0048] As a preferred solution, the high-dimensional space parameter vector The solution process is: through eigenvalue decomposition, the high-dimensional space parameters are converted into the expression of the kernel matrix, and then the decomposition result is inverted to obtain the parameter solution. The process is:

[0049] Equation 18: ;

[0050] Equation 19: .

[0051] As a preferred solution, the prediction formula for the target sample in the kernelized geographically weighted regression model is:

[0052] Equation 20: ;

[0053] or,

[0054] Equation 21: ;

[0055] in, is the objective function, It is a row vector consisting of the kernel function values ​​of the target sample and all training samples.

[0056] The present invention also provides a geographically weighted regression model based on nonlinear kernel mapping, which is obtained by any of the above-mentioned modeling methods.

[0057] The present invention also provides an application of a geographically weighted regression model based on nonlinear kernel mapping for geochemical anomaly identification, the process of which is as follows:

[0058] i) Based on the predicted values ​​obtained after reconstruction by the kernelized geographically weighted regression model, the difference between the reconstructed geochemical data and the original values ​​of each sampling point is calculated. The anomaly score is defined as the sum of squares of the reconstruction errors. Then, the geochemical anomaly score of the i-th sampling point is Calculated as:

[0059] Equation 22: ;

[0060] ii) Using the maximum Youden index to divide the geochemical anomaly score threshold, when the geochemical anomaly score When it is higher than the threshold, it is marked as an abnormal area;

[0061] iii) Using GIS technology, the anomaly identification results are superimposed with spatial data such as geological maps and structural lines to generate a visual anomaly distribution map.

[0062] Compared with the prior art, the beneficial technical effects of the technical solution of the present invention are:

[0063] 1) The model provided by this invention is based on a joint nonlinear system of kernel mapping and spatial weighting, achieving the technical goal of local linearization of nonlinear problems in geographic models. For the nonlinear correlation and spatial heterogeneity of multi-element geochemical data, it can achieve effective synergy between the local spatial characteristics of the data and nonlinear covariates, thereby improving the data utilization of the model, reducing the loss of nonlinear signals, and significantly improving the accuracy and recognition rate of the model.

[0064] 2) The modeling method provided by this invention achieves accurate identification of weak geochemical anomalies under complex geological conditions by integrating nonlinear modeling of kernel functions with spatial localization analysis of geographically weighted regression. This model implicitly maps the raw characteristics of geochemical elements into a high-dimensional space, effectively capturing the complex nonlinear synergistic mechanisms between elements while avoiding the direct display and calculation of high-dimensional features. This overcomes the limitations of traditional GWR linear assumptions and improves the accuracy of complex data characterization, such as non-normal distributions. Furthermore, based on the local weight mechanism of geographically weighted regression and adaptive bandwidth optimization, it dynamically adapts to differences in geological unit boundaries and scales, thereby enhancing the ability to characterize spatial heterogeneity.

[0065] 3) In the technical solution provided by the present invention, based on the excellent performance of the above-mentioned model, its application in identifying geochemical anomalies can significantly improve the accuracy and stability of anomaly detection, especially more accurately identifying weak anomaly areas, providing an efficient and accurate technical means for deep mineral resource exploration. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 This is a line graph of AUC corresponding to different bandwidths during the bandwidth adaptive optimization process in Example 1 of the present invention;

[0067] Figure 2 This is a line graph of AUC corresponding to different regularization coefficients during the adaptive optimization process of the regularization coefficient in Example 1 of the present invention;

[0068] Figure 3 This is a three-dimensional graph showing the AUC changes corresponding to the bandwidth and regularization coefficient combination in Example 1 of the present invention;

[0069] Figure 4 Schematic diagram of geochemical anomalies identified by the model obtained in Example 1 of the present invention;

[0070] Figure 5 This is a distribution diagram of regression coefficients of different elements identified by the model obtained in Example 1 of the present invention;

[0071] in, Figure 5 (a) is the regression coefficient distribution diagram of the Ag element identified by the obtained model, Figure 5 (b) is the regression coefficient distribution diagram of the As element identified by the obtained model. Figure 5 (c) is the regression coefficient distribution diagram of Bi element identified by the obtained model. Figure 5 (d) is the regression coefficient distribution diagram of the Pb element identified by the obtained model. Figure 5 (e) Distribution diagram of the regression coefficient of the Sb element identified by the obtained model. DETAILED DESCRIPTION

[0072] The technical solution of the present invention is further described in detail below with reference to specific embodiments and accompanying drawings. To facilitate understanding of the present invention, the present invention will be described in more comprehensive and detailed form below with reference to the accompanying drawings and preferred embodiments. It should be noted that the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0073] Example 1

[0074] This embodiment provides a geographically weighted regression model based on nonlinear kernel mapping, and its specific modeling process is as follows:

[0075] Step S1: Acquire geochemical sample data including spatial locations and construct a standardized sample set;

[0076] The geochemical sample data includes: multi-dimensional element feature vectors, spatial coordinates and sample labels.

[0077] The standardization process is to normalize the sample data to eliminate dimensional differences.

[0078] The geochemical sample data is , its sample feature matrix is:

[0079] Formula 1: ;

[0080] In formula 1, For the Sample No. element feature vector, is the total number of samples, is the feature dimension. In this embodiment =39.

[0081] Step S2: construct an initial geographically weighted regression model based on the standardized sample set, and solve the linear optimal solution of its parameter vector;

[0082] The linear optimal solution of the eigenvector in the initial geographically weighted regression model is obtained by estimating the regression coefficient of spatial variation through local weighted least squares to obtain the objective function of the target position, and then converting it into a matrix form of the loss function and performing differentiation to obtain the objective function;

[0083] Specifically, the mathematical expression of the objective function is:

[0084] Formula 2: ;

[0085] This function is mainly used to measure the model prediction value and the true label The goal is to find the difference that smallest ;

[0086] The matrix form of the loss function is:

[0087] Formula 3: ;

[0088] The process of finding the optimal solution is:

[0089] Formula 4: ;

[0090] Formula 5: ;

[0091] In formulas 2 to 5, The parameter vector to be estimated, is a matrix composed of the geochemical element characteristics of all sample points, The geochemical characteristic vector reconstructed for the target, is the spatial weight;

[0092] Step S3: Based on the standardized sample set, the local weights are calculated using a Gaussian weight function, and adaptive bandwidth optimization is performed to generate an adaptive spatial weight matrix for geographically weighted regression;

[0093] The generation process of the adaptive spatial weight matrix is ​​as follows: based on the sample spatial position , constructing the spatial weight matrix of geographically weighted regression using Gaussian weights , the process is:

[0094] Formula 6: ;

[0095] Formula 7: ;

[0096] In Equation 6 and Equation 7, The target sample and The Euclidean distance of samples, is bandwidth;

[0097] The bandwidth The adaptive optimization process is as follows: determine the initial bandwidth through cross-validation and Akaike Information Criterion, select the bandwidth interval for iteration, and calculate the AUC value corresponding to the bandwidth in the iterative process. When the AUC value is ≥90%, it is obtained;

[0098] Step S4: using a nonlinear kernel function to convert the original features in the sample set into a high-order space through implicit mapping to obtain a kernelized geographically weighted regression model;

[0099] The process of obtaining the kernelized geographically weighted regression model is as follows:

[0100] Step S4-1: Map the original features to the high-dimensional space through implicit mapping, and define the high-dimensional space loss function. The calculation process is:

[0101] Formula 8: ( ), ;

[0102] Formula 9: ;

[0103] Substituting Equation 9 into Equation 3, we get:

[0104] Formula 10: ;

[0105] Formula 11: ;

[0106] Step S4-2: Construct a kernel matrix based on the Gaussian kernel function and the implicit mapping described in step S4-1. , the process is:

[0107] Formula 12: ;

[0108] Formula 13: ;

[0109] Equation 14: ;

[0110] Equation 15: ;

[0111] Step S4-3: According to the reproducing kernel Hilbert space theory, the nonlinear Convert it into a linear combination of high-dimensional features, and then substitute it into the objective function to convert it into a kernel-only matrix And the kernelized objective function of the regularization term, the process is:

[0112] Equation 16: ;

[0113] Equation 17: ;

[0114] In formulas 8 to 17, Represents a nonlinear mapping relationship, is a high-dimensional feature matrix, is the high-dimensional space parameter vector,

[0115] is the coefficient to be optimized, which is , is the regularization parameter.

[0116] The present invention also performs eigenvalue decomposition and parameter solution on the above model, and the high-dimensional space parameter vector The solution process is: through eigenvalue decomposition, the high-dimensional space parameters are converted into the expression of the kernel matrix, and then the decomposition result is inverted to obtain the parameter solution. The process is:

[0117] First, due to the high-dimensional matrix It is difficult to directly invert, so eigenvalue decomposition is performed. Symmetric positive definite, so this is a symmetric positive definite decomposition:

[0118] Equation 18: ;

[0119] in is an orthogonal eigenvector matrix, is a diagonal matrix of eigenvalues. The modified decomposition converts the high-dimensional space parameters into a form that can be represented by the kernel matrix, avoiding explicit high-dimensional calculations and conforming to the core idea of ​​"implicit mapping" of the kernel method.

[0120] Next, we use the decomposition result to invert and obtain the parameter solution:

[0121] Equation 19: .

[0122] Furthermore, the prediction formula for the target sample in the kernelized geographically weighted regression model is:

[0123] Equation 20: ;

[0124] or,

[0125] Equation 21: ;

[0126] in, is the objective function, It is a row vector consisting of the kernel function values ​​of the target sample and all training samples.

[0127] To further illustrate the technical effect of the model provided in this embodiment, taking the Jiaoxibei gold mining area in Shandong Province as an example, the geochemical measurement values ​​of 39 elements of stream sediments in this area were selected as raw data. In a computer environment (Nvidia RTX 3090Ti 24G graphics card), the geographically weighted regression model based on nonlinear kernel mapping proposed in Example 1 was used to extract geochemical anomalies. The specific process is as follows:

[0128] First, the original data is transformed into ilr, and then the transformed data is input into the model for training.

[0129] Then determine the bandwidth and regularization coefficient, where 1km is the bandwidth selection interval, calculate the AUC corresponding to different bandwidths, and observe the trend of AUC changing with bandwidth. Figure 1 It can be seen that within a certain range, as the bandwidth increases, the corresponding AUC increases, and then as the bandwidth continues to increase, the AUC decreases instead of increasing. When the bandwidth is about 50km, the AUC reaches its maximum value; the same as the bandwidth determination process, with 1 as the regularization coefficient selection interval, calculate the AUC corresponding to different regularization coefficients, and observe the trend of AUC changing with bandwidth. Figure 2 It can be seen that AUC increases with the increase of coefficient within a certain range, but begins to decrease after a point and reaches its maximum value when the coefficient is around 10.

[0130] Furthermore, in order to obtain the best combination of bandwidth and regularization coefficient, the present invention combines the adjustable ranges of the two into a 70×100 point set matrix and calculates the AUC value corresponding to their combination. The results are as follows: Figure 3 As shown in the figure, when the bandwidth is 50 km and the coefficient is 10, the AUC is the highest, which is 0.914.

[0131] Then, the outliers obtained by interpolating the gold points in the gold mine are used as positive samples, and the randomly selected non-outliers are used as negative samples. When selecting negative samples, 100 negative sample sets are randomly selected, and the AUC corresponding to each negative sample set is calculated and the average is taken.

[0132] According to the predicted value obtained after reconstruction of the kernelized geographically weighted regression model, the difference between the reconstructed geochemical data and the original value of each sampling point is calculated. The anomaly score is defined as the sum of the squares of the reconstruction errors. Then, the geochemical anomaly score of the i-th sampling point is Calculated as:

[0133] Equation 22: ;

[0134] The maximum Youden index is used to divide the geochemical anomaly score threshold. When it is higher than the threshold, it is marked as an abnormal area.

[0135] Using GIS technology, the anomaly identification results are superimposed with geological maps, structural lines and other spatial data to generate a visual anomaly distribution map. The results are as follows: Figure 4 As shown by Figure 4 It can be seen that most gold spots are distributed in areas with higher outlier values, and most anomalies are distributed in the northeast and southeast along the main fault zone in the northwest of Jiaodong. The receiver operating characteristic curve (ROC) is used to test the performance of the model to measure the spatial correlation between the geochemical model and the known mineralization. In this example, the AUC of the model provided by the present invention is ≥0.91, which fully demonstrates that the model has excellent performance and there is a strong correlation between geochemical anomalies and mineral spots.

[0136] In order to further illustrate that the model provided by the present invention can effectively overcome the problem that the spatial regression model in the prior art cannot fully reflect the true characteristics of spatial data, and to explain the necessity of the model provided by the present invention to use a large amount of nonlinear data and reconstruct spatial weights, the present invention uses the spatial regression coefficient to detect the spatial correlation between other elements and the Au element background reconstruction under the above-mentioned optimal bandwidth and regularization coefficient combination.

[0137] Ag, As, Bi, Pb and Sb elements were randomly selected, and the spatial regression coefficients of each element at the sample point were calculated using the model provided in Example 1 and visualized. The results are shown in Figure 2. Figure 5 As shown. Figure 5It can be seen that the coefficients of different elements are distributed in different regions, with some regions having negative values ​​and some regions having positive values. This indicates that the geochemical distribution patterns of different elements have different spatial correlations and non-steady-state relationships, which is further reflected in the different impacts on the construction of the Au background value. Therefore, the model provided by the present invention, based on the joint nonlinear system of kernel mapping and spatial weights, has better authenticity, accuracy and recognition.

Claims

1. A geographically weighted regression modeling method based on nonlinear kernel mapping, characterized in that: include: Step S1: Acquire geochemical sample data including spatial locations and construct a standardized sample set; Step S2: construct an initial geographically weighted regression model based on the standardized sample set, and solve the linear optimal solution of its parameter vector; Step S3: Based on the standardized sample set, the local weights are calculated using a Gaussian weight function, and adaptive bandwidth optimization is performed to generate an adaptive spatial weight matrix for geographically weighted regression; Step S4: Based on the initial geographically weighted regression model, a nonlinear kernel function is used to implicitly map the original features in the sample set to a high-order space to obtain a kernelized geographically weighted regression model. The linear optimal solution of the eigenvector in the initial geographically weighted regression model is obtained by estimating the regression coefficient of spatial variation through local weighted least squares to obtain the objective function of the target position, and then converting it into a matrix form of the loss function and performing differentiation to obtain the objective function; The process of obtaining the kernelized geographically weighted regression model is as follows: Step S4-1, mapping the original features to a high-dimensional space through implicit mapping, and defining a high-dimensional space loss function; Step S4-2, constructing a kernel matrix according to the Gaussian kernel function and the implicit mapping described in step S4-1; Step S4-3: According to the reproducing kernel Hilbert space theory, the nonlinear high-dimensional space parameter vector is converted into a linear combination of high-dimensional features through regularization, and after substituting it into the objective function, it is converted into a kernelized objective function containing only the kernel matrix and the regularization term.

2. The geographically weighted regression modeling method based on nonlinear kernel mapping according to claim 1, characterized in that: The geochemical sample data includes: multi-dimensional element feature vectors, spatial coordinates and sample labels; the standardization process is: normalizing the sample data to eliminate dimensional differences; The geochemical sample data is , its sample feature matrix is: Formula 1: ; In formula 1, For the Sample No. element feature vector, is the total number of samples, is the feature dimension.

3. The geographically weighted regression modeling method based on nonlinear kernel mapping according to claim 1, characterized in that: The mathematical expression of the objective function is: Formula 2: ; The matrix form of the loss function is: Formula 3: ; The process of finding the optimal solution is: Formula 4: ; Formula 5: ; In formulas 2 to 5, The parameter vector to be estimated, is a matrix composed of the geochemical element characteristics of all sample points, The geochemical characteristic vector reconstructed for the target, is the spatial weight.

4. The geographically weighted regression modeling method based on nonlinear kernel mapping according to claim 3, characterized in that: The generation process of the adaptive spatial weight matrix is ​​as follows: based on the sample spatial position , constructing the spatial weight matrix of geographically weighted regression using Gaussian weights , the process is: Formula 6: ; Formula 7: ; In Equation 6 and Equation 7, The target sample and The Euclidean distance of samples, For bandwidth.

5. The geographically weighted regression modeling method based on nonlinear kernel mapping according to claim 4, characterized in that: The bandwidth The adaptive optimization process is as follows: determine the initial bandwidth through cross-validation and Akaike information criterion, select the bandwidth interval for iteration, and calculate the AUC value corresponding to the bandwidth in the iterative process. When the AUC value is ≥90%, the method is obtained.

6. The geographically weighted regression modeling method based on nonlinear kernel mapping according to claim 1, characterized in that: The process of obtaining the kernelized geographically weighted regression model is as follows: Step S4-1: Map the original features to the high-dimensional space through implicit mapping, and define the high-dimensional space loss function. The calculation process is: Formula 8: ( ), ; Formula 9: ; Substituting Equation 9 into Equation 3, we get: Formula 10: ; Formula 11: ; Step S4-2: Construct a kernel matrix based on the Gaussian kernel function and the implicit mapping described in step S4-1. , the process is: Formula 12: ; Formula 13: ; Equation 14: ; Equation 15: ; Step S4-3: According to the reproducing kernel Hilbert space theory, the nonlinear Convert it into a linear combination of high-dimensional features, and then substitute it into the objective function to convert it into a kernel-only matrix And the kernelized objective function of the regularization term, the process is: Equation 16: ; Equation 17: ; In formulas 8 to 17, Represents a nonlinear mapping relationship, is a high-dimensional feature matrix, is the high-dimensional space parameter vector, is the coefficient to be optimized, which is , is the regularization parameter.

7. The geographically weighted regression modeling method based on nonlinear kernel mapping according to claim 6, characterized in that: The high-dimensional space parameter vector The solution process is: through eigenvalue decomposition, the high-dimensional space parameters are converted into the expression of the kernel matrix, and then the decomposition result is inverted to obtain the parameter solution. The process is: Equation 18: ; Equation 19: ; The prediction formula for the target sample in the kernelized geographically weighted regression model is: Equation 20: ; or, Equation 21: ; in, is the objective function, It is a row vector consisting of the kernel function values ​​of the target sample and all training samples.

8. A geographically weighted regression model based on nonlinear kernel mapping, characterized by: Obtained by the modeling method according to any one of claims 1 to 7.

9. The application of the geographically weighted regression model based on nonlinear kernel mapping according to claim 8, characterized in that: For geochemical anomaly identification, the process is: i) Based on the predicted values ​​obtained after reconstruction by the kernelized geographically weighted regression model, the difference between the reconstructed geochemical data and the original values ​​of each sampling point is calculated. The anomaly score is defined as the sum of squares of the reconstruction errors. Then, the geochemical anomaly score of the i-th sampling point is Calculated as: Equation 22: ; ii) Using the maximum Youden index to divide the geochemical anomaly score threshold, when the geochemical anomaly score When it is higher than the threshold, it is marked as an abnormal area; iii) Using GIS technology, the anomaly identification results are superimposed with spatial data such as geological maps and structural lines to generate a visual anomaly distribution map.

Citation Information

Patent Citations

  • Stock index prediction method and device based on feature weighted support vector regression, and medium

    CN114529065A

  • Atmospheric XCO2 spatialization method based on multi-source remote sensing data

    CN118779828A