Geographical weighted regression model based on nonlinear kernel mapping and modeling method and application thereof
Through the geo-weighted regression model based on nonlinear nuclear mapping, the problem of difficult processing of nonlinear features of geochemical data in the prior art is solved, and efficient and accurate identification of geochemical anomalies under complex geological conditions is achieved, especially the accurate identification of weak anomalies.
Patent Information
- Application Number
- CN202510775991.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-11
AI Technical Summary
When existing geochemical exploration methods deal with non-normal distribution of geochemical data, it is difficult to effectively capture the nonlinear synergistic mechanism between elements, resulting in insufficient abnormal recognition accuracy and recognition rate, especially in complex geological conditions, which is difficult to identify weak anomalies.
A geographic weighted regression model based on nonlinear kernel mapping is adopted, and the local linearization of geochemical data is achieved through a joint nonlinear system of kernel mapping-spatial weights. Combined with Gaussian weighting function and adaptive bandwidth optimization, we dynamically adapt to the boundary and scale differences of geological units, and improve the accuracy and recognition rate of the model.
It significantly improves the detection accuracy and stability of geochemical anomalies, can accurately identify weak anomalies, and improves the model's data utilization and spatial heterogeneity characterization capabilities.
Smart Images

Figure CN120277638A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a geographically weighted regression model, and particularly to a geographically weighted regression model based on non-linear kernel mapping, its modeling method and application, belonging to the technical field of geochemical exploration. Background Art
[0002] The shallow mineral resources on the earth are gradually exhausted, and the exploration of deep mineral resources has attracted more and more attention. However, the detection of element concentration is easily affected by complex geological conditions, and efficient exploration of mineral resources has become a challenge in the industry. In the past few decades, many methods for identifying geochemical anomalies in exploration have been proposed by scholars, which are mainly divided into two categories: those based on data frequency distribution and those based on data frequency - spatial distribution. The methods based on data frequency distribution include classical statistical methods, multivariate statistical methods (such as principal component analysis), and data exploration analysis methods. These methods assume that the data follows a normal distribution or approximately a normal distribution through transformation. However, geochemical data often shows significant non-normal distribution characteristics (such as skewness, multi-modal distribution) due to the superposition of multi-source geological processes, resulting in estimation bias of statistical parameters and limited accuracy in anomaly identification. The methods based on data frequency - spatial distribution are represented by geostatistical models (such as Kriging interpolation), fractal / multifractal models. Although they can take into account the spatial correlation and local characteristics of the data, there are two limitations: insufficient description of the multi-scale characteristics of the data (such as the scale differences in element enrichment at different ore-forming stages) and complex spatial structures (such as the discontinuous effects of faults and lithological boundaries); secondly, they are still essentially based on linear assumptions or simple non-linear transformations and are difficult to adapt to the complex non-linear cooperative mechanisms among elements during the ore-forming process (such as the coupled enrichment of trace elements and main elements, the non-linear correlation of isotope fractionation).
[0003] In response to the non-linear characteristics of geochemical data, various local anomaly identification methods such as local singularity analysis and local domain statistics have emerged one after another, focusing on local anomalies through sliding windows or spatial weight matrices. However, these methods generally have the following problems: 1) relying on artificial setting of window sizes or weight functions, with insufficient adaptive ability to spatial non-stationarity; 2) not effectively integrating non-linear feature modeling, and being prone to missing key anomaly signals when dealing with the non-linear mapping relationship of element concentrations; therefore, there is an urgent need for a high-precision and high-identification-rate adaptive geographical model. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the first object of the present invention is to provide a geographically weighted regression model based on non - linear kernel mapping. Based on the joint non - linear system of kernel mapping - spatial weights, this model realizes the technical purpose of local linearization of non - linear problems in geographical models. For the multi - element non - linear correlation and spatial heterogeneity of geochemical data, it can effectively coordinate the local spatial characteristics of the data and non - linear covariates, thereby improving the data utilization rate of the model, reducing the loss of non - linear signals, and greatly improving the accuracy and recognition rate of the model.
[0005] The second object of the present invention is to provide a method for geographically weighted regression modeling based on non - linear kernel mapping. This method realizes the accurate identification of weak geochemical anomalies under complex geological conditions by integrating the non - linear modeling of kernel functions and the spatial locality analysis of geographically weighted regression. On the one hand, this model implicitly maps the original features of geochemical elements to a high - dimensional space. While effectively capturing the complex non - linear cooperation mechanism between elements, it can also avoid directly calculating high - dimensional features explicitly, breaking through the limitations of traditional GWR linear assumptions and improving the accuracy of depicting complex data such as non - normal distributions. On the other hand, based on the local weight mechanism and adaptive bandwidth optimization of geographically weighted regression, it dynamically adapts to the geological unit boundaries and scale differences, thereby enhancing the ability to depict spatial heterogeneity.
[0006] The third object of the present invention is to provide an application of a geographically weighted regression model based on non - linear kernel mapping for geochemical anomaly identification. Based on the excellent performance of the above - mentioned model, when it is used to identify geochemical anomalies, it can significantly improve the accuracy and stability of anomaly detection, especially it can more accurately identify weak anomaly regions, providing an efficient and accurate technical means for deep - layer mineral resource exploration.
[0007] In order to achieve the above - mentioned technical objects, the present invention provides a method for geographically weighted regression modeling based on non - linear kernel mapping, which is characterized by including: Step S1: Obtain geochemical sample data containing spatial positions and construct a standardized sample set; Step S2: According to the standardized sample set, construct an initial geographically weighted regression model and solve the linear optimal solution of its parameter vector; Step S3: According to the standardized sample set, calculate local weights using a Gaussian weight function and perform adaptive bandwidth optimization to generate an adaptive spatial weight matrix for geographically weighted regression; Step S4: According to the initial geographically weighted regression model, use a non - linear kernel function to convert the original features in the sample set to a high - dimensional space through implicit mapping to obtain a kernelized geographically weighted regression model.
[0008] As a preferred solution, the geochemical sample data includes: multi - dimensional element feature vectors, spatial coordinates, and sample labels.
[0009] As a preferred solution, the standardized processing process is as follows: normalizing the sample data to eliminate the dimensional difference.
[0010] As a preferred solution, the geochemical sample data is , and its sample feature matrix is: Equation 1: ; In Equation 1, is the th sample's th element feature vector, is the total number of samples, is the feature dimension.
[0011] As a preferred solution, the solving process of the linear optimal solution of the feature vector in the initial geographically weighted regression model is as follows: obtaining the objective function at its target position by locally weighted least squares estimation of the spatially varying regression coefficients, and then transforming it into the matrix form of the loss function and taking the derivative to obtain it.
[0012] As a preferred solution, the mathematical expression of the objective function is: Equation 2: ; The matrix form of the loss function is: Equation 3: ; The solving process of the optimal solution is: Equation 4: ; Equation 5: ; In Equations 2 to 5, is the parameter vector to be estimated, is the matrix composed of the geochemical element features of all sample points, is the geochemical feature vector of the target reconstruction, is the spatial weight.
[0013] As a preferred solution, the generation process of the adaptive spatial weight matrix is: based on the sample spatial position , constructing the spatial weight matrix of the geographically weighted regression through Gaussian weights, and its process is: Equation 6: ; Equation 7: ; In Equations 6 and 7, is the Euclidean distance between the target sample and the th sample. is the bandwidth.
[0014] As a preferred solution, the adaptive optimization process of the bandwidth is as follows: determine the initial bandwidth through cross-validation and the Akaike information criterion, select a bandwidth interval for iteration, and calculate the AUC value corresponding to the bandwidth during the iteration. When the AUC value ≥ 90%, it is obtained.
[0015] As a preferred solution, the acquisition process of the kernelized geographically weighted regression model is: Step S4-1: Map the original features to a high-dimensional space through implicit mapping and define the high-dimensional space loss function. Its calculation process is: Equation 8: ( ), ; Equation 9: ; Substitute Equation 9 into Equation 3 to get: Equation 10: ; Equation 11: ; Step S4-2: According to the Gaussian kernel function and the implicit mapping described in Step S4-1, construct the kernel matrix , and its process is: Equation 12: ; Equation 13: ; Equation 14: ; Equation 15: ; Step S4-3: According to the reproducing kernel Hilbert space theory, through regularization, transform the non-linear into a linear combination of high-dimensional features. After substituting it into the objective function, transform it into a kernelized objective function containing only the kernel matrix and the regularization term. Its process is: Equation 16: ; Equation 17: ; In Equations 8~17, represents the non-linear mapping relationship, is the high-dimensional feature matrix, is the high-dimensional space parameter vector, is the coefficient to be optimized, which becomes after optimization, is the regularization parameter.
[0016] As a preferred solution, the high-dimensional space parameter vector is solved as follows: Through eigenvalue decomposition, the high-dimensional space parameter is converted into the expression form of the kernel matrix, and then the inverse is obtained using the decomposition result. The process is as follows: Equation 18: ; Equation 19: .
[0017] As a preferred solution, the prediction formula for the target sample in the kernelized geographically weighted regression model is: Equation 20: ; Or, Equation 21: ; Wherein, is the objective function, is the row vector composed of the kernel function values of the target sample and all training samples.
[0018] The present invention also provides a geographically weighted regression model based on non-linear kernel mapping, obtained by the modeling method described in any one of the above.
[0019] The present invention also provides an application of the geographically weighted regression model based on non-linear kernel mapping for geochemical anomaly identification. The process is as follows: i) According to the predicted value obtained after reconstruction of the kernelized geographically weighted regression model, calculate the difference between the reconstructed geochemical data and the original value at each sampling point. The anomaly score is defined as the sum of squares of the reconstruction errors. Then, the geochemical anomaly score at the i-th sampling point is calculated as: Equation 22: ; ii) Use the maximum Youden index to divide the threshold of the geochemical anomaly score. When the geochemical anomaly score is higher than the threshold, it is marked as an anomaly area; iii) Use GIS technology to overlay the anomaly identification result with spatial data such as geological maps and tectonic lines to generate a visualized anomaly distribution map.
[0020] Compared with the prior art, the beneficial technical effects of the technical solution of the present invention are: 1) The model provided by the present invention is based on the joint non-linear system of kernel mapping - spatial weight, achieving the technical purpose of local linearization of non-linear problems in geographical models. For the multi-element non-linear correlation and spatial heterogeneity of geochemical data, it can effectively coordinate the local spatial characteristics of the data and non-linear covariates, thereby improving the data utilization rate of the model, reducing the loss of non-linear signals, and greatly improving the accuracy and recognition rate of the model; 2) The modeling method provided by the present invention realizes the accurate identification of weak geochemical anomalies under complex geological conditions by integrating the non-linear modeling of kernel functions and the spatial locality analysis of geographically weighted regression. On the one hand, this model implicitly maps the original features of geochemical elements into a high-dimensional space, effectively capturing the complex non-linear cooperation mechanism between elements, while also avoiding directly calculating high-dimensional features explicitly, breaking through the limitations of traditional GWR linear assumptions, and improving the characterization accuracy of complex data such as non-normal distributions. On the other hand, based on the local weight mechanism of geographically weighted regression and adaptive bandwidth optimization, it dynamically adapts to the geological unit boundaries and scale differences, thereby enhancing the ability to characterize spatial heterogeneity. 3) In the technical solution provided by the present invention, based on the excellent performance of the above model, when it is used to identify geochemical anomalies, it can significantly improve the accuracy and stability of anomaly detection. In particular, it can more accurately identify weak anomaly regions, providing an efficient and accurate technical means for deep mineral resource exploration. Description of the Drawings
[0021] Figure 1 It is a line graph of AUC corresponding to different bandwidths during the adaptive optimization process of the bandwidth in Embodiment 1 of the present invention; Figure 2 It is a line graph of AUC corresponding to different regularization coefficients during the adaptive optimization process of the regularization coefficient in Embodiment 1 of the present invention; Figure 3 It is a three-dimensional graph of the change of AUC corresponding to the combination of bandwidth and regularization coefficient in Embodiment 1 of the present invention; Figure 4 It is a schematic diagram of the geochemical anomalies identified by the model obtained in Embodiment 1 of the present invention; Figure 5 It is a distribution diagram of the regression coefficients of different elements identified by the model obtained in Embodiment 1 of the present invention; Among them, Figure 5 (a) is the distribution diagram of the regression coefficient of the Ag element identified by the obtained model, Figure 5 (b) is the distribution diagram of the regression coefficient of the As element identified by the obtained model, Figure 5 (c) is the distribution diagram of the regression coefficient of the Bi element identified by the obtained model, Figure 5 (d) is the distribution diagram of the regression coefficient of the Pb element identified by the obtained model, Figure 5 (e) is the distribution diagram of the regression coefficient of the Sb element identified by the obtained model. Detailed Embodiments
[0022] The technical solution of the present invention will be further described in detail below in conjunction with specific embodiments and the accompanying drawings. To facilitate the understanding of the present invention, the present invention will be described more comprehensively and in detail below in conjunction with the specification drawings and preferred embodiments. It should be noted that the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work shall fall within the scope of protection of the present invention.
[0023] Embodiment 1
[0024] This embodiment provides a geographically weighted regression model based on non-linear kernel mapping, and its specific modeling process is as follows: Step S1. Obtain geochemical sample data containing spatial positions and construct a standardized sample set; The geochemical sample data includes: multi-dimensional element feature vectors, spatial coordinates, and sample labels.
[0025] The standardization process is: perform normalization processing on the sample data to eliminate the dimensional difference.
[0026] The geochemical sample data is , and its sample feature matrix is: Equation 1: ; In Equation 1, is the th sample's th element feature vector, is the total number of samples, is the feature dimension, and in this embodiment = 39.
[0027] Step S2. According to the standardized sample set, construct an initial geographically weighted regression model and solve the linear optimal solution of its parameter vector; The solution process of the linear optimal solution of the feature vector in the initial geographically weighted regression model is: obtain the objective function of the target position by locally weighted least squares estimation of the spatially varying regression coefficients, and then transform it into the matrix form of the loss function and take the derivative to obtain it; Specifically, the mathematical expression of the objective function is: Equation 2: ; This function is mainly used to measure the difference between the model prediction value and the true label , and the goal is to find the that minimizes ; The matrix form of the loss function is: Equation 3: ; The solution process of the optimal solution is as follows: Equation 4: ; Equation 5: ; In Equations 2 to 5, the parameter vector to be estimated, is a matrix composed of the geochemical element characteristics of all sample points, is the geochemical feature vector of the target reconstruction, is the spatial weight; Step S3. According to the standardized sample set, calculate the local weight using the Gaussian weight function and perform adaptive bandwidth optimization to generate the adaptive spatial weight matrix of the geographically weighted regression; The generation process of the adaptive spatial weight matrix is: based on the sample spatial position , construct the spatial weight matrix of the geographically weighted regression through the Gaussian weight , and its process is: Equation 6: ; Equation 7: ; In Equations 6 and 7, is the Euclidean distance between the target sample and the th sample, is the bandwidth; The adaptive optimization process of the bandwidth is: determine the initial bandwidth through cross-validation and the Akaike information criterion, select the bandwidth interval for iteration, and calculate the AUC value corresponding to the bandwidth in the iteration process. When the AUC value ≥ 90%, it is obtained; Step S4. Adopt a non-linear kernel function to transform the original features in the sample set to a high-dimensional space through implicit mapping to obtain a kernelized geographically weighted regression model; The acquisition process of the kernelized geographically weighted regression model is: Step S4-1. Map the original features to a high-dimensional space through implicit mapping and define the high-dimensional space loss function. Its calculation process is: Equation 8: ( ), ; Equation 9: ; Substitute Equation 9 into Equation 3 to get: Equation 10: ; Equation 11: ; Step S4-2: Construct a kernel matrix according to the Gaussian kernel function and the implicit mapping described in Step S4-1. , and the process is as follows: Equation 12: ; Equation 13: ; Equation 14: ; Equation 15: ; Step S4-3: According to the theory of reproducing kernel Hilbert space, through regularization, the non-linear is transformed into a linear combination of high-dimensional features. After substituting it into the objective function, it is transformed into a kernelized objective function that only contains the kernel matrix and the regularization term. The process is as follows: Equation 16: ; Equation 17: ; In Equations 8 to 17, represents the non-linear mapping relationship, is the high-dimensional feature matrix, is the high-dimensional space parameter vector, is the coefficient to be optimized, and after optimization, it is , is the regularization parameter.
[0028] The present invention also performs eigenvalue decomposition and parameter solution on the above model. The solution process of the high-dimensional space parameter vector is as follows: Through eigenvalue decomposition, the high-dimensional space parameter is converted into the expression form of the kernel matrix, and then the inverse is obtained by using the decomposition result. The process is as follows: First, since it is difficult to directly invert the high-dimensional matrix , eigenvalue decomposition is performed. Since is symmetric positive definite, this is a symmetric positive definite decomposition: Equation 18: ; where is the orthogonal eigenvector matrix, is the eigenvalue diagonal matrix. This decomposition converts the high-dimensional space parameter into a form that can be represented by the kernel matrix, avoiding explicit high-dimensional calculations, which conforms to the core idea of the kernel method of "implicit mapping"; Then, use the decomposition result to find the inverse to obtain the parameter solution: Equation 19: .
[0029] Furthermore, the prediction formula for the target sample in the kernelized geographically weighted regression model is: Equation 20: ; Or, Equation 21: ; Wherein, is the objective function, is a row vector composed of kernel function values of the target sample and all training samples.
[0030] To further illustrate the technical effects of the model provided in this embodiment, taking the gold ore concentration area in the northwestern part of Jiaodong, Shandong Province as an example, the geochemical survey values of 39 elements in the water system sediments in this area are selected as the original data. Under the computer environment (Nvidia RTX 3090Ti 24G graphics card), the geochemical anomaly extraction is carried out by using the geographically weighted regression model based on non-linear kernel mapping proposed in Embodiment 1. The specific process is as follows: First, perform an isometric log-ratio (ilr) transformation on the original data, and then input the transformed data into the model for training.
[0031] Then, determine the bandwidth and regularization coefficient. Among them, taking 1km as the bandwidth selection interval, calculate the AUC corresponding to different bandwidths, and observe the trend of the change of AUC with the bandwidth. Through Figure 1 It can be seen that within a certain range, as the bandwidth increases, the corresponding AUC will increase. Then, continue to increase the bandwidth, and the AUC does not increase but decreases. When the bandwidth is about 50km, the AUC reaches the maximum value; the same as the process of determining the bandwidth, taking 1 as the regularization coefficient selection interval, calculate the AUC corresponding to different regularization coefficients, and observe the trend of the change of AUC with the bandwidth. Through Figure 2 It can be seen that the AUC increases with the increase of the coefficient within a certain range, but starts to decline after a certain point, and reaches the maximum value when the coefficient is about 10.
[0032] Furthermore, in order to obtain the optimal combination of the bandwidth and the regularization coefficient, the present invention combines the adjustable ranges of the two into a point set matrix of 70×100, and calculates the AUC values corresponding to their combinations. The results are as Figure 3 shown. When the bandwidth is 50km and the coefficient is 10, the AUC is the highest, which is 0.914.
[0033] Then, taking the outliers interpolated from the gold points in the gold ore as positive samples, and randomly selected non-outliers as negative samples. When selecting negative samples, randomly select 100 negative sample sets, calculate the AUC corresponding to each negative sample set, and take the average value.
[0034] Based on the predicted values obtained after reconstruction by the kernelized geographically weighted regression model, calculate the difference between the reconstructed geochemical data and the original values at each sampling point. The anomaly score is defined as the sum of the squares of the reconstruction errors. Then, the geochemical anomaly score at the \(i\)-th sampling point is calculated as: Equation 22: ; The maximum Youden index is used to divide the threshold of the geochemical anomaly score. When the geochemical anomaly score is higher than the threshold, it is marked as an anomaly area.
[0035] Using GIS technology, overlay the anomaly identification results with spatial data such as geological maps and tectonic lines to generate a visual anomaly distribution map. The results are as Figure 4 shown. As can be seen from Figure 4 , most of the gold points are distributed in areas with higher anomaly values, and the anomalies are mostly distributed in the northeast and southeast along the main fault zone in the northwest of Jiaodong. The performance of the model is tested using the receiver operating characteristic curve (ROC) to measure the spatial correlation between the geochemical model and the known mineralization. In this example, the AUC of the model provided by the present invention is ≥0.91, which fully proves that the model has excellent performance and there is a strong correlation between the geochemical anomalies and the ore points.
[0036] In order to further illustrate that the model provided by the present invention can effectively overcome the problem that the existing spatial regression model cannot fully reflect the true characteristics of spatial data, and to explain the necessity of using a large amount of non-linear data and reconstructing the spatial weights in the model provided by the present invention, the present invention uses the spatial regression coefficient to detect the spatial correlation between other elements and the reconstruction of the Au element background under the above combination of the optimal bandwidth and the regularization coefficient.
[0037] Randomly select elements of Ag, As, Bi, Pb, and Sb, use the model provided in Example 1 to calculate the spatial regression coefficients of each element at the sample points, and visualize them. The results are as Figure 5 shown. As can be seen from Figure 5 , the coefficient distributions of different elements are in different regions, some regions are negative and some regions are positive, which indicates that the geochemical distribution patterns of different elements have different correlations and non-steady relationships in space, and further reflects different influences in the process of constructing the Au background value. Therefore, the model provided by the present invention based on the joint non-linear system of kernel mapping - spatial weight has more excellent authenticity, accuracy, and recognition.
Claims
1. A geographical weighted regression modeling method based on non - linear kernel mapping, characterized in that, Including: Step S1: Obtain geochemical sample data containing spatial positions and construct a standardized sample set; Step S2: Based on the standardized sample set, construct an initial geographically weighted regression model and solve the linear optimal solution of its parameter vector; Step S3: Based on the standardized sample set, calculate local weights using a Gaussian weight function and perform adaptive bandwidth optimization to generate an adaptive spatial weight matrix for geographically weighted regression; Step S4: Based on the initial geographically weighted regression model, use a non-linear kernel function to transform the original features in the sample set to a high-dimensional space through implicit mapping to obtain a kernelized geographically weighted regression model.
2. The geographically weighted regression modeling method based on non-linear kernel mapping according to claim 1, wherein: The geochemical sample data includes: multi-dimensional element feature vectors, spatial coordinates, and sample labels; the standardization process is: perform normalization processing on the sample data to eliminate dimensional differences; The geochemical sample data is , and its sample characteristic matrix is: Formula 1: ; In Equation 1, is the th element feature vector of the th sample, is the total number of samples, is the feature dimension.
3. A geographically weighted regression modeling method based on non-linear kernel mapping according to claim 1, characterized in that: The process of solving the linear optimal solution of the feature vector in the initial geographically weighted regression model is: obtain the objective function at its target position by locally weighted least squares to estimate the spatially varying regression coefficients, and then transform it into the matrix form of the loss function and take the derivative to obtain it.
4. A geographically weighted regression modeling method based on non-linear kernel mapping according to claim 3, characterized in that: The mathematical expression of the objective function is: Formula 2: ; The matrix form of the loss function is: Formula 3: ; The solution process for the optimal solution is as follows: Formula 4: ; Formula 5: ; In Formulas 2 to 5, the parameter vector to be estimated, is a matrix composed of the geochemical element characteristics of all sample points, is the geochemical feature vector for target reconstruction, is the spatial weight.
5. A geographically weighted regression modeling method based on non-linear kernel mapping according to claim 4, characterized in that: The generation process of the adaptive spatial weight matrix is as follows: Based on the sample spatial positions , construct the spatial weight matrix of geographically weighted regression through Gaussian weights , and the process is as follows: Formula 6: ; Formula 7: ; In Equation 6 and Equation 7, is the Euclidean distance between the target sample and the th sample, is the bandwidth.
6. The geographically weighted regression modeling method based on non-linear kernel mapping according to claim 5, wherein: The bandwidth described The adaptive optimization process is as follows: Determine the initial bandwidth through cross-validation and the Akaike information criterion, select a bandwidth interval for iteration, and calculate the AUC value corresponding to the bandwidth during the iteration. When the AUC value ≥ 90%, it is obtained.
7. A geographically weighted regression modeling method based on non-linear kernel mapping according to claim 1, characterized in that: The process of obtaining the kernelized geographically weighted regression model is: Step S4-1: Map the original features to a high-dimensional space through implicit mapping and define the high-dimensional space loss function, and its calculation process is: Formula 8: ( ) ; Formula 9: ; Substitute Equation 9 into Equation 3 to get: Formula 10: ; Formula 11: ; Step S4-2: Construct a kernel matrix based on the Gaussian kernel function and the implicit mapping described in Step S4-1 , and the process is as follows: Formula 12: ; Formula 13: ; Formula 14: ; Formula 15: ; Step S4-3: According to the reproducing kernel Hilbert space theory, through regularization, the non-linear is transformed into a linear combination of high-dimensional features. After substituting it into the objective function, it is transformed into a kernelized objective function that only contains the kernel matrix and the regularization term. The process is as follows: Formula 16: ; Formula 17: ; In Formulas 8 to 17, represents a non-linear mapping relationship, is a high-dimensional feature matrix, is a high-dimensional space parameter vector, is the coefficient to be optimized, which becomes after optimization, and is the regularization parameter.
8. A geographic weighted regression modeling method based on non-linear kernel mapping according to claim 7, characterized in that: The high-dimensional space parameter vector The solution process is as follows: Through eigenvalue decomposition, the high-dimensional space parameters are converted into the expression form of the kernel matrix, and then the inverse is obtained using the decomposition result, that is, the parameter solution. The process is as follows: Formula 18: ; Formula 19: ; The prediction formula for the target sample in the kernelized geographically weighted regression model is: Formula 20: ; Or, Formula 21: ; Among them, is the objective function, is a row vector composed of the kernel function values of the target sample and all training samples.
9. A geographically weighted regression model based on non-linear kernel mapping, characterized in that: Obtained from the modeling method according to any one of claims 1 to 8.
10. Application of a geographically weighted regression model based on non-linear kernel mapping according to claim 9, characterized in that: For geochemical anomaly identification, the process is: i) Calculate the difference between the reconstructed geochemical data and the original value at each sampling point based on the predicted values obtained after reconstruction using the kernelized geographically weighted regression model. The anomaly score is defined as the sum of the squares of the reconstruction errors. Then, the geochemical anomaly score at the \(i\)-th sampling point is calculated as: Formula 22: ; ii) The maximum Youden index is used to divide the geochemical anomaly score threshold. When the geochemical anomaly score is higher than the threshold, it is marked as an anomaly area; iii) Use GIS technology to overlay the anomaly identification results with spatial data such as geological maps and tectonic lines to generate a visualized anomaly distribution map.
Citation Information
Patent Citations
Geospatial outlier detection method based on multivariate adaptive regression
CN107729293A
Departure sliding time prediction method based on local weighted support vector regression
CN110766064A
Method and device for constructing geographically weighted regression model for mineral exploration
CN112712276A
Air quality prediction method based on spatio-temporal bandwidth adaptive geographically weighted regression
CN112990609A
Stock index prediction method and device based on feature weighted support vector regression, and medium
CN114529065A