Geographic Weighting Weight Determination Method Based on Density Regression
Through the geographic weighted weight determination method based on density regression, the subjectivity and spatial heterogeneity of weight calculation in the prior art are solved, and the precise weight calculation of multiple data types is realized, and the application scope of geographic weighted regression is expanded.
Patent Information
- Application Number
- CN202310131697.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-02-15
AI Technical Summary
The existing geo-weighted regression technology is subjective in weight calculation, especially when processing line data and high-dimensional complex data, the distance definition is not accurate enough, and the anisotropy of spatial data is not effectively considered.
The geographic weighted weight determination method based on density regression is used to estimate the distance of the spatial data objectively described by spatial kernel density, and the optimal model bandwidth and weight are calculated considering its anisotropy.
It realizes accurate weight calculation of spatial data, which is suitable for a variety of data types, especially complex high-dimensional spatiotemporal data, and broadens the application scope of geographic weighted regression.
Smart Images

Figure CN116108674B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of geographic information, and particularly relates to a method for determining geographic weighted weights based on density regression. Background Art
[0002] In the application of geographic information data analysis, Geographically Weighted Regression (GWR) takes into account the geographical effect of data and realizes the analysis of spatial heterogeneity of regression coefficients through a spatial weight matrix; while Spatio-Temporal Geographically Weighted Regression (GTWR) that incorporates time information into the weight matrix based on GWR realizes the calculation of data heterogeneity in the time dimension; the above two models have been widely applied in the field of geographical analysis and achieved good results. Geographically weighted regression has been applied as an algorithm for industrial and agricultural applications in various analysis and processing processes.
[0003] However, current technological innovations in geographically weighted regression mainly focus on the improvement of distance measurement and weighting methods, with little involvement in the calculation of the weights themselves. Moreover, the spatio-temporal distance obtained by weighted averaging of time and spatial distance has a certain degree of subjectivity, and the anisotropy of spatial data is not considered in bandwidth selection; in addition, current geographically weighted regression mainly targets point data and surface data, and when dealing with spatial interaction data represented by line data or higher-dimensional complex data, the definition of distance also has a large degree of subjectivity. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for determining geographic weighted weights based on density regression for the deficiencies of the prior art. This method objectively describes the distance of spatial data through spatial kernel density estimation and takes into account the existence of its anisotropy, making geographically weighted regression more accurate.
[0005] To solve the above technical problems, the present invention adopts the following technical solutions:
[0006] A method for determining geographic weighted weights based on density regression includes the following steps:
[0007] Step 1, obtain sample data that requires geographically weighted regression;
[0008] Step 2, calculate the optimal model bandwidth according to the sample data;
[0009] Step 3, calculate the geographic weight of any sample according to the optimal model bandwidth.
[0010] Further, the sample data in Step 1 includes an independent variable X and a dependent variable y at a spatial position U, where U is a two-dimensional geographic coordinate, or time dimension information and position information in other dimensions are added.
[0011] Further, the structure of U is:
[0012]
[0013] Wherein, m represents the dimension;
[0014] And the independent variable X and the dependent variable y are organized into the following structure:
[0015]
[0016] Wherein, n represents the number of samples, and p represents the number of independent variables.
[0017] Furthermore, the method for calculating the optimal bandwidth in step 2 is as follows:
[0018] S21. Set the precision threshold e, the iteration upper limit T, and the initial bandwidth b0. Then, generate m initial points, b1, b2,..., b m , such that b i is larger than b0 on the i-th component, and the other components remain the same; set the iteration count value t = 0;
[0019] S22. Calculate the cross-validation value CV of the above m + 1 points;
[0020] S23. Sort the above points in ascending order according to their corresponding CV values CV(b i ) to obtain a new set of points If or t > T, then execute step S210; otherwise, execute step S24;
[0021] S24. Calculate the centroid c of the first m points and the reflection point r of the (m + 1)-th point with respect to the centroid c;
[0022] S25. If then calculate the expansion point s in this direction;
[0023] If CV(s) ≤ CV(r), let and continue to execute step S23; otherwise, let and execute step S22;
[0024] S26. If let and execute step S23;
[0025] S27. If let
[0026]
[0027] If CV(e1) ≤ CV(r), let and execute step S23; otherwise, execute step S29;
[0028] S28, if make
[0029]
[0030] if make And execute step S23, otherwise execute step S29;
[0031] S29, Order make And execute step S23;
[0032] S210, take b1 as the minimum point, CV(b1) as the minimum value; that is, take the model bandwidth
[0033] Furthermore, the CV(b) value with bandwidth b is calculated as:
[0034]
[0035] In the formula, n is the number of samples, y is i is the dependent variable sequence, is the estimated value of the dependent variable obtained by solving the model after excluding the target data point i;
[0036] The specific model solution calculation formula is:
[0037]
[0038]
[0039] In the formula, is the intercept term at the target point i, is the estimated value of the regression coefficient at the target point i, x ij is the observed value of the independent variable at the target point i, p represents the total number of independent variables, ε i represents the error term that follows a normal distribution; is the regression analysis coefficient vector at the target point i, that is W i is a weight matrix, the diagonal element value is the weight from each data point to the regression point; X is the independent variable observation matrix, y is the dependent variable observation vector; for each weight of W, its calculation formula is:
[0040]
[0041] Among them, u is the spatial position of the solution target, u i is the position of the i-th sample data within the bandwidth range, b is the bandwidth of the sample space, and K is the kernel function of the corresponding space.
[0042] Furthermore, the calculation formula for the centroid c is as follows:
[0043]
[0044] Furthermore, the calculation formula for the reflection point r of the (m + 1)-th point with respect to the centroid c is as follows:
[0045]
[0046] Furthermore, the calculation formula for the extended point s is as follows:
[0047]
[0048] Furthermore, the method for calculating the weight in step S3 is as follows:
[0049] According to the model bandwidth preferably obtained in step 2 Select a certain kernel function to estimate the weight w i (u) at any spatial position u, and its calculation formula is as follows:
[0050]
[0051] In the formula, n is the number of samples, and u i is the position of the i-th sample data within the bandwidth range, and K is the kernel function of the corresponding space.
[0052] Furthermore, the kernel function is Gaussian, Gaussian, and Local Periodical kernel functions.
[0053] Compared with the prior art, the beneficial effects of the present invention are as follows: Based on the definition of spatial position and spatial heterogeneity by conditional probability, the present invention uses the method of spatial kernel density estimation to calculate objective weights for different spatial positions, providing a basis for accurate geographically weighted regression, which is an extension of the geographically weighted regression method; in addition, the present invention also constructs a related bandwidth optimization method, enabling the calculation of geographically weighted weights based on density regression to be applicable to various data, especially complex high-dimensional spatio-temporal data, further broadening the application scope of geographically weighted regression. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flowchart of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0056] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0057] The present invention will be further described below in conjunction with specific embodiments, but it is not limited to the present invention.
[0058] As Figure 1 shown, the embodiments of the present invention disclose a method for determining geographical weighted weights based on density regression, including the following steps:
[0059] Step 1: Obtain sample data for geographical weighted regression;
[0060] In this step, the obtained samples are the independent variable X and the dependent variable y at the spatial position U, where U can be two-dimensional geographical coordinates, or time dimension information or position information of other dimensions can be added thereto, and its structure is:
[0061]
[0062] In the formula, m represents the dimension;
[0063] And the independent variable X and the dependent variable y are organized into the following structure:
[0064]
[0065] Among them, n represents the number of samples, and p represents the number of independent variables.
[0066] Step 2: Calculate the optimal model bandwidth according to the sample data;
[0067] The bandwidth is an important control parameter for calculating the weights of the geographical weighted regression model. It determines the rate at which the weights decay with the increase of the distance and the range of effective data points for each regression analysis point. The main goal of the bandwidth optimization process is: to find a bandwidth value such that the estimated values of the regression coefficients in all samples satisfy that the sample estimation residuals are as small as possible, and the variance of the sample estimation coefficients is as small as possible.
[0068] When the variable has multi-dimensional information, multiple bandwidths need to be calculated. At this time, the dimension m of the bandwidth b is greater than 1. To quickly obtain the optimal bandwidth value, the embodiments of the present invention adopt the following optimization algorithm steps:
[0069] A1. Set the precision threshold e and the iteration upper limit T, and set the initial bandwidth b0; then, generate m initial points, b1, b2, …, b m , such that b i is 5% larger than b0 on the i-th component, and the other components remain the same; set the iteration number value t = 0;
[0070] A2. Calculate the cross-validation (CV) values of the above m + 1 points. The calculation formula for the CV(b) value with bandwidth b is:
[0071]
[0072] In the formula, n is the number of samples, y i is the dependent variable sequence, is the estimated value of the dependent variable obtained by solving the model after excluding itself at the target data point i; among them, the specific model calculation formula is:
[0073]
[0074]
[0075] In the formula, is the intercept term at the target point i, is the estimated value of the regression coefficient at the target point i, x ij is the observed value of the independent variable at the target point i, p represents the total number of independent variables, ε i represents the error term subject to a normal distribution; is the regression analysis coefficient vector at the target point i, that is W i is the weight matrix, and the diagonal element values are the weights of each data point to the regression point; X is the matrix of observed values of independent variables, y is the vector of observed values of dependent variables. For each weight of W, its calculation formula is:
[0076]
[0077] Among them, u is the spatial position of the solution target, u i is the position of the i-th sample data within the bandwidth range, b is the bandwidth of the sample space, K is the kernel function of the corresponding space, and n is the number of samples.
[0078] A3. Sort the above points in ascending order according to their corresponding CV values CV(b i ) to obtain a new set of points If or t > T, then execute step A11, otherwise execute step A4;
[0079] A4. Calculate the centroid c of the first m points:
[0080]
[0081] A5. Calculate the reflection point r of the (m + 1)-th point with respect to the centroid c as:
[0082]
[0083] A6. If then calculate the extended point s in this direction as:
[0084]
[0085] Furthermore, if CV(s) ≤ CV(r), let and continue to execute step A3; otherwise let and execute step A2;
[0086] A7. If let and execute step A3;
[0087] A8. If let
[0088]
[0089] If CV(e1) ≤ CV(r), let and execute step A3, otherwise execute step A10;
[0090] A9. If let
[0091]
[0092] If let and execute step A3, otherwise execute step A10;
[0093] A10. Let Make and execute step A3;
[0094] A11. Take b1 as the minimum point and CV(b1) as the minimum value; that is, take the model bandwidth
[0095] This method for solving the optimal bandwidth does not require the function to be differentiable and can converge to the local minimum relatively quickly; the disadvantage is that the solution is unstable and an initial value needs to be provided; in practical operations, if a variable bandwidth in percentage form is adopted, the i-th component value with as the initial point can be selected; if a fixed bandwidth is adopted, The product of the maximum distance difference in the i-th dimension of the sample is the value of the i-th component of the initial point.
[0096] Step 3: Calculate the geographical weight of any sample according to the optimal model bandwidth.
[0097] In this step, according to the model bandwidth preferably obtained in Step 2 By selecting a certain kernel function, the weight w i (u) at any spatial position u can be estimated, and its calculation formula is:
[0098]
[0099] In the formula, n is the number of samples, and u i is the position of the i-th sample data within the bandwidth range, and K is the kernel function of the corresponding space.
[0100] By using the above method, the weight of any spatial position can be calculated. For the weight calculation of different dimensions, it can be achieved by defining different kernel functions, so it is applicable to multi-dimensional data.
[0101] To better illustrate the present invention, it is further described with a specific implementation case. In this case, the air quality index (AQI) data of domestic air quality monitoring stations is used, and according to the fact that air pollutants may be related to the average temperature and precipitation, the average temperature and precipitation data in the same period are obtained; the above data is aggregated on a monthly basis, and the average values of AQI, average temperature, and precipitation are taken as experimental data.
[0102] Step 1: Input the experimental data. The spatio-temporal data X includes geographical coordinates u, v and the number of months t corresponding to the time. The dependent variable is the AQI index data, and the independent variable data is the average temperature and precipitation.
[0103] Step 2: For the experimental data X, Gaussian, Gaussian, and LocalPeriodical kernel functions are respectively used in three dimensions to calculate the CV value, and the bandwidth is preferably selected according to the bandwidth preference method to obtain the optimal bandwidth b = [b u b v b t of the experimental data X in the u, v, t dimensions.
[0104] (1), According to step A1, set the precision threshold e = 0.1 and the iteration upper limit T = 100,000, select a fixed bandwidth, and combine the maximum sample distances of each component obtained by calculation to set the initial bandwidth value b0 = [964013.783 1144676.192 21.630]. Since the data dimension is 3, 3 initial points are generated according to b0, and the values of b1, b2, and b3 are: b1 = [1012214.472 1144676.192 21.630], b2 = [964013.783 1201910.002 21.630], b3 = [964013.783 1144676.192 22.712]; the number of iterations is t = 0;
[0105] (2), According to step A2, calculate the CV values of the above points b0, b1, b2, and b3. The calculation results are: CV(b0) = 15998727.51, CV(b1) = 15998858.33, CV(b2) = 16001371.55, CV(b3) = 16059871.12;
[0106] (3), According to step A3, reorder b0, b1, b2, and b3 according to the magnitude of the CV values in S2 to obtain Judge And t = 0 < T, continue to execute step A4;
[0107] (4), According to step A4, calculate the centroid c = [980080.679 1163754.129 21.630];
[0108] (5), According to step A5, calculate the reflection point r = [996147.575 1182832.065 20.549], and at the same time calculate CV(r) = 15929318.32;
[0109] (6), Make a judgment to get Therefore, enter step A6, calculate the expansion point s = [1012214.472 1201910.002 19.467], and calculate CV(s) = 15846178.12;
[0110] (7), Make a judgment to get CV(s) = 15846178.12 < CV(r) = 15929318.32, so let Enter the next loop, the number of iterations t = 1, and execute step A3;
[0111] (8), According to step A3, reorder according to the magnitude of the CV value to obtain a new Among them, Judge and when t = 1 < T, continue to execute step A4;
[0112] (9), According to step A4, the calculated centroid c = [990791.943 1157394.817 21.270];
[0113] (10), According to step A5, the calculated reflection point r = [1017570.104 1112879.631 20.909], and at the same time calculate CV(r) = 15950870.04;
[0114] (11), After judgment, it is obtained that Therefore, execute step A7, and let Enter the next loop, the iteration number t = 2, and execute step A3;
[0115] ……
[0116] Repeat the above process. According to the bandwidth optimization method proposed by the present invention, until the accuracy or iteration number condition is met, the optimization ends, and the optimal bandwidth value is obtained as
[0117] Step 3, for the data at a certain position, according to the calculated optimal bandwidth, calculate the weight value according to kernel density estimation. The expression is:
[0118]
[0119] where n is the number of samples, u i is the position value of the i-th sample, including the information of u, v, and t dimensions; is the bandwidth optimized in step 2, including the bandwidth values of u, v, and t dimensions, that is, b u = 1650445.998, b v = 145469.286, b t = 2.016, K is the kernel function of the corresponding dimension, that is, Gaussian, Gaussian, and Local Periodical kernel functions.
[0120] Finally, the weight matrix of all the target positions to be calculated can be calculated. Taking the calculation result of the weight matrix of Beijing in January 2016 as an example, see Table 1.
[0121] Table 1 Results of the weight matrix of the data in Beijing in January 2016
[0122]
[0123] As can be seen from the weight matrix results in Table 1, the algorithm of the present invention can assign objective values to the data weights of different cities and times according to the distance, construct the weight matrix, and provide a basic basis for subsequent geographically weighted regression analysis and modeling.
[0124] In this case, according to the above steps, the weights of all target positions to be calculated can finally be obtained according to the algorithm of the present invention, and the weight values are objectively true and reliable.
[0125] In summary, the present invention objectively calculates the weights of spatial positions based on spatial kernel density estimation, providing a basis for further geographically weighted regression; integrating the idea and method of density regression into geographically weighted regression, enabling its weight calculation to be applicable to various data, especially complex high-dimensional spatio-temporal data, further broadening the application scope of geographically weighted regression and strengthening the data analysis and application ability of geographically weighted regression.
[0126] The above are only preferred embodiments of the present invention, and do not limit the implementation manners and protection scope of the present invention. For those skilled in the art, it should be realized that all equivalent replacements and obvious changes made by using the content of the specification of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for determining geographical weighted weights based on density regression, characterized in that It includes the following steps: Step 1: Obtain the sample data for geographically weighted regression; Step 2: Calculate the optimal model bandwidth based on the sample data; Step 3: Calculate the geographical weights of any sample based on the optimal model bandwidth; Among them, the method for calculating the optimal bandwidth in Step 2 is: S21. Set the precision threshold , the iteration upper limit and the initial bandwidth , then, according to generate initial points such that the i th component is larger than , and the other components remain the same; set the initial iteration number value t = 0; S22: Calculate the cross-validation value CV of the above m + 1 points; S23. Sort the above points in ascending order according to their corresponding CV values CV(b i ) to obtain a new set of points If or , then execute step S210; otherwise, execute step S24. S24. Before calculation the centroid of the previous m points, and the reflection point of the (m + 1)-th point with respect to the centroid ; ; S25. If then calculate the extended point in the direction of the reflection point ; If , let and continue to execute step S23; otherwise, let and execute step S22; S26. If Let and execute step S23; S27. If Let ; If CV(e1) ≤ CV(r), let and execute step S23; otherwise, execute step S29; S28. If Let If Let and execute step S23; otherwise, execute step S29; S29. Let , where , so that and execute step S23; S210. Obtain as the minimum point, and CV(b1) as the minimum value; that is, obtain the model bandwidth 2. The method for determining the geographical weighted weight based on density regression according to claim 1, wherein The sample data in step 1 includes independent variables U at the spatial position X and dependent variables y, wherein 𝑼 is a two-dimensional geographical coordinate or coordinate information with time dimension information added.
3. The method for determining geographic weighted weights based on density regression according to claim 2, wherein The structure of 𝑼 is: Wherein, represents the dimension; And organize the independent variable and the dependent variable into the following structure: Among them, represents the number of samples, represents the number of independent variables.
4. The method for determining geographical weighted weights based on density regression according to claim 1, wherein The bandwidth is of The value calculation formula is: In the formula, is the number of samples, is the dependent variable sequence, is the estimated value of the dependent variable obtained by solving the model after excluding itself at the target data point ; Among them, the model solution calculation formula is: wherein, is the intercept term at the target point i , is the estimated value of the regression coefficient at the target point i , is the observed value of the independent variable at the target point i , p represents the total number of independent variables represents the error term subject to a normal distribution; is the regression analysis coefficient vector at the target point i , that is is the weight matrix, and the diagonal element values are the weights of each data point to the regression point; is the matrix of observed values of the independent variable is the vector of observed values of the dependent variable; for each weight, its calculation formula is: Among them, is the spatial position of the solution target, is the position of the -th sample data within the bandwidth range, is the bandwidth of the space where the sample is located, is the kernel function of the corresponding space.
5. The method for determining geographical weighted weights based on density regression according to claim 1, wherein Centroid The calculation formula is as follows:
6. The method for determining geographical weighted weights based on density regression according to claim 1, wherein The reflection point of the (m + 1)-th point about the centroid is calculated by the formula:
7. The method for determining geographical weighted weights based on density regression according to claim 1, wherein Expansion point The calculation formula is as follows:
8. The method for determining geographical weighted weights based on density regression according to claim 1, wherein The method for calculating the weight in Step S3 is: The model bandwidth preferably obtained according to Step 2 Select a certain kernel function to estimate the weight at any spatial position The weight , and its calculation formula is: In the formula, is the number of samples, is the position of the th sample data within the bandwidth range, is the kernel function of the corresponding space.
9. The method for determining the geographical weighted weight based on density regression according to claim 8, wherein The kernel function is Gaussian, Local Periodical kernel function.
Citation Information
Patent Citations
Method for identifying and predicting bus passenger flow influence factor based on geographically and temporally weighted regression
CN107103392A
Crime risk factor analysis method considering spatial heterogeneity and excessive discrete phenomenon
CN111860999A