Real Population Distribution Estimation Method Based on Spatial Extended Kalman Filtering Technology

Through the method based on spatially extended Kalman filtering technology, the problems of high timeliness and cost, complexity of data interpretation, difficulty in privacy protection and high technical threshold in the prior art are solved, and efficient and accurate population distribution estimation and uncertainty analysis are achieved.

CN118747266BActive Publication Date: 2025-06-10PEKING UNIV LAND & SPACE PLANNING & DESIGN INST (BEIJING) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410795481.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-19
Publication Date
2025-06-10
Estimated Expiration
2044-06-19

AI Technical Summary

Technical Problem

The prior art has problems such as high timeliness and cost, complexity of data interpretation, difficulty in privacy protection and high technical threshold in population distribution estimation.

Method used

Using a method based on spatially extended Kalman filtering technology, a geospatial data set is constructed, state variables and observation functions are defined, state transfer functions are constructed, and population distribution estimation model is optimized using spatially extended Kalman filtering algorithm, and model parameters are calibrated, and the model is verified using historical data sets.

Benefits of technology

The use of uncorrected observation data to make population distribution estimates as close to the true as possible, improves the timeliness and accuracy of the estimates, lowers the technical threshold, and provides a quantitative analysis of estimation uncertainty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118747266B_ABST
    Figure CN118747266B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for estimating real population distribution based on spatial extended Kalman filtering technology, which relates to the field of real population distribution estimation. The method includes the following steps: constructing a geospatial data set and preprocessing it to obtain a historical data set that satisfies an approximate normal distribution; defining state variables and an observation function, and constructing a state transition function based on the geospatial relationship between demographic units to initialize the population distribution estimation model; using the spatial extended Kalman filtering algorithm to optimize the population distribution estimation model and calibrate the parameters of the population distribution estimation model, and using the historical data set to verify the calibrated model; based on the verified model, inputting real-time observation data and outputting the population distribution estimation result. The present invention proposes a spatial extended Kalman filtering technology based on the extended Kalman filtering method and combines geospatial laws, and realizes its application in population estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of real population distribution estimation, and more specifically, to a method for estimating real population distribution based on spatial extended Kalman filtering technology. Background Art

[0002] Population distribution estimation plays a crucial role in fields such as urban planning, people's livelihood improvement, and economic and social research. For urban planning, accurate population data can help planners reasonably arrange infrastructure construction, such as housing, transportation, education, and medical care, to ensure the effective allocation of resources; in terms of people's livelihood improvement, understanding the population size and its distribution helps the government formulate more precise social security policies and improve the quality of life of citizens; in addition, for economic and social research, accurate population estimation data is the basis for analyzing economic development and formulating effective policies.

[0003] In the field of real population distribution estimation, existing technical methods can generally be divided into three categories: traditional statistical methods, remote sensing technology methods, and big data-based methods. Each method has specific technical characteristics and application scenarios, and at the same time, there are also their own limitations. Among them, traditional statistical methods include population census and sampling surveys. The population census requires a large amount of manpower, material resources, and financial resources, and has a long cycle, making it difficult to conduct frequently. Sampling surveys may have sampling biases, and their accuracy and representativeness are limited by the sampling design; remote sensing technology methods for estimating the population rely on indirect indicators, and the relationship between these indirect indicators and the population size needs to be established through models, which has certain uncertainties. In addition, the data processing process is complex, and in some areas, the estimation accuracy may be affected by factors such as cloud cover and image resolution; big data-based methods involve issues of personal privacy and data security, requiring strict data protection measures. Moreover, they often rely on specific user groups, which may not fully represent the overall population and there may be selection biases. In addition, the sources of big data are diverse, and the data quality is uneven, and the cleaning, integration, and analysis of data require complex technical support.

[0004] In summary, although traditional statistical methods provide basic and reliable population data, they face challenges in terms of timeliness and cost. Although remote sensing technology methods and big data-based methods can make up for the deficiencies of traditional statistical methods, they also have their own limitations, such as the complexity of data interpretation, the challenges of privacy protection, and the high technical threshold. Therefore, there is an urgent need for a method for estimating real population distribution based on spatial extended Kalman filtering technology, which can linearize non-linear functions, fuse multi-source data, and thus make as accurate an estimate of the real population distribution as possible using uncorrected observation data to support more scientific and reasonable decision-making.

[0005] In response to the problems in the related art, no effective solution has been proposed yet. Summary of the Invention

[0006] (1) Technical Problem to be Solved:

[0007] Aiming at the deficiencies of the prior art, the present invention provides a method for estimating real population distribution based on spatial extended Kalman filtering technology, which has the advantages of being able to combine geographical spatial laws, better grasp the relationship between observed data and real values, and making an estimated population distribution as close to the real one as possible using uncorrected observed data. Furthermore, it solves the challenges of traditional statistical methods in timeliness and cost in the prior art, as well as the limitations of remote sensing technology methods and big data-based methods in terms of data interpretation complexity, privacy protection difficulty, and technical threshold height.

[0008] (2) Technical Solution:

[0009] To achieve the above advantages of better grasping the relationship between observed data and real values, and thus making an estimated population distribution as close to the real one as possible using uncorrected observed data, the specific technical solution adopted by the present invention is as follows:

[0010] A method for estimating real population distribution based on spatial extended Kalman filtering technology, comprising the following steps:

[0011] S1. Construct a geospatial dataset and preprocess it to obtain a historical dataset that satisfies an approximate normal distribution;

[0012] S2. Define state variables and an observation function, and construct a state transition function based on the geospatial relationship between demographic units to initialize the population distribution estimation model;

[0013] S3. Use the spatial extended Kalman filtering algorithm to optimize the population distribution estimation model, calibrate the parameters of the population distribution estimation model, and verify the calibrated model using the historical dataset;

[0014] S4. Based on the verified model, input real-time observed data and output the estimated population distribution result.

[0015] Further, constructing a geospatial dataset and preprocessing it to obtain a historical dataset that satisfies an approximate normal distribution includes the following steps:

[0016] S11. Integrate multi-source data to establish a geospatial dataset, where the multi-source data includes geographical coordinates, observed data, auxiliary variables, and real value data;

[0017] S12. Use numerical transformation to process the geospatial dataset to satisfy an approximate normal distribution, where the expression of the numerical transformation is:

[0018] Y' = log(Y + c);

[0019] Where Y represents the original data, Y' represents the transformed data, and c represents a positive constant used to ensure that the value inside the logarithm is positive when Y is 0;

[0020] S13. Calculate the distance matrix between all demographic units based on the geographical coordinates of different demographic units.

[0021] Further, define the state variable and the observation function, and based on the geospatial relationship between demographic units, construct the state transition function. Initializing the population distribution estimation model includes the following steps:

[0022] S21. Define the state variable and establish the observation function in combination with the auxiliary variable;

[0023] S22. Based on the spatial dependence function and the spatial weight matrix, establish the state transition function;

[0024] S23. Conduct initial state estimation based on the observed data, and initialize the auxiliary variable and the error covariance matrix.

[0025] Further, based on the spatial dependence function and the spatial weight matrix, establishing the state transition function includes the following steps:

[0026] S221. Analyze the spatial connection between cities and construct the spatial dependence function. Among them, the expression of the spatial dependence function is:

[0027]

[0028] In the formula, g(x t ) represents the spatial dependence function, which is used to capture how the spatial connection with other cities affects the population growth of a city;

[0029] w ij represents the spatial connection strength between prefecture-level city i and prefecture-level city j;

[0030] x jt represents the observed population size of city j at time t;

[0031] S222. Based on the spatial adjacency weights, establish the spatial weight matrix;

[0032] S223. Based on the logarithmic model characteristics of population growth, construct the state transition function. Among them, the expression of the state transition function is:

[0033]

[0034] In the formula, x t represents the state vector;

[0035] c t represents the auxiliary variable;

[0036] u t represents the process noise;

[0037] r t represents the basic natural population growth rate;

[0038] K represents the environmental carrying capacity;

[0039] represents the impact of spatial relationships on population size changes;

[0040] B represents the coefficient matrix.

[0041] Furthermore, initial state estimation is performed based on the observed data, and the auxiliary variables and error covariance matrix are initialized, including the following steps:

[0042] S231. Perform initial state estimation based on the population observation data of the earliest year in the observed data;

[0043] S232. Initialize the auxiliary variables based on the state vector;

[0044] S233. Set the initial covariance and initialize the error covariance matrix.

[0045] Furthermore, the spatial extended Kalman filter algorithm is used to optimize the population distribution estimation model, and the parameters of the population distribution estimation model are calibrated. Verifying the calibrated model using the historical data set includes the following steps:

[0046] S31. Use the spatial extended Kalman filter algorithm to linearize the state transition function and the observation function, and define the prediction step and update step of the population distribution estimation model;

[0047] S32. Define the likelihood function, and based on the true value data, optimize the parameters of the population distribution estimation model with the goal of maximizing the likelihood function;

[0048] S33. Verify the optimized model based on the historical data set and perform sensitivity analysis.

[0049] Furthermore, using the spatial extended Kalman filter algorithm to linearize the state transition function and the observation function, and defining the prediction step and update step of the population distribution estimation model includes the following steps:

[0050] S311. Use the Jacobian matrix to linearize the state transition function and the observation function;

[0051] S312. Perform state prediction and error covariance prediction based on each time step;

[0052] S313. Calculate the Kalman gain based on each time step, and perform state update and error covariance update.

[0053] Furthermore, using the Jacobian matrix, the linearized expression of the state transition function is:

[0054]

[0055] In the formula, F t,ij is an element of the Jacobian matrix, representing the partial derivative of the state transition function with respect to the state vector;

[0056] i and j represent the indices of cities;

[0057] w ij represents the spatial weight between city i and city j;

[0058] x jt represents the population size of city j at time t;

[0059] r t represents the population growth rate;

[0060] Δt represents the time step.

[0061] Furthermore, define the likelihood function, and based on the real value data, with the goal of maximizing the likelihood function, the steps to optimize the parameters of the population distribution estimation model are as follows:

[0062] S321. Solve the likelihood function based on the observation residual and the observation noise covariance matrix, where the expression for solving the likelihood function is:

[0063]

[0064] In the formula, Z represents the observation data set;

[0065] θ represents the set of model parameters;

[0066] T represents the number of observation periods;

[0067] k represents the dimension of the observation vector;

[0068] represents the observation residual at time t, z t is the actual observed value, is the observed value predicted according to the model parameter θ;

[0069] |R| represents the determinant of R;

[0070] S322. With the goal of maximizing the likelihood function, calculate the negative logarithm of the minimization of the likelihood function;

[0071] S323. Minimize the negative logarithm based on the likelihood function and use the quasi - Newton method to solve the optimal parameters of the population distribution estimation model.

[0072] Further, based on the historical data set, verify the optimized model and conduct a sensitivity analysis, including the following steps:

[0073] S331. Split the historical data set to obtain a calibration set and a validation set;

[0074] S332. Use the validation set to verify the optimized population distribution estimation model;

[0075] S333. Based on the sensitivity analysis, evaluate the sensitivity of the output of the population distribution estimation model to different parameters.

[0076] (III) Beneficial effects:

[0077] Compared with the prior art, the present invention provides a method for estimating the real population distribution based on the spatial extended Kalman filtering technology. The beneficial effects of the present invention are as follows:

[0078] (1) Based on the extended Kalman filtering method, the present invention combines the geospatial law to propose the technology of spatial extended Kalman filtering and realizes its application in population estimation, reflecting the efficient processing ability for nonlinear systems and dynamic systems; at the same time, with the continuous progress of technology and the increasing richness of data resources, the present invention is expected to play a greater role in fields such as urban planning, people's livelihood improvement, and economic and social research, and contribute to the construction of a more accurate and efficient population management and service system.

[0079] (2) By utilizing the characteristics of the spatial extended Kalman filtering technology, the present invention can better grasp the relationship between the observed data and the true value, so as to make a population size estimate as close to the real value as possible using uncorrected observed data, supporting more scientific and reasonable decision - making; at the same time, the present invention can use new observed data to update the population estimation results in real - time, providing more timely population dynamic monitoring.

[0080] (3) By linearizing the nonlinear function, the present invention can handle complex nonlinear relationships, such as the spatial dependence between cities, which is of great significance for understanding the population distribution pattern and the interaction between cities; in addition, the present invention can integrate data from different sources, such as population census, remote sensing images, social media, and mobile phone data, by setting auxiliary variables, and improve the accuracy and comprehensiveness of the estimation through effective data fusion.

[0081] (4) The present invention provides a quantitative analysis of the estimation uncertainty, providing important reference information for policy - making and decision - making. Description of the Drawings

[0082] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0083] Figure 1 It is a schematic flowchart of a method for estimating the real population distribution based on the spatial extended Kalman filtering technology according to an embodiment of the present invention;

[0084] Figure 2 It is a flowchart framework of a method for estimating the real population distribution based on the spatial extended Kalman filtering technology according to an embodiment of the present invention;

[0085] Figure 3 It is a specific implementation diagram of a method for estimating the real population distribution based on the spatial extended Kalman filtering technology according to an embodiment of the present invention. Detailed implementation manners

[0086] To further illustrate each embodiment, the present invention provides drawings. These drawings are a part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be combined with the relevant descriptions in the specification to explain the operating principle of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are usually used to represent similar components.

[0087] According to an embodiment of the present invention, a method for estimating the real population distribution based on the spatial extended Kalman filtering technology is provided.

[0088] Now, the present invention will be further described in combination with the drawings and specific implementation manners. As Figures 1-3 shown, a method for estimating the real population distribution based on the spatial extended Kalman filtering technology according to an embodiment of the present invention includes:

[0089] S1. Construct a geospatial dataset and preprocess it to obtain a historical dataset that satisfies an approximate normal distribution;

[0090] S2. Define state variables and an observation function, and based on the geospatial relationship between demographic units, construct a state transition function to initialize the population distribution estimation model;

[0091] S3. Use the spatial extended Kalman filtering algorithm to optimize the population distribution estimation model, calibrate the parameters of the population distribution estimation model, and verify the calibrated model using the historical dataset;

[0092] S4. Based on the verified model, input real-time observation data and output the population distribution estimation result.

[0093] Specifically, based on the calibrated and verified population distribution estimation model, real-time observation data is input, the population distribution estimation result is output, and the output result is evaluated, and finally it is applied to the analysis and application of decision-making results.

[0094] Specifically, based on the calibrated and tested model, using the current latest population observations, the latest true population sizes of each statistical unit are predicted through the S-EKF model.

[0095] Specifically, combining the subjective experience of the operator, a detailed analysis is carried out on the estimated true population sizes of each prefecture-level city provided by the S-EKF, including the accuracy of the estimation, the trend of the time series, and possible outliers.

[0096] Specifically, based on the accurate population estimation results, it is used to support decision-making in multiple aspects such as urban planning, resource allocation, and policy formulation.

[0097] In one embodiment, constructing a geospatial dataset and preprocessing it to obtain a historical dataset that satisfies an approximate normal distribution includes the following steps:

[0098] S11. Integrate multi-source data to establish a geospatial dataset, where the multi-source data includes geographical coordinates, observation data, auxiliary variables, and true value data;

[0099] Specifically, data preparation includes geographical coordinate verification, observation data collection, auxiliary variable construction, and true value data acquisition, and the above data is preprocessed. Among them, geographical coordinates are used to ensure that each sample point has an associated (x, y) coordinate; observation data is used to collect current and historical observation data (y t ) for each sample point; auxiliary variables are used to collect current and historical auxiliary variable data (c t ) for each sample point, which can improve the effect of filtering calculation and effectively improve the estimation accuracy of the true population. Auxiliary variables generally use variables in the city that are correlated with the population size, such as GDP, night light brightness, construction land scale, road network density, etc.; true value data is used to prepare a historical true value dataset, which is specifically used to calibrate the model.

[0100] S12. Use numerical transformation to process the geospatial dataset to satisfy an approximate normal distribution, where the expression of the numerical transformation is:

[0101] Y' = log(Y + c);

[0102] In the formula, Y represents the original data, Y' represents the transformed data, and c represents a positive constant, which is used to ensure that the value inside the logarithm is positive when Y is 0;

[0103] Specifically, since urban data (such as population size, GDP, night-time light brightness, etc.) generally do not satisfy the normal distribution but follow a power-law distribution, in order to meet the normal distribution assumption of the S-EKF, the data can be appropriately transformed. This transformation aims to make the data distribution closer to the normal distribution, thereby improving the accuracy and reliability of model estimation.

[0104] Specifically, in this embodiment, logarithmic transformation is used to achieve the purpose of approximately satisfying the normal distribution of the data.

[0105] Specifically, after the transformation, the Shapiro-Wilk test is used to perform a normality test on the results. If the results meet the normal distribution characteristics, they can be used for subsequent calculations.

[0106] S13. Calculate the distance matrix between all demographic units based on the geographical coordinates of different demographic units.

[0107] Specifically, calculate the distance matrix between all statistical units based on the geographical coordinates of the demographic units for spatial analysis. The great circle distance D is calculated using the Haversine Formula, and the expression is:

[0108]

[0109] In the formula, Δφ represents the i-1 , P i latitude difference between;

[0110] Δλ represents the i-1 , P i longitude difference between;

[0111] φ 1 and φ 2 respectively represent the i-1 latitude of P i and P;

[0112] r represents the radius of the earth (average radius = 6371 km).

[0113] In one embodiment, defining the state variables and the observation function, and based on the geospatial relationship between the demographic units, constructing the state transition function, and initializing the population distribution estimation model includes the following steps:

[0114] S21. Define the state variables and establish the observation function in combination with the auxiliary variables;

[0115] S22. Based on the spatial dependence function and the spatial weight matrix, establish the state transition function;

[0116] S23. Perform initial state estimation based on the observed data, and initialize the auxiliary variables and the error covariance matrix.

[0117] Specifically, define the non - linear state transition and observation functions, including defining the state variables, the observation function, and the state transition function, and introduce the geospatial relationship between statistical units in this process.

[0118] Specifically, the state variables include the state vector (x t ) and the observation vector (z t ). The state vector (x t ) is the population size of each prefecture - level city, and the observation vector (z t ) contains the observed population data of each prefecture - level city at time t.

[0119] Specifically, the observation function reflects how to obtain the observation vector from the state vector, and also considers the direct influence of the auxiliary variables. The expression is:

[0120] h(x t , c t ) = Dx t + Ec t + v t ;

[0121] In the formula, Dx t represents the direct influence of the state vector x t on the observed value;

[0122] Ec t represents the direct influence of the auxiliary variable on the observed value, where E is the coefficient matrix;

[0123] v t represents the observation noise.

[0124] In one embodiment, based on the spatial dependence function and the spatial weight matrix, establishing the state transition function includes the following steps:

[0125] S221. Analyze the spatial connection between cities, and construct the spatial dependence function. Among them, the expression of the spatial dependence function is:

[0126]

[0127] In the formula, g(x t ) represents the spatial dependence function, which is used to capture how the spatial connection with other cities affects the population growth of a city;

[0128] w ij represents the spatial connection strength between prefecture - level city i and prefecture - level city j;

[0129] x jtDenote the observed population size of city j at time t;

[0130] Specifically, g(x t ) represents the spatial dependence function. In this embodiment, the logarithmic spatial dependence function is adopted. It is expected that the population interaction effect between cities is related to the logarithm of the population size. This form assumes that the marginal impact of the population size decreases with the increase of the population size.

[0131] S222. Based on the spatial adjacency weights, establish a spatial weight matrix;

[0132] Specifically, W is the spatial weight matrix, where each element w of the matrix ij represents the spatial connection strength between prefecture-level city i and prefecture-level city j. In this embodiment, the spatial adjacency weights are adopted. If the statistical units i and j are adjacent, then w ij = 1, otherwise 0.

[0133] S223. Based on the logarithmic model characteristics of population growth, construct a state transition function. Among them, the expression of the state transition function is:

[0134]

[0135] In the formula, x t represents the state vector;

[0136] c t represents the auxiliary variable;

[0137] u t represents the process noise;

[0138] r t represents the basic natural population growth rate;

[0139] K represents the environmental carrying capacity;

[0140] represents the influence of spatial relationship on the change of population size;

[0141] B represents the coefficient matrix.

[0142] Specifically, considering the logarithmic model characteristics of population growth, it is assumed that the population growth rate changes with the change of the population size and is affected by the auxiliary variable and spatial relationship. A state transition function is constructed. In the expression of this function, represents the influence of spatial relationship on the change of population size. Here, it is assumed that the spatial relationship influence is adjusted through the function g(x t ) of the population size x t ), and Bc t represents the influence of the auxiliary variable on the change of population size, and B represents the coefficient matrix.

[0143] In one embodiment, the initial state estimation is performed based on the observed data, and the initialization of the auxiliary variables and the error covariance matrix includes the following steps:

[0144] S231. Perform an initial state estimation based on the population observation data of the earliest year in the observed data;

[0145] S232. Initialize the auxiliary variables based on the state vector;

[0146] S233. Set the initial covariance and initialize the error covariance matrix.

[0147] Specifically, the initialization of the population distribution estimation model includes the initial state estimation, the initialization of the auxiliary variables, and the initialization of the covariance matrix.

[0148] Specifically, the initial state estimation Take the population observation data of the earliest year in the available observed data as the basis of the initial state. If the state vector x t includes the population size of each prefecture-level city, then the expression for the initial state estimation is:

[0149]

[0150] In the formula, P 2018,i represents the observed population size of the i-th prefecture-level city in the earliest year, and n is the total number of prefecture-level cities.

[0151] Specifically, for the initialization of the auxiliary variables, for the auxiliary variables (such as GDP growth rate, night light brightness, etc.) included in the state vector, their initial estimates can be obtained based on the relevant data of the earliest year, and the expression is:

[0152]

[0153] In the formula, c 2018,j represents the observed value of the j-th auxiliary variable in the earliest year, and m is the number of auxiliary variables.

[0154] Specifically, for the initialization of the error covariance matrix, the initial error covariance matrix represents the quantification of the uncertainty of the initial state estimation. In this embodiment, by setting a relatively large initial covariance to reflect the uncertainty of the initial state estimation, the expression is:

[0155] P 0∣0 = σ 2 I;

[0156] In the formula, σ 2 is a relatively large constant, and I is the identity matrix, whose size is the same as that of the state vector x thave the same dimension. This setting assumes that the uncertainty of the initial estimate is the same in all dimensions and that the state variables are independent; in practical applications, the specific value of σ 2 can be adjusted according to the availability and quality of the data. If you have high confidence in the initial estimate of certain variables, you can assign a smaller initial covariance to these variables. Conversely, if you are less certain about the initial estimate of certain variables, a larger initial covariance should be assigned.

[0157] In one embodiment, the spatially extended Kalman filter algorithm is used to optimize the population distribution estimation model and calibrate the parameters of the population distribution estimation model. Verifying the calibrated model using the historical dataset includes the following steps:

[0158] S31. Use the spatially extended Kalman filter algorithm to linearize the state transition function and the observation function, and define the prediction step and the update step of the population distribution estimation model;

[0159] S32. Define the likelihood function, and based on the true value data, optimize the parameters of the population distribution estimation model with the goal of maximizing the likelihood function;

[0160] S33. Verify the optimized model based on the historical dataset and perform a sensitivity analysis.

[0161] In one embodiment, using the spatially extended Kalman filter algorithm to linearize the state transition function and the observation function, and defining the prediction step and the update step of the population distribution estimation model includes the following steps:

[0162] S311. Use the Jacobian matrix to linearize the state transition function and the observation function;

[0163] Specifically, using the S-EKF method, set up each link of the EKF model. First, linearize the process function and the observation function, and then define the prediction step and the update step respectively.

[0164] Specifically, for each time step t, calculate the Jacobian matrices F at t and H t of the state transition function f and the observation function h. This is the key to the S-EKF for dealing with nonlinear systems, that is, through the first-order Taylor expansion approximation. The Jacobian matrix is a matrix composed of partial derivatives with respect to the state vector, and it is used in S-EKF to linearize nonlinear functions.

[0165] Specifically, the Jacobian matrix plays a crucial role in S-EKF. The Jacobian matrix is a concept in mathematics used to describe the derivative (or partial derivative) of a vector-valued function with respect to another vector. In S-EKF, this concept is used to process and linearize the state transition and observation models of nonlinear models.

[0166] Specifically, assume the state vector is x t , then the Jacobian matrix F of the state transition function f with respect to x t can be expressed as: t

[0167]

[0168] In the formula, represents the partial derivative of the function state transition function f with respect to its state vector x t . This matrix is used to approximate the linear behavior of the function near the current state.

[0169] Specifically, for the state transition function with spatial dependence, the expression is:

[0170] f(x t ) = x t + Δt·(r t ·x t + W·g(x t ));

[0171] In the formula, g(x t ) is the spatial dependence function between cities. The calculation of the Jacobian matrix will involve the partial derivatives of x t and g(x t ).

[0172] In one embodiment, using the Jacobian matrix, the linearized expression of the state transition function is:

[0173]

[0174] In the formula, F t,ij is an element of the Jacobian matrix, representing the partial derivative of the state transition function with respect to the state vector;

[0175] i and j represent the indices of cities;

[0176] w ij represents the spatial weight between city i and city j.

[0177] x jt represents the population size of city j at time t.

[0178] r t represents the population growth rate;

[0179] Δt represents the time step.

[0180] Specifically, the present invention adopts logarithmic spatial dependence, and the final Jacobian matrix is as shown in the above expression.

[0181] ​Specifically, for the case of diagonal elements (i = j), in addition to the basic growth rate r t we also need to add the influence of all other cities on city i. For non-diagonal elements (i ≠ j), only the direct influence of city j on city i is considered. This expression can accurately reflect the influence of the interaction between the populations of each city when considering spatial dependence.

[0182] Specifically, assume the observation vector is z t , then the Jacobian matrix H of the observation function h with respect to x t can be expressed as: t

[0183]

[0184] If the observation function h directly maps the state variable x t to the observation space, for example:

[0185] h(x t , c t ) = Dx t + Ec t + v t ;

[0186] where D is a mapping matrix, then the Jacobian matrix H t is D.

[0187] S312. Perform state prediction and error covariance prediction based on each time step;

[0188] Specifically, for each time step t, the expression for state prediction is:

[0189]

[0190] The expression for error covariance prediction is:

[0191]

[0192] In the formula, Q t represents the process noise covariance matrix.

[0193] ​Specifically, in the S-EKF, the process noise covariance matrix Q is a key parameter of the model. It represents the uncertainty in the state transition process or the magnitude of the system's internal noise. Correctly setting Q is crucial for ensuring the performance of the S-EKF. According to the specific data and requirements mentioned above, the steps to set the process noise covariance matrix Q in the EKF are as follows: Assume we have a historical dataset including the population growth rate and auxiliary variables (such as the GDP growth rate). First, calculate the volatility of each variable during the historical period, i.e., the standard deviation, to estimate the magnitude of the process noise. For the population growth rate, calculate the average population growth rate

[0194]

[0195] where r t is the population growth rate in the t-th year and N is the total number of years.

[0196] Specifically, the expression for calculating the standard deviation (σ r ) of the population growth rate is:

[0197]

[0198] The expression for constructing the process noise covariance matrix Q is:

[0199]

[0200] Here, each diagonal element represents the process noise variance of the corresponding state variable. For example, if the standard deviation σ r of the population growth rate is 0.01 and the standard deviation σ g of the GDP growth rate is 0.02, then the diagonal elements of the Q matrix will be set to the squares of these values to reflect the magnitudes of these process noises.

[0201] S313. Based on each time step, calculate the Kalman gain and perform state update and error covariance update.

[0202] Specifically, for each time step t, calculate the Kalman gain, and the expression is:

[0203]

[0204] where R tis the observation noise covariance matrix. In the S-EKF framework, the observation noise covariance matrix R represents the uncertainty or noise level in the observation process. Correctly setting R is crucial for ensuring the accuracy of S-EKF estimation. According to the specific data provided (including historical population observation data, GDP, night light brightness, and other auxiliary variable data) within a certain time period and requirements, the steps to set the observation noise covariance matrix R in S-EKF are as follows: First, evaluate the uncertainty of historical observation data. Analyze the historical observation data to evaluate its uncertainty. For population data and auxiliary variables (such as GDP, night light brightness, road network density), calculate their volatility or standard deviation during the historical period. For example, the standard deviation of each variable can be calculated using the following formula:

[0205]

[0206] where, Z i,t is the observed value of variable i at time t, is the average observed value of variable i, and N is the number of observation periods.

[0207] Specifically, then set the diagonal elements of R. Based on the calculated standard deviation, initial values can be set for the diagonal elements of R. If the observation noise is assumed to be isotropic and independent, then R can be set as a diagonal matrix, where each diagonal element is an estimate of the noise variance of the corresponding observed variable:

[0208]

[0209] In the above embodiment, if the volatility (standard deviation) of population data within a certain time period is 1000 people, and the volatility of GDP is 2% (assuming a standard deviation of 0.02), then these values can be directly used as the initial values of the corresponding elements of the R matrix.

[0210] Specifically, the expression for state update is:

[0211]

[0212] The expression for error covariance update is:

[0213] P t∣t =(I - K t H t )P t∣t-1 ;

[0214] Through the above steps, S-EK can update the population size estimate of each prefecture-level city according to the latest observation data at each time step, and take into account the non-linear characteristics of population growth. This process relies on the linear approximation of non-linear functions, which is carried out through the Jacobian matrix, enabling the algorithm to adapt to the characteristics of non-linear systems.

[0215] In one embodiment, a likelihood function is defined, and based on the true value data, with the goal of maximizing the likelihood function, optimizing the parameters of the population distribution estimation model includes the following steps:

[0216] S321. Solve the likelihood function based on the observation residuals and the observation noise covariance matrix. Among them, the expression for solving the likelihood function is:

[0217]

[0218] In the formula, Z represents the observation data set;

[0219] θ represents the set of model parameters;

[0220] T represents the number of observation periods;

[0221] k represents the dimension of the observation vector;

[0222] represents the observation residual at time t, z t is the actual observed value, is the observed value predicted according to the model parameter θ;

[0223] |R| represents the determinant of R;

[0224] S322. With the goal of maximizing the likelihood function, calculate the minimized negative logarithm of the likelihood function;

[0225] S323. Based on the minimized negative logarithm of the likelihood function, use the quasi - Newton method to solve the optimal parameters of the population distribution estimation model.

[0226] Specifically, in practical applications, the parameters of the model, such as the natural population growth rate r tParameters such as the environmental carrying capacity \(K\) need to be estimated through data. The estimation of these parameters can be carried out outside the framework of the S-EKF. Based on real historical population data, the maximum likelihood estimation (MLE) method is used to estimate the parameters of the model with non-linear parameters proposed in the present invention; using the maximum likelihood estimation (MLE) method to calibrate the S-EKF means that a set of values of the model parameters needs to be found, and this set of values can maximize the probability of the observed data appearing under the given model. The model parameters considered here may include the population growth rate, the environmental carrying capacity, and the influence coefficients of auxiliary variables (such as GDP growth rate, night light brightness, etc.) on population growth. The detailed steps for calibrating the S-EKF model based on the maximum likelihood estimation method are as follows: First, define the likelihood function, that is, the probability of the observed data appearing when the model parameters are given. For the S-EKF model, the likelihood function is usually defined based on the observation residuals (the difference between the observed value and the model prediction value) and the observation noise covariance matrix \(R\). Assuming that the observation residuals follow a multivariate normal distribution, the expression of the likelihood function \(L\) is obtained.

[0227] Specifically, the goal of calibrating the model is to find a set of parameters \(\theta\) that maximizes the likelihood function \(L\), which is equivalent to minimizing the negative log-likelihood function (NLL), and the expression is:

[0228] \(NLL(\theta|X)=-\ln L(\theta|Z)\);

[0229] Substitute the likelihood function and simplify to obtain:

[0230]

[0231] Specifically, the quasi-Newton method is used to find the parameter \(\theta\) that minimizes the NLL. This involves calculating the derivative (gradient) and second derivative (Hessian) of the NLL with respect to \(\theta\), and iteratively updating \(\theta\) until convergence. The specific steps are as follows: First, perform initialization, select an initial parameter estimate \(\theta\) (0) , set the approximation \(H\) of the initial Hessian matrix (0) , usually the identity matrix, and set the iteration count \(k = 0\).

[0232] Specifically, then repeat the following steps until convergence:

[0233] a. Calculate the gradient: Calculate the gradient of the negative log-likelihood function at the current parameter \(\theta\) (k)

[0234] b. Calculate the search direction: Determine the descent direction \(p\) (k) \(=-H\) (k) \(g\) (k) , where \(H\) (k) is the approximation of the Hessian matrix at the \(k\)-th step;

[0235] c. Line search: Select a step size α (k) , such that after moving along the direction p (k) , the NLL function value decreases the most. This can be achieved through a line search algorithm, such as backtracking line search;

[0236] d. Update parameters: Update the parameter θ according to the step size (k+1) = θ (k) + α (k) p (k) ;

[0237] e. Update the Hessian approximation: Use the selected quasi-Newton method to update the approximation H of the Hessian matrix (k+1) . In the present invention, the BFGS (Broyden-Fletcher-Goldfarb-Shanno) method is used to update the Hessian matrix, and the update formula is as follows:

[0238] s (k) = θ (k+1) - θ (k)

[0239] y (k) = g (k+1) - g (k)

[0240]

[0241] H (k+1) = (I - ρ (k) s (k) y (k)T ) H (k) (I - ρ (k) y (k) s (k)T ) + ρ (k) s (k) s (k)T

[0242] where I is the identity matrix, s (k) is the parameter update amount, y (k) is the gradient update amount, and ρ (k) is a scalar;

[0243] f. Check for convergence: If ||g (k+1) || is small enough or ||θ (k+1) - θ (k) || is small enough, or the maximum number of iterations is reached, then stop the iteration.

[0244] Specifically, finally output the final parameter θ * as the optimal solution.

[0245] In one embodiment, based on the historical data set, the optimized model is verified and sensitivity analysis is performed, including the following steps:

[0246] S331. Split the historical data set to obtain a calibration set and a validation set;

[0247] S332. Use the validation set to verify the optimized population distribution estimation model;

[0248] S333. Based on sensitivity analysis, evaluate the sensitivity of the output of the population distribution estimation model to different parameters.

[0249] Specifically, for the implementation and verification of the model, first, the results of model calibration are verified using the validation set data, and then sensitivity analysis is performed.

[0250] Specifically, after parameter estimation and model calibration are completed, an independent data set should be used to verify the model to evaluate its prediction accuracy, which can be done by comparing the error between the prediction results of the model and the actual observations.

[0251] Specifically, first perform data splitting, randomly dividing the real historical population data into two parts:

[0252] Calibration set (training set): accounting for 70% of the total data, used for parameter calibration of the model;

[0253] Validation set (test set): accounting for 30% of the total data, used for checking the calibration results.

[0254] Specifically, use the validation set data to evaluate the prediction performance of the model, calculate the difference between the predicted value of the model on the validation set and the actual observation value. Common evaluation metrics include mean squared error (MSE), root mean squared error (RMSE), and coefficient of determination (R 2 ), and the expression for calculating the mean squared error (MSE) is:

[0255]

[0256] In the formula, T test is the number of time points in the validation set, is the population prediction value of the model for time point t, and P t is the actual observation value.

[0257] Specifically, the expression for the root mean squared error (RMSE) is:

[0258]

[0259] The expression for the coefficient of determination (R 2 ) is:

[0260]

[0261] In the formula, is the average value of all observations in the validation set.

[0262] Specifically, according to the performance of the model on the validation set, analyze the accuracy and reliability of the model prediction. If the prediction performance of the model does not meet the expectations, it is necessary to return to the model calibration stage, adjust the model structure or re-estimate the parameters.

[0263] Specifically, conduct a sensitivity analysis to evaluate the sensitivity of the model output to different parameters, which helps to identify which parameters have the greatest impact on the model prediction results, so as to provide a basis for further research and policy making. First, determine the analysis objective, determine the output of interest, and select one or more model outputs for sensitivity analysis. In this embodiment, it is the estimated true population value of each current statistical unit. Then select the key parameters and determine the parameters that may have a significant impact on the model output, such as the population growth rate, environmental carrying capacity, GDP growth rate, etc.

[0264] Specifically, when designing the experiment, the present invention adopts the single-parameter variation method, that is, change each selected parameter one by one, keep other parameters unchanged, observe the change of the model output, and then execute the experiment. For each set of parameter settings, run the model and record the corresponding output. This patent adopts the parameter scanning method, that is, fix other parameters, select one parameter to vary within a certain range, perform the model operation step by step, and finally analyze the results.

[0265] Specifically, first conduct a quantitative analysis. Based on the sensitivity index, calculate the sensitivity index of the model output to the change of each parameter. For example, for the parameter θ i , the sensitivity index can be expressed as:

[0266]

[0267] In the formula, Y is the model output, Δθ i is the change amount of the parameter, and ΔY is the change amount of the output caused by the parameter change.

[0268] Specifically, for the normalized sensitivity index, in order to compare between different parameters, the normalized sensitivity index can be calculated. Considering the dimensions of the parameters and the output, the obtained expression is:

[0269]

[0270] Specifically, secondly, qualitative analysis is carried out to analyze the influence mode of parameter changes on the model output, such as linear, non-linear or threshold effects. Finally, the sensitivity analysis results are applied to improve the model. According to the sensitivity analysis results, parameters that have a greater impact on the model output can be identified and focused on, for more accurate estimation or collection of more data, and for decision support to help decision-makers understand the degree of influence of different parameters on the prediction results, so as to make more informed decisions in the face of uncertainty.

[0271] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for estimating the real population distribution based on spatial extended Kalman filtering technology, characterized in that: include: S1. Construct and preprocess the geospatial data set to obtain a historical data set that satisfies an approximate normal distribution; S2, define state variables and observation functions, and construct state transfer functions based on the geospatial relationship between demographic units to initialize the population distribution estimation model; S3. Use the spatial extended Kalman filter algorithm to optimize the population distribution estimation model, calibrate the parameters of the population distribution estimation model, and verify the calibrated model using historical data sets; S4, based on the verified model, input real-time observation data and output population distribution estimation results; The S1 includes: S11, integrating multi-source data to establish a geospatial data set, wherein the multi-source data includes geographic coordinates, observation data, auxiliary variables and real value data; S12. Process geospatial datasets using numerical transformations to satisfy an approximate normal distribution; S13, calculating a distance matrix between all demographic units based on the geographic coordinates of different demographic units; The S2 includes: S21, define state variables, and establish observation functions in combination with auxiliary variables; S22, establishing a state transfer function based on the spatial dependency function and the spatial weight matrix; S23, performing initial state estimation based on the observed data, and initializing auxiliary variables and error covariance matrix; The S21 includes: S221. Analyze the spatial connection between cities and construct a spatial dependency function, wherein the expression of the spatial dependency function is: ; In the formula, represents the spatial dependence function, which is used to capture how the spatial connections with other cities affect the population growth of a city; Indicates prefecture-level city and prefecture-level cities The strength of the spatial connection between Indicates the city j exist t Observation population size at time; S222, establishing a spatial weight matrix based on the spatial adjacency weight; S223. Construct a state transfer function based on the logarithmic model characteristics of population growth.

2. The method for estimating real population distribution based on spatial extended Kalman filtering technology according to claim 1, characterized in that: The expression of the numerical transformation is: ; In the formula, Represents the original data, represents the transformed data, c Represents a positive constant, used to ensure that When it is 0, the value in the logarithm is positive; The expression of the state transfer function is: ; In the formula, represents the state vector; Represents auxiliary variables; represents process noise; represents the natural growth rate of the basic population; K Indicates environmental carrying capacity; Represents the impact of spatial relationships on changes in population size; B represents the coefficient matrix.

3. The method for estimating real population distribution based on spatial extended Kalman filtering technology according to claim 1, characterized in that: The initial state estimation based on the observed data and initializing the auxiliary variables and the error covariance matrix include: S231. Estimation of the initial state based on the population observation data of the earliest year among the observation data; S232, initializing auxiliary variables based on the state vector; S233. Set the initial covariance and initialize the error covariance matrix.

4. The method for estimating real population distribution based on spatial extended Kalman filtering technology according to claim 1, characterized in that: The method of optimizing the population distribution estimation model by using the spatial extended Kalman filter algorithm, calibrating the parameters of the population distribution estimation model, and verifying the calibrated model by using the historical data set includes: S31. Using the spatial extended Kalman filter algorithm, linearize the state transfer function and the observation function, and define the prediction step and the update step of the population distribution estimation model; S32, defining a likelihood function, and optimizing the parameters of the population distribution estimation model based on the real value data with the goal of maximizing the likelihood function; S33. Based on the historical data set, verify the optimized model and conduct sensitivity analysis.

5. The method for estimating real population distribution based on spatial extended Kalman filtering technology according to claim 4, characterized in that: The method of using the spatial extended Kalman filter algorithm to linearize the state transfer function and the observation function and define the prediction step and the update step of the population distribution estimation model includes: S311. Linearize the state transfer function and the observation function using the Jacobian matrix; S312, performing state prediction and error covariance prediction based on each time step; S313. Based on each time step, calculate the Kalman gain, and perform state update and error covariance update.

6. The method for estimating real population distribution based on spatial extended Kalman filtering technology according to claim 5, characterized in that: The expression of the linearized state transfer function using the Jacobian matrix is: ; In the formula, is the element of the Jacobian matrix, representing the partial derivative of the state transfer function with respect to the state vector; and An index representing a city; Indicates the city With the city The spatial weight between Indicates time Time City population size; represents the population growth rate; Represents the time step.

7. The method for estimating real population distribution based on spatial extended Kalman filtering technology according to claim 4, characterized in that: The definition of the likelihood function and the optimization of the population distribution estimation model parameters based on the real value data and maximizing the likelihood function as the goal include: S321, based on the observation residual and the observation noise covariance matrix, solving the likelihood function, wherein the expression for solving the likelihood function is: ; In the formula, represents an observational dataset; Represents a collection of model parameters; represents the number of observation periods; represents the dimension of the observation vector; Indicates at time The observed residuals, is the actual observed value, According to the model parameters The predicted observed values; represents the observation noise covariance matrix The determinant of ; S322, taking the maximization of the likelihood function as the goal, calculating the minimized negative logarithm of the likelihood function; S323. Based on minimizing the negative logarithm of the likelihood function, the quasi-Newton method is used to solve the optimal parameters of the population distribution estimation model.

8. The method for estimating real population distribution based on spatial extended Kalman filtering technology according to claim 4, characterized in that: The process of verifying the optimized model based on the historical data set and conducting sensitivity analysis includes: S331, split the historical data set to obtain a calibration set and a validation set; S332. Use the validation set to validate the optimized population distribution estimation model; S333. Based on sensitivity analysis, evaluate the sensitivity of the output of the population distribution estimation model to different parameters.