Crime factor correlation analysis method considering spatial heterogeneity and zero expansion phenomenon
Through the zero-inflated negative binomial model and latent variable weight optimization, the zero-inflation and spatial heterogeneity problems of drug crime data were solved, the impact of environmental factors on cases was quantified, and high-precision crime risk assessment and prevention and control strategy optimization were achieved.
Patent Information
- Application Number
- CN202510695960.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-23
AI Technical Summary
When processing drug crime case data, existing technologies ignore zero inflation and spatial heterogeneity, resulting in parameter estimation bias, inability to accurately assess the impact of environmental factors, and failure to quantify the symbiotic response of drug use and drug trafficking cases.
A zero-inflated negative binomial model combined with latent variable weights is used. The parameters are optimized alternately through the expectation maximization algorithm and the Newton-Raphson method to distinguish between structural zero values and random zero values, quantify the differentiated impact of environmental factors on cases, and generate a hotspot distribution map combined with kernel density estimation.
It has achieved high-precision assessment of drug abuse and drug trafficking cases, provided a scientific basis for urban safety planning and police resource allocation, and improved the accuracy of crime prevention and control strategies.
Smart Images

Figure CN120687760A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of case activity prediction and urban security planning, and in particular relates to a crime factor correlation analysis method taking into account spatial heterogeneity and zero-inflation phenomena. Background Art
[0002] Due to their illegality and concealment, drug crimes often show significant spatial clustering characteristics, and their distribution is closely related to the urban built environment. However, traditional case analysis methods have multiple limitations in revealing the complex relationship between the two. Existing studies often use Poisson regression or ordinary negative binomial models to process crime count data, but these methods ignore the zero inflation phenomenon that is prevalent in the data. The proportion of zero values in drug case data is extremely high, and their generation mechanism can be divided into two categories: structural zero values (such as the objective absence of crime conditions in strictly controlled areas) and random zero values (such as occasional unobserved cases). Traditional models assume that zero values are generated by only a single process, which leads to parameter estimation bias and cannot distinguish the differentiated driving factors of the two types of zero values, thereby affecting the accurate assessment of the impact of environmental factors.
[0003] Furthermore, while the theory of criminal symbiosis suggests that drug use and drug trafficking share a spatially coexisting distribution, existing methods fail to quantify the heterogeneous responses of the two to environmental factors. For example, the same environmental variable (such as the density of entertainment venues) may have a positive impact on drug use but an insignificant or even negative impact on drug trafficking, yet traditional models struggle to capture such differences.
[0004] Crime Preventive Environmental Design (CPTED) theory emphasizes suppressing crime opportunities through spatial optimization, but its application is often limited to the macroscale and lacks detailed modeling of the microscopic built environment. Existing research focuses on macro-level principles such as "natural surveillance" and "enhanced sense of territory," but fails to analyze the local mechanisms by which specific facilities (such as convenience stores and bus stops) drive crime hotspots. Furthermore, CPTED empirical analyses often rely on linear regression models, which struggle to account for the spatial heterogeneity of crime data (the same variable can have opposite effects in different regions) and zero-inflation characteristics, resulting in a lack of scientific basis for policy formulation.
[0005] In summary, the existing technologies have the following key problems: Insufficient zero-inflated data modeling: Traditional counting models cannot distinguish between structural zeros and random zeros, resulting in distorted estimates of the effects of environmental factors; Ignoring spatial heterogeneity: Global models assume that variable effects are spatially constant, making it difficult to reveal the differences in local driving mechanisms; Lack of quantification of crime symbiosis: The spatial symbiosis of drug use and drug trafficking cases has not been converted into a quantitative difference analysis of environmental responses. Therefore, there is an urgent need for an analysis method that takes into account both spatial heterogeneity and zero-inflation phenomena, so as to accurately quantify the differentiated impact of the urban built environment on criminal activities and provide scientific support for crime prevention and control and urban planning. The present invention solves the above technical bottlenecks by integrating the zero-inflated negative binomial model with spatial analysis technology, and achieves a high-precision assessment of the correlation between crime factors. Summary of the Invention
[0006] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology, provide a crime factor correlation analysis method that takes into account spatial heterogeneity and zero inflation phenomenon, solve the technical problem of inaccurate correlation assessment of risk factors and activities in drug abuse and drug trafficking cases in the existing technology, and more accurately quantify the differentiated impact of the urban built environment on drug abuse and drug trafficking case activities, thereby providing a scientific basis for urban security planning, police resource allocation and prevention.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A crime factor correlation analysis method that takes into account spatial heterogeneity and zero inflation phenomena includes the following steps:
[0009] Collecting spatiotemporal data and sociodemographic data on drug use and trafficking cases in each community within the target area; the spatiotemporal data on drug use and trafficking cases includes case type, geographic location, and time information;
[0010] Obtain initial environmental factors related to drug abuse and drug trafficking cases, conduct collinearity diagnosis on the initial environmental factors, screen out initial environmental factors with weak correlation, and perform standardization to obtain environmental factors with weak correlation as built environmental factors;
[0011] The global S-index and local S-index of drug use and drug trafficking cases were calculated based on the spatial point pattern test to analyze the similarity of the spatial distribution of drug use and drug trafficking cases.
[0012] Each community in the target area is divided into grids as analysis units. The number of drug use and trafficking cases, built environment factors, and sociodemographic data are divided into grids. The number of drug use and trafficking cases is used as the modeling object, and a zero-inflated negative binomial model is constructed for each analysis unit. The zero-inflated negative binomial model is expressed as:
[0013]
[0014] Where Γ(·) is the gamma function, which is used to normalize the probability mass function of the negative binomial distribution, y i is the number of drug abuse and drug trafficking cases in the i-th analysis unit, is the probability of structural zero value of the i-th analysis unit, λ i is the expected value of the number of non-zero cases in the ith analysis unit, α is the overdispersion parameter of the negative binomial distribution, when y i = 0, the probability of the caseload being zero is the probability of the zero expansion part Together with the zero value probability of the negative binomial distribution; when y i When > 0, the caseload is generated entirely by the negative binomial part;
[0015] The expectation-maximization algorithm and the Newton-Raphson method were used to iteratively optimize the zero-inflated partial regression coefficient, negative binomial partial regression coefficient, and overdispersion parameter in the zero-inflated negative binomial model until the parameters converged. Based on the converged zero-inflated partial regression coefficient and negative binomial partial regression coefficient, the incidence rate ratio and odds ratio of the explanatory variables between each built environment factor and sociodemographic data were calculated.
[0016] The optimized zero-inflated negative binomial model was used to estimate the regression coefficients of built environment factors;
[0017] The impact of each built environment factor on drug use and trafficking was analyzed based on the regression coefficients of built environment factors and the incidence and odds ratios of explanatory variables between each built environment factor and socio-demographic data. Kernel density estimation was used to generate kernel density distribution maps and visualize hotspot distribution maps of drug use and trafficking cases, based on the geographic location and caseload of drug use and trafficking cases.
[0018] Targeted urban environmental intervention strategies are proposed based on the model results, which include:
[0019] The converged negative binomial partial regression coefficient and overdispersion parameter, zero-inflated partial regression coefficient, hotspot distribution map, global S index, local S index, incidence rate ratio and odds ratio.
[0020] As a preferred technical solution, the analysis of the spatial distribution similarity of drug abuse and drug trafficking cases is specifically as follows:
[0021] The geographic location data of drug abuse cases and drug trafficking cases in the target area are converted into spatial point data respectively, forming two independent spatial point distribution datasets, including the spatial point distribution dataset of drug abuse cases and the spatial point distribution dataset of drug trafficking cases;
[0022] The global S index is calculated by spatial point pattern testing: two sets of independent spatial point distribution data sets are iteratively resampled multiple times, and a% of sample points are extracted in each iteration to generate the reference distribution of drug use and drug trafficking cases; the proportion S of spatial points of drug use cases that fall within the confidence interval of the kernel density estimate of drug trafficking cases is counted. global If S global If the ratio is greater than or equal to the set threshold, it is considered that there is significant global similarity between the spatial point distribution datasets of drug use cases and drug trafficking cases, which supports the theory of criminal symbiosis.
[0023] Calculate the local S index to identify spatial heterogeneity: divide the target area into grids as local analysis units; repeat the above spatial point pattern test in each grid, and generate the local S index S of each grid through multiple iterative resampling. lobal and its corresponding confidence interval; if the global S index S of a grid lobal If the confidence interval of the grid is exceeded, it indicates that there are significant spatial distribution differences in the region, and the driving mechanism needs to be further analyzed in combination with environmental factors;
[0024] A local S-index distribution map was generated, and the hotspot overlapping areas and heterogeneous areas of drug use and drug trafficking cases were compared through spatial visualization to verify the differentiated impact of environmental factors on the two types of cases.
[0025] As a preferred technical solution, the calculation of the incidence ratio and odds ratio of the explanatory variables between each built environment factor and socio-demographic data is specifically as follows:
[0026] Get the joint log-likelihood function for the zero-inflated negative binomial model, expressed as:
[0027]
[0028] Where n is the total number of grids divided in the target area, that is, the total number of analysis units; y i is the number of drug use and drug trafficking cases in the ith analysis unit, β is the regression coefficient of the negative binomial part, γ is the regression coefficient of the zero-inflated part, α is the overdispersion parameter of the negative binomial distribution, is the probability of structural zero value of the i-th analysis unit, λ i is the expected value of the number of non-zero cases in the ith analysis unit, f NB (y i ∣λ i ,α) is the probability mass function of negative binomial; I(y i =0) is the indicator function. When the number of drug abuse and drug trafficking cases y i =0, the value is 1, otherwise it is 0; I(y i >0) is an indicator function, when the number of drug abuse and drug trafficking cases yi When it is greater than 0, the value is 1, otherwise it is 0;
[0029] Introduce latent variable z into each analysis unit i Distinguish between structural zero values and random zero values, and use the expectation maximization algorithm to calculate the latent variable weight w i ;
[0030] The Newton-Raphson method is used to optimize the regression coefficient γ of the zero-inflated part, the regression coefficient β of the negative binomial part, and the overdispersion parameter α;
[0031] Repeatedly iteratively execute the expectation maximization algorithm and the Newton-Raphson method until the parameters converge; the condition for the parameter convergence is that the relative changes of the regression coefficient γ of the zero-inflated part, the regression coefficient β of the negative binomial part, and the overdispersion parameter α are all less than a preset threshold ∈;
[0032] Based on the converged parameters, the incidence rate ratios and odds ratios of the explanatory variables between each built environment factor and socio-demographic data were calculated.
[0033] As a preferred technical solution, the probability of the structural zero value of the i-th analysis unit is The calculation formula is:
[0034]
[0035] Among them, γ T is the transpose of the regression coefficient of the zero-inflated part, is the vector of explanatory variables between built environment factors and socio-demographic data in the case of structural zero value, which is used to quantify the impact of built environment factors on the probability of a region becoming a structural zero value;
[0036] The expected value λ of the number of non-zero cases in the i-th analysis unit i The calculation formula is:
[0037]
[0038] Among them, β T is the transpose of the regression coefficient of the negative binomial part, is the vector of explanatory variables between built environment factors and sociodemographic data under the condition of non-zero caseload, which is used to quantify the marginal effect of built environment factors on the intensity of non-zero caseload;
[0039] The probability mass function f of the negative binomial NB (y i ∣λ i ,α) is expressed as:
[0040]
[0041] As a preferred technical solution, the hidden variable z i Indicates whether the number of drug abuse and drug trafficking cases in the i-th analysis unit is zero due to structural zero or random zero; if it is from structural zero, then z i =1, if it comes from random zero value then z i =0;
[0042] The expectation maximization algorithm is used to calculate the latent variable weight w i Specifically:
[0043] For each analysis unit, if the number of drug use and drug trafficking cases is zero, the weight of the latent variable that the number of drug use and drug trafficking cases in the analysis unit belongs to structural zero is calculated using the expectation maximization algorithm; if the number of drug use and drug trafficking cases is greater than zero, the weight of the latent variable is reset to zero;
[0044] The calculation formula of the latent variable weight is:
[0045]
[0046] Among them, w i is the latent variable weight.
[0047] As a preferred technical solution, the Newton-Raphson method is used to optimize the regression coefficient γ of the zero-inflated part, the regression coefficient β of the negative binomial part, and the overdispersion parameter α, specifically:
[0048] With latent variable weight w i A weighted binary log-likelihood function is constructed for the input, and the regression coefficient γ of the zero-inflated part is optimized using the Newton-Raphson method. The probability of the structural zero value of the explanatory variables between each built environment factor and the socio-demographic data as the zero-inflated part is calculated. the impact of;
[0049] With the number of non-zero cases and the complementary latent variable weight 1-w i A weighted negative binomial log-likelihood function was constructed as the input, and the Newton-Raphson method was used to jointly optimize the regression coefficient β and overdispersion parameter α of the negative binomial part. The marginal effect of the explanatory variables between each built environment factor and socio-demographic data as the negative binomial part on the non-zero case rate was calculated.
[0050] As a preferred technical solution, the Newton-Raphson method is used to optimize the regression coefficient γ of the zero-inflated part as follows:
[0051] With latent variable weight w i Construct a weighted binary log-likelihood function for the input as the objective function to optimize the regression coefficient γ of the zero-inflated part, expressed as:
[0052]
[0053] The Newton-Raphson method is used to iteratively optimize the objective function to obtain the regression coefficient γ of the optimized zero-inflated part. The iterative optimization formula is:
[0054]
[0055] Among them, γ (t+1) is the regression coefficient of the zero-inflated part after the t+1th iteration, γ (t) is the regression coefficient of the zero-expansion part after the tth iteration, H(γ (t) ) is the Hessian matrix of the objective function after the tth iteration; is the gradient of the objective function after the tth iteration.
[0056] As a preferred technical solution, the Newton-Raphson method is used to jointly optimize the regression coefficient β and the overdispersion parameter α of the negative binomial part as follows:
[0057] With non-zero case volume and complementary latent variable weight 1-w i A weighted negative binomial log-likelihood function is constructed for the input as the objective function to optimize the regression coefficient β and overdispersion parameter α of the negative binomial part, which is expressed as:
[0058]
[0059] The Newton-Raphson method is used to iteratively jointly optimize the objective function to obtain the regression coefficient β and overdispersion parameter α of the optimized negative binomial part. The iterative joint optimization formula is:
[0060]
[0061] Among them, β (t+1) is the regression coefficient of the negative binomial part after the t+1th iteration, α (t+1) is the overdispersion parameter of the negative binomial part after the t+1th iteration, β (t) is the regression coefficient of the negative binomial part after the tth iteration, α (t) is the overdispersion parameter of the negative binomial part after the tth iteration, H(β (t) ,α (t) ) is the Hessian matrix of the objective function after the tth iteration; is the gradient of the objective function after the tth iteration.
[0062] As a preferred technical solution, the condition for parameter convergence is described as:
[0063]
[0064] Among them, γ (t+1) is the regression coefficient of the zero-inflated part after the t+1th iteration, γ(t) is the regression coefficient of the zero-inflated part after the tth iteration, β (t+1) is the regression coefficient of the negative binomial part after the t+1th iteration, α (t+1) is the overdispersion parameter of the negative binomial part after the t+1th iteration, β (t) is the regression coefficient of the negative binomial part after the tth iteration, α (t) is the overdispersion parameter of the negative binomial part after the tth iteration;
[0065] The formula for calculating the incidence ratio of the explanatory variables between the built environment factors and the socio-demographic data is:
[0066] IRR k =exp(β k ),
[0067] Among them, β k is the regression coefficient of the explanatory variable between the kth built environment factor and socio-demographic data in the negative binomial part, IRR K It represents the multiple by which the expected incidence rate of non-zero cases is multiplied for every 1 unit increase in the explanatory variable between the k-th built environment factor and socio-demographic data;
[0068] The formula for calculating the odds ratio of the explanatory variables between the built environment factors and the socio-demographic data is:
[0069] OR k =exp(γ k ),
[0070] Among them, γ k is the regression coefficient of the explanatory variable between the kth built environment factor and socio-demographic data in the zero-inflated part, OR K It represents the multiple by which the probability of the k-th built environment factor and the socio-demographic data becoming a structural zero value is multiplied for each 1 unit increase in the explanatory variable.
[0071] As a preferred technical solution, the kernel density estimation is used to generate the kernel density distribution map, specifically:
[0072] Extract geographic location information of drug use and drug trafficking cases, convert it into spatial point data, and unify the coordinate system;
[0073] The Gaussian kernel function is used to calculate the bandwidth parameters of drug abuse and drug trafficking cases based on the Silverman criterion. The formula is:
[0074]
[0075] Where σ is the standard deviation of the spatial distribution of drug use and drug trafficking cases, IQR is the interquartile range, and n is the number of drug use and drug trafficking cases;
[0076] The analysis unit is used as the output raster of the kernel density calculation;
[0077] The kernel density contribution value of all drug use and drug trafficking case points in the neighborhood of each analysis unit center point is calculated using the formula:
[0078]
[0079] Among them, d i is the Euclidean distance between the ith drug use and drug trafficking case point and the center point of the current analysis unit, and K(·) is the Gaussian kernel function;
[0080] Perform maximum-minimum normalization on the kernel density values, map them to the [0,1] interval, and generate a heat map using gradient color scales;
[0081] Overlay the heat map with the spatial distribution data of built environment factors to identify the spatial correlation between built environment factors and case hotspots;
[0082] The spatial coordinates and density levels of hot spots are marked to generate independent kernel density distribution maps and comparison maps of drug use and drug trafficking cases.
[0083] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0084] Compared with the existing technology, the present invention has achieved a multi-dimensional technological breakthrough in the analysis of the correlation between crime factors. Traditional methods use Poisson regression or ordinary negative binomial models, which ignore the zero-inflated characteristics of crime data (such as the mixing of structural zero values and random zero values) and spatial heterogeneity (such as the effect of environmental variables varies with regions), resulting in parameter estimation bias and lack of analysis of local driving mechanisms. The present invention uses the zero-inflated negative binomial model (ZINB) combined with latent variable weights to distinguish the zero-value generation mechanism, and uses the expectation maximization algorithm and the Newton-Raphson method to alternately optimize parameters, significantly improving the model accuracy; at the same time, based on the crime symbiosis theory, it quantifies the spatial distribution similarity of drug use and drug trafficking cases, combines the crime prevention environmental design (CPTED) theory to screen key environmental variables, and analyzes the differentiated impact of environmental factors on the two types of crime through gridded spatial unit division and spatial point pattern test (SPPT). Finally, the kernel density estimation (KDE) is used to visualize the symbiotic hotspots and heterogeneous areas, combining the crime symbiosis law with the CPTED spatial optimization principle to provide a scientific basis for preventive intervention measures. This invention overcomes the limitations of traditional methods in zero-value processing, spatial modeling and theoretical integration, and achieves the coordinated optimization of accurate crime risk assessment and prevention and control strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0086] Figure 1 4 is an overall flow chart of a crime factor correlation analysis method that takes into account spatial heterogeneity and zero inflation phenomenon in an embodiment of the present invention.
[0087] Figure 2 This is a local S-index distribution diagram of the spatial point pattern test of drug abuse and drug trafficking cases in an embodiment of the present invention.
[0088] Figure 3 4 is a kernel density distribution map of drug abuse and drug trafficking cases in an embodiment of the present invention. DETAILED DESCRIPTION
[0089] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0090] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0091] like Figure 1 As shown, this embodiment provides a crime factor correlation analysis method that takes into account spatial heterogeneity and zero inflation, including the following steps:
[0092] S1. Collect spatiotemporal data and sociodemographic data on drug use and trafficking cases in each community within the target area. The spatiotemporal data on drug use and trafficking cases include case type, geographic location, and time information.
[0093] S2. Obtain initial environmental factors related to drug abuse and drug trafficking cases, conduct collinearity diagnosis on the initial environmental factors, screen out initial environmental factors with weak correlation, and perform standardization to obtain environmental factors with weak correlation as built environmental factors.
[0094] For the acquisition of initial environmental factors, the present invention is based on crime geography theories such as the crime symbiosis theory and the crime prevention environment design theory, as well as relevant research literature on drug abuse and drug trafficking cases, combined with the situation of drug abuse and drug trafficking case data acquisition and the community economic environment of the target area for selection. Of course, other ways of obtaining initial environmental factors are also possible, and the present invention will not go into details. When performing collinearity diagnosis, binary variable correlation analysis is used to test the collinearity between the initial environmental factors. Specifically: the Pearson correlation coefficient matrix between all initial environmental factors is calculated, and the threshold is set to |r|<0.7. If the correlation coefficient between a pair of initial environmental factors exceeds the set threshold, it is considered that there is strong collinearity between the two and further screening is required. For strongly correlated pairs of initial environmental factors, priority is given to retaining initial environmental factors that are more strongly associated with the theory of drug abuse and drug trafficking cases (such as those supported by literature based on crime symbiosis theory and CPTED theory), or retaining initial environmental factors with higher data integrity. Calculate the VIF value for the remaining initial environmental factors, and exclude initial environmental factors with VIF≥5 to control multicollinearity. The initial environmental factors with weak correlation after screening were Z-score standardized to eliminate dimensional differences, and the environmental factors with weak correlation were obtained as built environmental factors.
[0095] In this embodiment, based on the number of drug abuse and drug trafficking cases, environmental factor data and census data of a certain city from 2015 to 2019, hotels, inns, convenience stores, shopping malls and supermarkets, entertainment venues, parking lots, bus stops, parks and squares, the proportion of main roads, the proportion of branch roads, the proportion of residential area, the proportion of migrant population, the average night light value and police control points were selected as initial environmental influencing factors. After performing collinearity diagnosis on the data of the initial environmental influencing factors in the area, it was found that the above 14 environmental influencing factors did not have a strong correlation, so these 14 environmental factors were selected as built environment factors.
[0096] S3. Calculate the global S index and local S index of drug use and drug trafficking cases based on spatial point pattern testing, and analyze the similarity of the spatial distribution of drug use and drug trafficking cases.
[0097] The Spatial Point Pattern Test (SPPT) is a nonparametric statistical method used to quantify the similarity between the distributions of two spatial points. Its core concept is to compare the degree of overlap in the spatial distribution of two groups of point events (e.g., drug use and drug trafficking) through resampling techniques. Based on the SPPT, global and local S-indexes are calculated to analyze the spatial similarity of drug use and drug trafficking cases. This aims to validate the theory of criminal symbiosis, which states that the two types of events may be spatially driven by the same environmental factors:
[0098] S3.1. Convert the geographic location data of drug abuse cases and drug trafficking cases in the target area into spatial point data respectively, forming two independent spatial point distribution data sets, including the spatial point distribution data set of drug abuse cases and the spatial point distribution data set of drug trafficking cases.
[0099] S3.2. Calculate the global S index through spatial point pattern testing: Perform multiple iterative resampling on two independent sets of spatial point distribution data sets, extracting a% of sample points each time to generate reference distributions of drug use and drug trafficking cases; calculate the proportion S of spatial points of drug use cases that fall within the confidence interval of the kernel density estimate of drug trafficking cases. global If S global If the ratio is greater than or equal to the set ratio threshold, it is considered that there is significant global similarity between the spatial point distribution datasets of drug use cases and the spatial point distribution datasets of drug trafficking cases, which supports the crime symbiosis theory.
[0100] S3.3. Calculate the local S index to identify spatial heterogeneity: Divide the target area into grids as local analysis units; repeat the above spatial point pattern test in each grid, and generate the local S index S of each grid through multiple iterative resampling. lobal and its corresponding confidence interval; if the global S index S of a grid lobal If the confidence interval of the grid is exceeded, it indicates that there are significant spatial distribution differences in the grid area, and the driving mechanism needs to be further analyzed in combination with environmental factors.
[0101] S3.4. Generate a local S-index distribution map and compare the hotspot overlap areas and heterogeneous areas of drug use and drug trafficking cases through spatial visualization to verify the differentiated impact of environmental factors on the two types of cases.
[0102] In this example, two independent sets of spatial point distribution data sets were resampled 200 times, with 85% of the sample points extracted each time. The global S index of the spatial point pattern test method was set to a ratio threshold of 0.80. If the global S index is greater than or equal to 0.80, the two sets of distributions are considered to have significant global similarity. The local S index distribution diagram of the spatial point pattern test of drug abuse and drug trafficking cases is shown in the figure below. Figure 2 As shown in the figure, it can also be seen that the spatial distribution of drug abuse and drug trafficking cases is relatively similar.
[0103] S4. Divide each community in the target area into a grid as an analysis unit. Assign drug use and trafficking caseloads, built environment factors, and sociodemographic data to each grid. Use drug use and trafficking caseloads as the modeling object and construct a zero-inflated negative binomial model for each analysis unit. The zero-inflated negative binomial model (ZINB) is expressed as:
[0104]
[0105] Where Γ(·) is the gamma function, which is used to normalize the probability mass function, y i is the number of drug abuse and drug trafficking cases in the i-th analysis unit, is the probability of structural zero value of the i-th analysis unit, λ i is the expected value of the number of non-zero cases in the ith analysis unit, α is the overdispersion parameter of the negative binomial distribution; when y i = 0, the probability of the caseload being zero is the probability of the zero expansion part Together with the zero value probability of the negative binomial distribution; when y i >0, the caseload is entirely generated by the negative binomial part. In this embodiment, each community in the target area is divided into a grid of 150*150, and the length and width of the grid are in meters. The reason for selecting the zero-inflated negative binomial model (ZINB) is that it can effectively solve the problems of zero inflation, overdispersion and spatial heterogeneity in crime data. The model quantifies the impact of environmental factors on the structural zero probability through the zero-inflated part, and analyzes the driving mechanism of the non-zero caseload through the negative binomial part, thereby improving the accuracy of parameter estimation, enhancing explanatory power, and supporting the formulation of refined crime intervention strategies based on hotspot distribution and key environmental factors.
[0106] S5. The expectation-maximization algorithm and the Newton-Raphson method are used to iteratively optimize the zero-inflated partial regression coefficient, the negative binomial partial regression coefficient, and the overdispersion parameter in the zero-inflated negative binomial model until the parameters converge. Based on the converged zero-inflated partial regression coefficient and the negative binomial partial regression coefficient, the incidence rate ratio and the odds ratio of the explanatory variables between each built environment factor and the socio-demographic data are calculated to quantify their marginal effects on the non-zero case incidence rate and the structural zero value probability.
[0107] The model parameters need to be optimized because the present invention introduces hidden variables into the zero-inflated negative binomial model, and the model itself contains complex parameters. It is difficult to directly maximize the joint likelihood function, and it is necessary to iteratively approximate the optimal solution. Therefore, the present invention uses the expectation maximization algorithm (EM) and the Newton-Raphson method to iteratively and alternately optimize the inflated partial regression coefficient, the negative binomial partial regression coefficient and the over-dispersion parameter in the ZINB. Among them, the EM algorithm: calculates the weights of the hidden variables (distinguishing the zero value type) through E steps, and optimizes the parameters through M steps to solve the missing data problem caused by the hidden variables and ensure stable convergence. The Newton-Raphson method: uses the gradient and the Hessian matrix to converge quickly, is suitable for explicit parameter optimization, and is more efficient when approaching the optimal solution than the gradient descent method. The process of iterative optimization is:
[0108] S5.1. First, obtain the joint log-likelihood function of the zero-inflated negative binomial model, expressed as:
[0109]
[0110] Where n is the total number of grids divided in the target area, that is, the total number of analysis units; y i is the number of drug use and drug trafficking cases in the ith analysis unit, β is the regression coefficient of the negative binomial part, γ is the regression coefficient of the zero-inflated part, α is the overdispersion parameter of the negative binomial distribution, is the probability of structural zero value of the i-th analysis unit, λ i is the expected value of the number of non-zero cases in the ith analysis unit, f NB (y i ∣λ i ,α) is the probability mass function of negative binomial; I(y i =0) is the indicator function. When the number of drug abuse and drug trafficking cases y i =0, the value is 1, otherwise it is 0; I(y i >0) is an indicator function, when the number of drug abuse and drug trafficking cases y i When it is greater than 0, the value is 1, otherwise it is 0.
[0111] Furthermore, the probability of structural zero value of the i-th analysis unit is The calculation formula is:
[0112]
[0113] Among them, γ T is the transpose of the regression coefficient of the zero-inflated part, is the vector of explanatory variables between built environment factors and socio-demographic data in the case of structural zero value, which is used to quantify the impact of built environment factors on the probability of a region becoming a structural zero value;
[0114] The expected value of the number of non-zero cases in the i-th analysis unit λ i The calculation formula is:
[0115]
[0116] Among them, β T is the transpose of the regression coefficient of the negative binomial part, is the vector of explanatory variables between built environment factors and sociodemographic data under the condition of non-zero caseload, which is used to quantify the marginal effect of built environment factors on the intensity of non-zero caseload;
[0117] The probability mass function f of the negative binomial NB (y i ∣λ i ,α) is expressed as:
[0118]
[0119] S5.2. Next, introduce the latent variable z into each analysis unit. i Distinguish between structural zero values and random zero values, and use the expectation maximization algorithm to calculate the latent variable weight w i .
[0120] Furthermore, the latent variable z i Indicates whether the number of drug abuse and drug trafficking cases in the i-th analysis unit is zero due to structural zero or random zero; if it is from structural zero, then z i =1, if it comes from random zero value then z i =0;
[0121] The expectation maximization algorithm is used to calculate the latent variable weight w i Specifically:
[0122] For each analysis unit, if the number of drug use and drug trafficking cases is zero, the expectation maximization algorithm is used to calculate the latent variable weight of the structural zero value of the drug use and drug trafficking case volume of the analysis unit; if the number of drug use and drug trafficking cases is greater than zero, the latent variable weight is reset to zero; therefore, the calculation formula of the latent variable weight is:
[0123]
[0124] Among them, w i is the latent variable weight (the expected value of the zero-inflated probability). This step solves the zero-inflated problem by quantifying the generation mechanism of zero values through probability weights.
[0125] S5.3. Subsequently, the Newton-Raphson method is used to optimize the parameters in two parts, namely, optimizing the regression coefficient γ of the zero-inflated part and the regression coefficient β and overdispersion parameter α of the negative binomial part respectively.
[0126] Furthermore, the optimization steps are:
[0127] S5.3.1. Parameter optimization of zero-expansion part: taking latent variable weight w i A weighted binary log-likelihood function was constructed for the input, and the regression coefficient γ of the zero-inflated part was optimized using the Newton-Raphson method. The impact of the explanatory variables between each built environment factor and socio-demographic data as the zero-inflated part on the probability of structural zero value was calculated.
[0128] More specifically, the zero-inflated parameter optimization process is:
[0129] With latent variable weight w i Construct a weighted binary log-likelihood function for the input as the objective function to optimize the regression coefficient γ of the zero-inflated part, expressed as:
[0130]
[0131] The Newton-Raphson method is used to iteratively optimize the objective function to obtain the regression coefficient γ of the optimized zero-inflated part. The iterative optimization formula is:
[0132]
[0133] Among them, γ (t+1) is the regression coefficient of the zero-inflated part after the t+1th iteration, γ (t) is the regression coefficient of the zero-expansion part after the tth iteration, H(γ (t) ) is the Hessian matrix of the objective function after the tth iteration; is the gradient of the objective function after the tth iteration.
[0134] In the zero-expansion part, the gradient is used to measure the sensitivity of the log-likelihood function to the change of the parameter γ, indicating the direction of parameter update to maximize the likelihood function. The gradient formula is:
[0135]
[0136] Among them, the vector X of explanatory variables between each built environment factor and socio-demographic data is i Refers to the built environment factors and socio-demographic variables after standardized processing in the i-th analysis unit.
[0137] The Hessian matrix is the second-order derivative matrix of the objective function, which is used to quantify the curvature of the log-likelihood function in the parameter space, thereby adjusting the step size of the parameter update to improve the convergence efficiency; the Hessian matrix is expressed as:
[0138]
[0139] in, is the weight term of the second-order derivative of logistic regression; X i X i T It is the outer product matrix of explanatory variables between each built environment factor and socio-demographic data.
[0140] S5.3.2. Negative binomial partial parameter optimization: non-zero case volume and complementary latent variable weight 1-w i A weighted negative binomial log-likelihood function was constructed as the input, and the Newton-Raphson method was used to jointly optimize the regression coefficient β and overdispersion parameter α of the negative binomial part. The marginal effect of the explanatory variables between each built environment factor and socio-demographic data as the negative binomial part on the non-zero case rate was calculated.
[0141] More specifically, the negative binomial parameter optimization process is:
[0142] With the number of non-zero cases and the complementary latent variable weight 1-w iA weighted negative binomial log-likelihood function is constructed for the input as the objective function to optimize the regression coefficient β and overdispersion parameter α of the negative binomial part, which is expressed as:
[0143]
[0144] The Newton-Raphson method is used to iteratively jointly optimize the objective function to obtain the regression coefficient β and overdispersion parameter α of the optimized negative binomial part. The iterative joint optimization formula is:
[0145]
[0146] Among them, β (t+1) is the regression coefficient of the negative binomial part after the t+1th iteration, α (t+1) is the overdispersion parameter of the negative binomial part after the t+1th iteration, β (t) is the regression coefficient of the negative binomial part after the tth iteration, α (t) is the overdispersion parameter of the negative binomial part after the tth iteration, H(β (t) ,α (t) ) is the Hessian matrix of the objective function after the tth iteration; is the gradient of the objective function after the tth iteration.
[0147] For the gradient in the negative binomial part, it is divided into two parts, one of which is the partial derivative of the regression coefficient β of the negative binomial part:
[0148]
[0149] The other part is the partial derivative of the overdispersion parameter α of the negative binomial part:
[0150]
[0151] where ψ(·) is the digamma function:
[0152]
[0153] z represents the parameter of the Gamma function, specifically y i +α -1 or α -1 , used for the normalized calculation of the probability mass function; d is the differential operator, which represents the derivative operation of the function; d z It is the differential of the input variable z, that is, the first or second derivative of the input white energy z of the Gamma function.
[0154] For the Hessian matrix of the negative binomial part, it contains the following sub-blocks:
[0155] The second derivative of the regression coefficient β with one part being the negative binomial part:
[0156]
[0157] The other part is the second derivative of the overdispersion parameter α of the negative binomial part:
[0158]
[0159] where ψ′(·) is the Trigamma function:
[0160]
[0161] Another part is the cross derivative (the second-order derivative of β and α):
[0162]
[0163] S5.4. Repeat the iterative execution of the expectation maximization algorithm and the Newton-Raphson method until the parameters converge; the condition for parameter convergence is that the relative changes of the regression coefficient γ of the zero-inflated part, the regression coefficient β of the negative binomial part, and the overdispersion parameter α are all less than the preset threshold ∈.
[0164] Furthermore, the condition for parameter convergence can be described as:
[0165]
[0166] Among them, γ (t+1) is the regression coefficient of the zero-inflated part after the t+1th iteration, γ (t) is the regression coefficient of the zero-inflated part after the tth iteration, β (t+1) is the regression coefficient of the negative binomial part after the t+1th iteration, α (t+1) is the overdispersion parameter of the negative binomial part after the t+1th iteration, β (t) is the regression coefficient of the negative binomial part after the tth iteration, α (t) is the overdispersion parameter of the negative binomial part after the tth iteration.
[0167] S5.5. Based on the converged parameters, calculate the incidence rate ratio and odds ratio of the explanatory variables between each built environment factor and socio-demographic data.
[0168] Furthermore, based on the converged parameters, the incidence rate ratios (IRRs) and odds ratios (ORs) of the explanatory variables between each built environment factor and the sociodemographic data were calculated. In the zero-inflated negative binomial model, the incidence rate ratios and odds ratios correspond to the negative binomial part and the zero-inflated part of the model, respectively, and are used to quantify the differential impact of the explanatory variables on the number of cases. The calculation formula for the incidence rate ratios (IRRs) of the explanatory variables between each built environment factor and the sociodemographic data is:
[0169] IRR k =exp(β k ),
[0170] Among them, β k is the regression coefficient of the explanatory variable between the kth built environment factor and socio-demographic data in the negative binomial part, IRR k It represents the multiple by which the expected incidence rate of non-zero cases is multiplied for every 1 unit increase in the explanatory variable between the k-th built environment factor and socio-demographic data.
[0171] The odds ratio (OR) calculation formula between each built environment factor and socio-demographic data explanatory variables is:
[0172] OR k =exp(γ k ),
[0173] Among them, γ k is the regression coefficient of the explanatory variable between the kth built environment factor and socio-demographic data in the zero-inflated part, OR k It represents the multiple by which the probability of the k-th built environment factor and the socio-demographic data becoming a structural zero value is multiplied for each 1 unit increase in the explanatory variable.
[0174] S6. Use the optimized zero-inflated negative binomial model to estimate the regression coefficients of built environment factors.
[0175] In this embodiment, the regression coefficients of 14 built environment factors are estimated based on the optimized zero-inflated negative binomial model. The results are shown in Table 1 below:
[0176] Table 1 Estimation results of built environment factors by the optimized zero-inflated negative binomial model
[0177]
[0178] Note: *** P < 0.001; ** P < 0.01; * P<0.05, P is the significance level.
[0179] In the table, constant is the intercept term of the model, Beta is the regression coefficient β of the negative binomial component, and Gamma is the regression coefficient γ of the zero-inflated component. The table shows that hotels, shopping malls, supermarkets, and entertainment venues have significant positive effects on both drug use and drug trafficking, increasing the likelihood of both. Convenience stores have a significant promoting effect on drug use but no significant effect on drug trafficking. Furthermore, these environmental factors (such as convenience stores and shopping malls) significantly reduce the probability of both drug use and drug trafficking occurring within a grid by reducing the structural zero probability, negatively impacting the zero-case state. On the other hand, parking lots inhibit drug use but increase the likelihood of drug use, but have no effect on drug trafficking. The proportion of residential area has no significant effect on either drug use or drug trafficking, but increases the likelihood of both. The proportion of main roads favors drug use but not drug trafficking, having no significant effect on non-drug use but increasing the likelihood of drug trafficking. In contrast, the proportion of branch roads inhibits both drug use and drug trafficking but increases the likelihood of both.
[0180] S7. The impact of each built environment factor on drug use and trafficking was analyzed based on the regression coefficients of built environment factors and the incidence and odds ratios of explanatory variables between each built environment factor and sociodemographic data. Kernel density estimation (KDE) was used to generate kernel density distribution maps and visualize the hotspot distribution maps of drug use and trafficking cases based on the geographical location and case volume of drug use and trafficking cases.
[0181] Furthermore, the specific steps of using kernel density estimation to generate a kernel density distribution map are as follows:
[0182] S7.1. Extract geographic location information of drug use and drug trafficking cases, convert it into spatial point data, and unify the coordinate system.
[0183] S7.2. Use the Gaussian kernel function and the Silverman criterion to calculate the bandwidth parameter for drug use and drug trafficking cases. The formula is:
[0184]
[0185] Where σ is the standard deviation of the spatial distribution of drug use and drug trafficking cases, IQR is the interquartile range, and n is the number of drug use and drug trafficking cases.
[0186] S7.3. Use the analysis unit as the output raster of the kernel density calculation.
[0187] S7.4. For each analysis unit center point, calculate the kernel density contribution of all drug use and drug trafficking cases in its neighborhood using the following formula:
[0188]
[0189] Among them, d i is the Euclidean distance between the ith drug use and drug trafficking case and the center of the current analysis unit, and K(·) is the Gaussian kernel function.
[0190] S7.5. Perform maximum-minimum normalization on the kernel density values, map them to the [0, 1] interval, and generate a heat map using a gradient color scale.
[0191] S7.6. Overlay the heat map with the spatial distribution data of built environment factors to identify the spatial correlation between built environment factors and case hotspots.
[0192] S7.7. Mark the spatial coordinates and density levels of hotspot areas to generate independent kernel density distribution maps and comparison maps of drug use and drug trafficking cases.
[0193] The kernel density distribution diagram of drug abuse and drug trafficking cases obtained in this embodiment is as follows: Figure 3 As shown in the figure, drug use and drug trafficking cases are clearly clustered spatially, forming an "inverted V"-shaped spatial distribution. This pattern extends in a belt-like pattern along the northwest-southwest and northwest-southeast directions. Although the spatial distribution is generally similar, there are also significant differences. For example, while drug use rates are high in Districts C4 and F4, the drug trafficking hotspot in District C4 has shifted southeastward, and the high-incidence area in District F4 has transitioned to a moderate-incidence area.
[0194] S8. Propose targeted urban environmental intervention strategies based on model results. Model results include: converged negative binomial partial regression coefficients and overdispersion parameters, zero-inflated partial regression coefficients, hotspot distribution maps, global S-index, local S-index, incidence rate ratios, and odds ratios.
[0195] In summary, the zero-inflated negative binomial model constructed in this invention can more accurately assess the correlation between case risk and activities by taking into account spatial heterogeneity and zero-inflation phenomena, and can further more accurately quantify the differentiated impact of the urban built environment on case activities, thereby providing a scientific basis for urban security planning, police resource allocation and prevention.
[0196] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.
[0197] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0198] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A crime factor correlation analysis method that takes into account spatial heterogeneity and zero inflation, characterized by: The steps include: Collecting spatiotemporal data and sociodemographic data on drug use and trafficking cases in each community within the target area; the spatiotemporal data on drug use and trafficking cases includes case type, geographic location, and time information; Obtain initial environmental factors related to drug abuse and drug trafficking cases, conduct collinearity diagnosis on the initial environmental factors, screen out initial environmental factors with weak correlation, and perform standardization to obtain environmental factors with weak correlation as built environmental factors; The global S-index and local S-index of drug use and drug trafficking cases were calculated based on the spatial point pattern test to analyze the similarity of the spatial distribution of drug use and drug trafficking cases. Each community in the target area is divided into grids as analysis units. The number of drug use and trafficking cases, built environment factors, and sociodemographic data are divided into grids. The number of drug use and trafficking cases is used as the modeling object, and a zero-inflated negative binomial model is constructed for each analysis unit. The zero-inflated negative binomial model is expressed as: Where Γ(·) is the gamma function, which is used to normalize the probability mass function of the negative binomial distribution, y i is the number of drug abuse and drug trafficking cases in the i-th analysis unit, is the probability of structural zero value of the i-th analysis unit, λ i is the expected value of the number of non-zero cases in the ith analysis unit, α is the overdispersion parameter of the negative binomial distribution, when y i = 0, the probability of the caseload being zero is the probability of the zero expansion part Together with the zero value probability of the negative binomial distribution; when y i When > 0, the caseload is generated entirely by the negative binomial part; The expectation-maximization algorithm and the Newton-Raphson method were used to iteratively optimize the zero-inflated partial regression coefficient, negative binomial partial regression coefficient, and overdispersion parameter in the zero-inflated negative binomial model until the parameters converged. Based on the converged zero-inflated partial regression coefficient and negative binomial partial regression coefficient, the incidence rate ratio and odds ratio of the explanatory variables between each built environment factor and sociodemographic data were calculated. The optimized zero-inflated negative binomial model was used to estimate the regression coefficients of built environment factors; The impact of each built environment factor on drug use and trafficking was analyzed based on the regression coefficients of built environment factors and the incidence and odds ratios of explanatory variables between each built environment factor and socio-demographic data. Kernel density estimation was used to generate kernel density distribution maps and visualize hotspot distribution maps of drug use and trafficking cases, based on the geographic location and caseload of drug use and trafficking cases. Targeted urban environmental intervention strategies are proposed based on the model results, which include: The converged negative binomial partial regression coefficient and overdispersion parameter, zero-inflated partial regression coefficient, hotspot distribution map, global S index, local S index, incidence rate ratio and odds ratio.
2. The crime factor correlation analysis method taking into account spatial heterogeneity and zero inflation phenomenon according to claim 1 is characterized in that: The analysis of the spatial distribution similarity of drug abuse and drug trafficking cases is as follows: The geographic location data of drug abuse cases and drug trafficking cases in the target area are converted into spatial point data respectively, forming two independent spatial point distribution datasets, including the spatial point distribution dataset of drug abuse cases and the spatial point distribution dataset of drug trafficking cases; The global S index is calculated by spatial point pattern testing: two sets of independent spatial point distribution data sets are iteratively resampled multiple times, and a% of sample points are extracted in each iteration to generate the reference distribution of drug use and drug trafficking cases; the proportion S of spatial points of drug use cases that fall within the confidence interval of the kernel density estimate of drug trafficking cases is counted. global If S global If the ratio is greater than or equal to the set threshold, it is considered that there is significant global similarity between the spatial point distribution datasets of drug use cases and drug trafficking cases, which supports the theory of criminal symbiosis. Calculate the local S index to identify spatial heterogeneity: divide the target area into grids as local analysis units; repeat the above spatial point pattern test in each grid, and generate the local S index S of each grid through multiple iterative resampling. lobal and its corresponding confidence interval; if the global S index S of a grid lobal If the confidence interval of the grid is exceeded, it indicates that there are significant spatial distribution differences in the region, and the driving mechanism needs to be further analyzed in combination with environmental factors; A local S-index distribution map was generated, and the hotspot overlapping areas and heterogeneous areas of drug use and drug trafficking cases were compared through spatial visualization to verify the differentiated impact of environmental factors on the two types of cases.
3. The crime factor correlation analysis method taking into account spatial heterogeneity and zero inflation phenomenon according to claim 1 is characterized in that: The calculation of the incidence rate ratio and odds ratio of the explanatory variables between each built environment factor and socio-demographic data is as follows: Get the joint log-likelihood function for the zero-inflated negative binomial model, expressed as: Where n is the total number of grids divided in the target area, that is, the total number of analysis units; y i is the number of drug use and drug trafficking cases in the ith analysis unit, β is the regression coefficient of the negative binomial part, γ is the regression coefficient of the zero-inflated part, α is the overdispersion parameter of the negative binomial distribution, is the probability of structural zero value of the i-th analysis unit, λ i is the expected value of the number of non-zero cases in the ith analysis unit, f NB (y i ∣λ i ,α) is the probability mass function of negative binomial; I(y i =0) is the indicator function. When the number of drug abuse and drug trafficking cases y i =0, the value is 1, otherwise it is 0; I(y i >0) is an indicator function, when the number of drug abuse and drug trafficking cases y i When it is greater than 0, the value is 1, otherwise it is 0; Introduce latent variable z into each analysis unit i Distinguish between structural zero values and random zero values, and use the expectation maximization algorithm to calculate the latent variable weight w i ; The Newton-Raphson method is used to optimize the regression coefficient γ of the zero-inflated part, the regression coefficient β of the negative binomial part, and the overdispersion parameter α; Repeatedly iteratively execute the expectation maximization algorithm and the Newton-Raphson method until the parameters converge; the condition for the parameter convergence is that the relative changes of the regression coefficient γ of the zero-inflated part, the regression coefficient β of the negative binomial part, and the overdispersion parameter α are all less than a preset threshold ∈; Based on the converged parameters, the incidence rate ratios and odds ratios of the explanatory variables between each built environment factor and socio-demographic data were calculated.
4. The crime factor correlation analysis method taking into account spatial heterogeneity and zero inflation according to claim 3 is characterized in that: The probability that the i-th analysis unit is structurally zero The calculation formula is: Among them, γ T is the transpose of the regression coefficient of the zero-inflated part, is the vector of explanatory variables between built environment factors and socio-demographic data in the case of structural zero value, which is used to quantify the impact of built environment factors on the probability of a region becoming a structural zero value; The expected value λ of the number of non-zero cases in the i-th analysis unit i The calculation formula is: Among them, β T is the transpose of the regression coefficient of the negative binomial part, is the vector of explanatory variables between built environment factors and sociodemographic data under the condition of non-zero caseload, which is used to quantify the marginal effect of built environment factors on the intensity of non-zero caseload; The probability mass function f of the negative binomial NB (y i ∣λ i ,α) is expressed as:
5. The crime factor correlation analysis method taking into account spatial heterogeneity and zero inflation phenomenon according to claim 3 is characterized in that: The hidden variable z i Indicates whether the number of drug abuse and drug trafficking cases in the i-th analysis unit is zero due to structural zero or random zero; if it is from structural zero, then z i =1, if it comes from random zero value then z i =0; The expectation maximization algorithm is used to calculate the latent variable weight w i Specifically: For each analysis unit, if the number of drug use and drug trafficking cases is zero, the weight of the latent variable that the number of drug use and drug trafficking cases in the analysis unit belongs to structural zero is calculated using the expectation maximization algorithm; if the number of drug use and drug trafficking cases is greater than zero, the weight of the latent variable is reset to zero; The calculation formula of the latent variable weight is: Among them, w i is the latent variable weight.
6. The crime factor correlation analysis method taking into account spatial heterogeneity and zero inflation according to claim 3, characterized in that: The Newton-Raphson method is used to optimize the regression coefficient γ of the zero-inflated part, the regression coefficient β of the negative binomial part, and the overdispersion parameter α, specifically: With latent variable weight w i A weighted binary log-likelihood function is constructed for the input, and the regression coefficient γ of the zero-inflated part is optimized using the Newton-Raphson method. The probability of the structural zero value of the explanatory variables between each built environment factor and the socio-demographic data as the zero-inflated part is calculated. the impact of; With the number of non-zero cases and the complementary latent variable weight 1-w i A weighted negative binomial log-likelihood function was constructed as the input, and the Newton-Raphson method was used to jointly optimize the regression coefficient β and overdispersion parameter α of the negative binomial part. The marginal effect of the explanatory variables between each built environment factor and socio-demographic data as the negative binomial part on the non-zero case rate was calculated.
7. The crime factor correlation analysis method taking into account spatial heterogeneity and zero inflation according to claim 6, characterized in that: The Newton-Raphson method is used to optimize the regression coefficient γ of the zero-inflated part: With latent variable weight w i Construct a weighted binary log-likelihood function for the input as the objective function to optimize the regression coefficient γ of the zero-inflated part, expressed as: The Newton-Raphson method is used to iteratively optimize the objective function to obtain the regression coefficient γ of the optimized zero-inflated part. The iterative optimization formula is: Among them, γ (t+1) is the regression coefficient of the zero-inflated part after the t+1th iteration, γ (t) is the regression coefficient of the zero-expansion part after the tth iteration, H(γ (t) ) is the Hessian matrix of the objective function after the tth iteration; is the gradient of the objective function after the tth iteration.
8. The crime factor correlation analysis method taking into account spatial heterogeneity and zero inflation according to claim 6, characterized in that: The Newton-Raphson method is used to jointly optimize the regression coefficient β and the overdispersion parameter α of the negative binomial part as follows: With the number of non-zero cases and the complementary latent variable weight 1-w i A weighted negative binomial log-likelihood function is constructed for the input as the objective function to optimize the regression coefficient β and overdispersion parameter α of the negative binomial part, which is expressed as: The Newton-Raphson method is used to iteratively jointly optimize the objective function to obtain the regression coefficient β and overdispersion parameter α of the optimized negative binomial part. The iterative joint optimization formula is: Among them, β (t+1) is the regression coefficient of the negative binomial part after the t+1th iteration, α (t+1) is the overdispersion parameter of the negative binomial part after the t+1th iteration, β (t) is the regression coefficient of the negative binomial part after the tth iteration, α (t) is the overdispersion parameter of the negative binomial part after the tth iteration, H(β (t) ,α (t) ) is the Hessian matrix of the objective function after the tth iteration; is the gradient of the objective function after the tth iteration.
9. The crime factor correlation analysis method taking into account spatial heterogeneity and zero inflation according to claim 3, characterized in that: The condition for the parameter convergence is described as: Among them, γ (t+1) is the regression coefficient of the zero-inflated part after the t+1th iteration, γ (t) is the regression coefficient of the zero-inflated part after the tth iteration, β (t+1) is the regression coefficient of the negative binomial part after the t+1th iteration, α (t+1) is the overdispersion parameter of the negative binomial part after the t+1th iteration, β (t) is the regression coefficient of the negative binomial part after the tth iteration, α (t) is the overdispersion parameter of the negative binomial part after the tth iteration; The formula for calculating the incidence ratio of the explanatory variables between the built environment factors and the socio-demographic data is: IRR k =exp(β k ), Among them, β k is the regression coefficient of the explanatory variable between the kth built environment factor and socio-demographic data in the negative binomial part, IRR K It represents the multiple by which the expected incidence rate of non-zero cases is multiplied for every 1 unit increase in the explanatory variable between the k-th built environment factor and socio-demographic data; The formula for calculating the odds ratio of the explanatory variables between the built environment factors and the socio-demographic data is: OR k =exp(γ k ), Among them, γ k is the regression coefficient of the explanatory variable between the kth built environment factor and socio-demographic data in the zero-inflated part, OR K It represents the multiple by which the probability of the k-th built environment factor and the socio-demographic data becoming a structural zero value is multiplied for each 1 unit increase in the explanatory variable.
10. The crime factor correlation analysis method taking into account spatial heterogeneity and zero inflation phenomenon according to claim 1, characterized in that: The kernel density estimation is used to generate the kernel density distribution map, specifically: Extract geographic location information of drug use and drug trafficking cases, convert it into spatial point data, and unify the coordinate system; The Gaussian kernel function is used to calculate the bandwidth parameters of drug abuse and drug trafficking cases based on the Silverman criterion. The formula is: Where σ is the standard deviation of the spatial distribution of drug use and drug trafficking cases, IQR is the interquartile range, and n is the number of drug use and drug trafficking cases; The analysis unit is used as the output raster of the kernel density calculation; The kernel density contribution value of all drug use and drug trafficking case points in the neighborhood of each analysis unit center point is calculated using the formula: Among them, d i is the Euclidean distance between the ith drug use and drug trafficking case point and the center point of the current analysis unit, and K(·) is the Gaussian kernel function; Perform maximum-minimum normalization on the kernel density values, map them to the [0,1] interval, and generate a heat map using gradient color scales; Overlay the heat map with the spatial distribution data of built environment factors to identify the spatial correlation between built environment factors and case hotspots; The spatial coordinates and density levels of hot spots are marked to generate independent kernel density distribution maps and comparison maps of drug use and drug trafficking cases.